Tools Reference
colabhive-mcp exposes two families of tools, with two naming conventions that live side by side:
- Manifest / action tools — the hive's models and endpoints as invocable tools. Their slug is
the endpoint name (
endpoint_namein the DB), kebab-case, e.g.qwen-2.5-7b-instruct-public,embeddings-public,web-fetch-public. This list is 100% DB-driven — it is whatever your account can see, fetched live fromGET /api/builder/v1/mcp/manifest. Plus two static platform operations (training.merge,training.retrain_on). - Built-in control-plane tools — 19 operational tools (snake_case) for account preferences, data-protection preflight, datasets, training, and model management. These are baked into the server, not derived from the manifest.
server.list_tools() returns the filtered manifest tools plus the 19 built-ins.
Because the manifest is per-account and dynamic, this page does not hardcode a catalog of model slugs. For the authoritative live list, call
GET /api/builder/v1/mcp/manifest,GET /api/builder/v1/actions?visibility=public, or browse the Model Catalog. The example slugs below are illustrative.
Platform operations (kind=operation)
Two static tools, injected into every authenticated account's manifest (before the model tools) so
agents can build new capabilities: merge trained models/adapters and retrain on top of any model.
Both create jobs that count against your training quota, so they carry costHint: "training_quota"
and sideEffects: ["creates_job", "creates_model_version", "consumes_training_quota"].
| Slug | Purpose | Backing endpoint | latencyClass |
|---|---|---|---|
training.merge | Merge N trained models/adapters into a base → a new servable model_version (weight arithmetic, CPU). | POST /training/merges | minutes |
training.retrain_on | Fine-tune on top of an existing trained/merged model_version (or HF/job/storage) as the base. | POST /training/runs | long |
Inputs reference models by ArtifactRef
(hf | model_version | job | storage).
training.merge input (summary):
{
"name": "qwen-domain-base",
"base": { "type": "hf", "repo_id": "Qwen/Qwen2.5-7B-Instruct" },
"method": "adapter_merge",
"adapters": [ { "type": "model_version", "version_id": "..." } ],
"sources": [],
"weights": [],
"precision": "bf16",
"hardware_preference": "cpu_only"
}
method ∈ adapter_merge (default) | slerp | ties | dare. Only name is required.
training.retrain_on input (summary):
{
"job_name": "domain-expert-v2",
"model_config_id": "...",
"dataset_id": "...",
"base": { "type": "model_version", "version_id": "..." },
"parent_job_id": null,
"hyperparameters": { "learning_rate": 0.0002, "epochs": 3 }
}
Required: job_name, model_config_id, dataset_id, base.
These deliver the generic model flywheel — merge capabilities, retrain-on-top, repeat. See the flywheel concept and the merge & retrain how-to.
Built-in control-plane tools
These 20 tools (snake_case) wrap the platform's operational surface — account preferences, data-protection reads, memory status, datasets, training runs, and model/endpoint management. They are always present, independent of the manifest. Agents such as SuperClaw / OpenCode use exactly these names.
Account preferences
| Tool | Purpose | Required args |
|---|---|---|
get_account_preferences | Read the authenticated account's preferences, including the Share Hive opt-in state. | — |
set_share_hive_opt_in | Explicitly enable or disable Share Hive spare-capacity lending. Disabled by default; changing it requires the account owner role. | enabled |
Memory
| Tool | Purpose | Required args |
|---|---|---|
get_memory_status | Read whether ColabHive memory is on for the authenticated account and whether a task submitted now would carry it. Answers enabled, ready and, when it is not usable, a detail naming why. | — |
The credential selects the account, so the tool takes no arguments: an agent cannot ask about an account it does not hold a key for. There is no MCP tool to turn memory on, to name a namespace or to read what memory contains — the platform owns the binding and the caller names nothing. See Memory.
Data protection
| Tool | Purpose | Required args |
|---|---|---|
get_effective_data_protection | Read the effective node-eligibility value, binding_level, full resolution trace, and can_relax_to. | task_exec_id, or workload_type + workload_id |
preflight_data_protection | Read whether eligible capacity exists under the current or a hypothetical policy. Returns relaxation_hint as structured data and never changes policy. | workload_type, workload_id |
The MCP surface deliberately has no account-policy mutation tool. Raising an entire account's placement ceiling requires an explicitly authorized API or Console action.
Discovery & inference
| Tool | Purpose | Required args |
|---|---|---|
list_endpoints | List invocable inference endpoints (base models + your deployed models) with readiness. | — (optional: include_readiness, task_type, search, limit) |
get_endpoint | Get an endpoint's details and readiness (warm/cached/cold). | endpoint_id |
run_inference | Run sync inference on an endpoint. input is the model's input object (for LLMs: {messages, tools?, tool_choice?, max_tokens?}; for specialists, the endpoint's schema). Returns the model output (OpenAI-shaped for LLMs, including tool_calls and, on a thinking model, reasoning — see Thinking models). | endpoint_id, input |
list_models | List your trained/registered models. | — |
Training lifecycle
| Tool | Purpose | Required args |
|---|---|---|
list_trainable_models | List base model configs available to fine-tune/train (id, name, category, framework). | — |
get_model_schema | Get the hyperparameter schema for a trainable model config. | model_config_id |
list_datasets | List datasets registered for your account. | — |
create_dataset | Register a new dataset (metadata). | name |
create_training_run | Start a real training run. operating_mode ∈ economy | balanced | performance. Returns the run id. | job_name, model_config_id, dataset_id |
list_training_runs | List your training runs with status. | — (optional: limit) |
get_training_run | Get a run's status/details. | run_id |
get_training_metrics | Get metrics (loss curves, eval scores, feature importance when produced). | run_id |
get_training_logs | Get logs for a run. | run_id |
cancel_training_run | Cancel a running training run. | run_id |
register_for_inference | Deploy a finished run as an inference endpoint (it becomes an invocable model tool). | run_id |
Model & endpoint tools (manifest-driven)
Every model/endpoint your account can see becomes a tool whose slug is its endpoint_name. The
kind is computed server-side (never hardcoded) from the endpoint's category/task:
kind | What it is | Example slugs (illustrative — check the live catalog) |
|---|---|---|
llm | Instruction-tuned language models (vLLM / transformers) | qwen-2.5-7b-instruct-public, mistral-7b-instruct-public, phi-3.5-mini-public, deepseek-distilled-7b-public, hf-google-gemma-2-9b-it |
specialist | Single-purpose ML endpoints | embeddings-public, rerank-public, stt-public, moderate-public, ocr-public |
tool | Network-enabled utilities (sideEffects: ["network"]) | web-fetch-public, web-search-public, web-scrape-public, geo-geocode-public |
generative | Image / audio / video / speech generation | z-image-turbo, hf-black-forest-labs-FLUX.1-dev, wan22-video, kokoro-tts, hf-facebook-musicgen-small |
model | Trainable/forecasting models (tabular, time series) | patchtst-forecasting, prophet-forecasting, xgboost-regression, random-forest-regression |
trained_model | Your trained/merged models (per-account) | appears only in your manifest, after you train/register |
Each tool's full manifest (input/output schemas, examples, side-effects, cost/latency hints) is
fetched at runtime via GET /api/builder/v1/mcp/manifest/{slug}. To invoke one:
run_inference (built-in), or the agent calls the tool by its slug directly.
Thinking models
run_inference against a thinking model returns the reasoning trace in its own field, separate
from the answer:
{"choices": [{"message": {
"content": "7 + 5 = 12",
"reasoning": "The user asks a simple sum in Spanish...",
"tool_calls": []
}}]}
Read reasoning. reasoning_content carries the same value as a legacy alias, but a client that
reads only the old name can get an empty string in silence from engines that renamed the field.
This matters for an agent more than for a human reader: the tool-call parser only looks for calls
in content, never in the reasoning. If the trace were not separated, a thinking model that
decides to call a tool would emit that call inside its reasoning, tool_calls would come back
empty, and the agent would see prose instead of an action. ColabHive enforces the invariant that
prevents it — an endpoint serving tool calling always serves a reasoning parser — so you do not
configure anything here. If tool_calls is empty and content reads like a train of thought, see
Reasoning.
Give them budget: the trace spends the same max_tokens as the answer, so a tight limit can produce
an empty content with finish_reason: "length" and no tool call.
Common input shapes
- LLMs:
{"messages": [{"role": "user", "content": "..."}], "temperature": 0.7, "max_tokens": 512} - Embeddings:
{"text": "..."}→{"embedding": [...]} - Rerank:
{"query": "...", "documents": [...]}→{"scores": [...]} - CLM scorer (
clm-v0.1-8b,kind=model; the endpoint of a head you trained,kind=trained_model): exactly one of{"system_one": {"state": ..., "questions": {...}}}or{"rank": {"context": ..., "question": "...", "answers": [...]}}— see CLM scorer results. - Tools (web-fetch etc.):
{"url": "..."}— see the tool's manifest for the exact schema.
Context windows, dimensions, prices, and readiness are per-endpoint and change over time — read them from the tool's manifest or the Model Catalog, not from this page.
CLM scorer results
A CLM scorer generates no text, so its answer is rendered for
the agent instead of passed as raw JSON: ranked candidates as 1. <candidate> (p=0.993) lines, best
first; System One answers as one line per question — P(true) for noul, the chosen option for
choice, the expected level for score, each with its confidence — plus the head's sha256 when the
model reports it. A second text block carries the whole answer as JSON for exact values. The
p=0.993 above shows the format only.
The input schema of a CLM tool admits exactly one of system_one and rank. The scorer does not
chat: a messages input is refused with model_not_chat.
Trained models (kind=trained_model)
These are per-account: each appears in your manifest only after you train it via the
Training API and it is registered as an endpoint. A trained CLM head is one of
them: its tool takes the CLM scorer input above and runs on the clm-v0.1-8b replicas. Its manifest carries a
lineage block pointing back to its base and dataset. Once published (visibility=public) and
reviewed, a trained model can also appear in other accounts' manifests.
Filtering the tools your agent sees
Use config flags — useful for cost control and reducing context bloat:
# Only LLMs and your own trained models
export COLABHIVE_ALLOW_KINDS="llm,trained_model"
# Only management + specialists (exclude LLMs, generative, tools)
export COLABHIVE_ALLOW_KINDS="specialist,trained_model"
# Only stable tools (exclude experimental/beta)
export COLABHIVE_STABILITY=stable
# Allowlist by slug glob
export COLABHIVE_ALLOW_TOOLS="qwen-*,embeddings-public,web-fetch-public"
Filters apply to the manifest tools; the built-in control-plane tools are always available.
See also
- Manifests — full schema for tool metadata
- Model Catalog — per-model documentation
- Actions API list endpoint — programmatic discovery