OpenAI-Compatible API
ColabHive exposes an OpenAI-compatible chat endpoint so you can point existing OpenAI clients at the platform with only a base-URL and API-key change.
POST https://api.colabhive.com/v1/chat/completions
GET https://api.colabhive.com/v1/models
GET https://api.colabhive.com/v1/models/{model}
Note the prefix: this surface is mounted at /v1 (the API root), not under /api/builder/v1.
The same root also serves the two routes of the CLM scorer, POST /v1/systemone and POST /v1/rank.
They are not OpenAI routes; see CLM scorer routes.
Authenticate with your ColabHive API key (hive_…) — see Authentication.
Errors are returned as OpenAI-style error objects ({"error": {"message": ..., "type": ...}}), so the
official SDKs raise their normal exception types (AuthenticationError on 401, NotFoundError on
404).
Request
{
"model": "qwen-2.5-7b-instruct",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "What is machine learning?"}
],
"max_tokens": 500,
"temperature": 0.7
}
| Field | Type | Notes |
|---|---|---|
model | string | Required. A ColabHive endpoint UUID — the id of an entry in GET /v1/models. That is the canonical identifier and always routes to the right deployment. A name is also accepted, tried in this order: a catalog model name (resolved to an active base model), then the endpoint name, then the readable name that /v1/models returns. When two endpoints you can see carry the same label, the one in your account wins; if the label is still ambiguous the answer is 409 with code: "ambiguous_model" and the candidate ids. Unknown values return 404. |
messages | array | Required, non-empty. OpenAI chat messages (role + content); tool role is accepted. |
max_tokens | int | Optional. |
temperature | number | Optional. |
stream | bool | Optional. See Streaming — note the current behavior. |
Passthrough fields. Additional OpenAI fields are forwarded verbatim to the model server:
tools, tool_choice, parallel_tool_calls, response_format, and other unknown fields. Tool calling
requires the target model to have been started with tool-choice support; when available, the response
includes tool_calls and finish_reason: "tool_calls".
Request size
A request body may be up to 11.9 MiB (12517376 bytes) of compact UTF-8 JSON — well over a
million tokens of ASCII context, so an agent carrying a 200k-token conversation fits many times
over. A larger body gets 400 with the limit in the message:
{"error": {"message": "Input validation failed: Data exceeds maximum encoded JSON size (max: 12517376 bytes)", "type": "invalid_request_error"}}
The number is not arbitrary: the whole request travels inside one WebSocket frame to the GPU node
that serves it, and the bound is that frame (12 MiB) minus the notification's envelope. Admitting
more would not deliver more — it would turn an error you can act on into a failed dispatch. Compact
the conversation against the model's max_model_len before you reach it.
Non-ASCII text is measured the same way: the budget counts UTF-8 bytes, so a character outside ASCII spends 2–4 of them, never one.
Response
A standard OpenAI chat.completion object:
{
"id": "chatcmpl-8f2a...",
"object": "chat.completion",
"created": 1770000000,
"model": "qwen-2.5-7b-instruct",
"choices": [
{
"index": 0,
"message": {"role": "assistant", "content": "Machine learning is..."},
"finish_reason": "stop"
}
],
"usage": {"prompt_tokens": 0, "completion_tokens": 0, "total_tokens": 0}
}
When the underlying model server emits tool_calls, finish_reason, and token usage, those pass
through unchanged. If the served model does not report token counts, usage fields are 0.
Without stream, the request is synchronous: nothing is sent until the whole completion is ready,
and the API edge closes a response that stays silent for 300 s. The gateway answers before that
limit. If the model is still not ready by then, the request is cancelled and is not billed, and the
answer is HTTP 503 with a Retry-After header. Only a request the model completes at the very
moment the cancellation reaches it counts as completed: the gateway returns that result when it still
can.
{"error": {"message": "The model was not ready within the 300 s ...", "type": "api_error", "param": null, "code": "no_warm_replica"}}
codeis why the request was still waiting, in the vocabulary ofreason_codein Inference:no_warm_replica(the model has no ready replica yet; one is being started),resident_capacity_full(its replicas are busy), or another reason listed there.Retry-Aftersays when to retry, in seconds. While the model loads it is the rest of its measured load time, counted from when the load started, and at least1. When the request waited for a busy replica it is a few seconds. It is0when the replica came up just as the limit was reached. The model keeps loading after the cancellation, so a retry after that delay finds it ready.- The OpenAI SDKs retry a
503twice by default, but they waitRetry-Afteronly when it is between 1 and 60 seconds. With a longer (or0) value they use their own backoff, 0.5 s doubling up to 8 s, so they can use up their retries before a model that takes minutes to load is ready. WaitRetry-Afteryourself, or usestream: true.
If the model was ready but a non-streamed answer takes longer than the limit, the request is cancelled
the same way and the answer is HTTP 504 with x-should-retry: false: send it with stream: true.
Streamed requests are not cut at 300 s (see Streaming): the gateway keeps them alive
while the model loads, so "stream": true is the way to wait for a model whose cold load takes
minutes. They still have a ceiling, load and generation included: 20 minutes when model is an
endpoint id or its name, and 1.25× the model's worst measured load time (at least 2 minutes) when it
is a catalog model name. A client with its own shorter timeout — OpenCode's default is 5 minutes —
gives up first.
Requests that set colabhive_execution (the Cohort preview) are not covered by this yet: they keep
their own wait, up to the model's measured load time, and can still get the edge's own 504 from a
model whose load takes longer than 300 s.
Streaming
Passing "stream": true returns incremental, token-by-token SSE (Content-Type: text/event-stream): the deltas are forwarded from the model server as they are produced, ending with
data: [DONE]. Use it exactly as you would with OpenAI.
stream = client.chat.completions.create(
model="5d21e32a-3bbb-4040-9c34-3b06c4415b84",
messages=[{"role": "user", "content": "Write a short paragraph about distributed computing."}],
max_tokens=800,
stream=True,
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="", flush=True)
If the node serving your request runs a node-runtime older than 0.10.195, it cannot emit incremental
deltas. In that case the gateway falls back to the previous behaviour — the full completion arrives as
one chat.completion.chunk followed by data: [DONE]. The response is still valid SSE and the
official SDKs consume it without changes, so a mixed-version fleet never breaks a client; you simply
get the answer all at once instead of gradually.
Listing models
GET /v1/models returns the standard OpenAI model list. The id of each entry is a ColabHive
endpoint UUID — the exact value to pass as model in a chat request, so whatever a client lists is
always something it can actually call.
{
"object": "list",
"data": [
{
"id": "5d21e32a-3bbb-4040-9c34-3b06c4415b84",
"object": "model",
"created": 1767654203,
"owned_by": "colabhive",
"name": "gpt-oss-20b",
"task_type": "text-generation",
"max_model_len": 131072,
"max_model_len_source": "resident"
}
]
}
name, task_type, model_config_id, tool_calling_configured and supports_tool_calls are
ColabHive extensions — the official SDKs ignore unknown fields, while UIs
that understand them can show a readable label instead of a raw UUID. name is also accepted as
model, but the id is the identifier that never changes: prefer it in configuration files.
Context window
max_model_len is the context window, in tokens, you should compact your conversation against. It is
the window the replica actually runs, not the one the catalogue declares: a model that declares
262144 may be served on a board where only 32768 fits in the KV cache, and compacting against the
declaration would fail the session halfway through.
max_model_len_source says where the number came from:
| Value | Meaning |
|---|---|
"resident" | Measured: a replica is serving this endpoint right now and runs that window. When several replicas serve it, the smallest of their windows is published — you do not choose the replica. |
"configured" | No replica is serving yet; this is the window the catalogue declares, which is what a cold start will ask for. |
null | Unknown. max_model_len is null too. Nothing is ever invented here — treat it as "no published window" rather than as a number. |
What you see. Models owned by your account, plus the approved public catalogue. The list is scoped
to models this surface can actually serve (chat-capable LLMs); trained tabular, forecasting and
specialist endpoints are not listed here — reach those through the
Inference API. CLM scorer endpoints are never listed and never accepted by
/v1/chat/completions: call them through the CLM scorer routes.
Tool calling fields
tool_calling_configuredistruewhen exactly one tool-call parser is declared for the model and auto tool choice is enabled; otherwise it isnull. It says the server can parse tool calls, not that the model uses them well.supports_tool_callsis alwaysnull: the platform does not claim a capability from configuration. Check it with a real request — see Coding Agents: OpenCode setup.model_config_ididentifies the model configuration behind the endpoint. Two endpoints with the samemodel_config_idserve the same model.
Retrieve a model
GET /v1/models/{model} retrieves a single entry. It accepts the endpoint UUID, or a label: the
endpoint name or the readable name, with the same precedence as chat completions (your account's
endpoints first) and the same 409 ambiguous_model when a label matches several. A catalog model name
that is not also an endpoint label is not accepted here. Unknown or invisible values return 404
with an OpenAI error object.
Reasoning models
Some models (for example gpt-oss-20b, Qwen3, DeepSeek-R1, GLM-4.7) reason before answering and
return that trace in its own field. The user-facing text is always in the usual
choices[0].message.content.
| field | what it holds |
|---|---|
choices[0].message.content | the answer |
choices[0].message.reasoning | the reasoning trace. Canonical name — read this one. |
choices[0].message.reasoning_content | the same trace, kept populated as a legacy alias |
reasoning, not reasoning_contentThe engine renamed this field. A client that reads only the old name can get an empty string in
silence while the new one is populated, and conclude the model does not reason when it does.
ColabHive keeps both names populated so no existing client breaks, but new code should read
reasoning and fall back:
message = response["choices"][0]["message"]
reasoning = message.get("reasoning") or message.get("reasoning_content") or ""
In a stream
The trace arrives in its own delta key, interleaved with the answer's. Append them to separate buffers:
data: {"choices":[{"delta":{"reasoning":"Let me check the table..."}}]}
data: {"choices":[{"delta":{"content":"7 + 5 = 12"}}]}
A client that appends every delta into one string ends up with the trace glued to the answer. The non-streaming message ColabHive rebuilds from a stream carries both field names with the same value, so the streamed and non-streamed shapes agree.
include_reasoning
Some builds omit the trace unless it is asked for. Send include_reasoning: true in the request body
to be explicit; a build that always returns the trace ignores the flag, so sending it is safe either
way.
Reasoning and tool_calls
The tool-call parser only looks for calls in content. It never looks in the reasoning. If a
thinking model emits its call inside the trace and the trace is not separated, tool_calls comes
back empty and it looks like the model wrote prose instead of using its tools.
ColabHive enforces the invariant that prevents it — an endpoint serving tool calling always serves a
reasoning parser — so on the hosted platform you should never see this. If you self-host and
tool_calls is empty while content reads like a train of thought, that is the configuration, not
the model. See Reasoning for how to confirm it.
max_tokensThe reasoning trace consumes the same token budget as the answer. With a small max_tokens, the model
can spend it all reasoning and return an empty content with finish_reason: "length". If you get
blank answers from a reasoning model, raise max_tokens before anything else.
Examples
OpenAI Python SDK
from openai import OpenAI
client = OpenAI(
base_url="https://api.colabhive.com/v1",
api_key="hive_...",
)
# Discover what you can call — each `id` is usable as `model` verbatim.
for m in client.models.list():
print(m.id, m.name)
resp = client.chat.completions.create(
model="5d21e32a-3bbb-4040-9c34-3b06c4415b84",
messages=[{"role": "user", "content": "What is machine learning?"}],
max_tokens=500,
)
print(resp.choices[0].message.content)
cURL
curl "https://api.colabhive.com/v1/models" \
-H "Authorization: Bearer hive_..."
curl -X POST "https://api.colabhive.com/v1/chat/completions" \
-H "Authorization: Bearer hive_..." \
-H "Content-Type: application/json" \
-d '{
"model": "5d21e32a-3bbb-4040-9c34-3b06c4415b84",
"messages": [{"role": "user", "content": "Hello!"}],
"max_tokens": 500
}'
CLM scorer routes: /v1/systemone and /v1/rank
POST https://api.colabhive.com/v1/systemone
POST https://api.colabhive.com/v1/rank
These routes call a CLM scorer: a model that generates no
text and scores the candidates you send. They keep the request and response format of the upstream
CLM server (Contrastive-LM/CLM) and add one required field,
model. Authenticate as for chat completions (Authorization: Bearer hive_…); the account comes from
the key.
model
A ColabHive CLM endpoint your account may call: its UUID, or its endpoint_name or
display_name (the public one is clm-v0.1-8b; a head you trained is served by the endpoint you
registered for it). Labels resolve with the same precedence as /v1/models: your account's endpoints
first, endpoint_name before display_name; a label that still matches several CLM endpoints is
409 ambiguous_model with their ids. model is never:
- a Hugging Face repo id (
400 hf_repo_not_accepted); - a runtime key containing
:trained:(400 runtime_key_not_accepted) — such a key is not bound to an account; - an endpoint that is not a CLM scorer (
400 model_not_clm).
POST /v1/systemone
| Field | Type | Notes |
|---|---|---|
model | string | Required. See above. |
state | string, object or array | Required. The context. A list of chat messages only for heads trained on chat-formatted states. |
questions | object | Required, non-empty. {id: question}; each question is {"type": "noul" | "choice" | "score", "instructions", "criteria"}. At most 64 per request. |
temperature | number | Optional, greater than 0 and at most 100; default 1. |
{
"model": "clm-v0.1-8b",
"state": "Customer: my invoice was charged twice and nobody answers the phone!",
"questions": {
"urgency": {"type": "noul", "instructions": "Is this urgent?"},
"department": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "Charges, invoices, refunds", "technical": "Bugs and outages"}}
}
}
Response: {"model", "answers": {id: answer}, "usage", "colabhive"}. A noul answer is
{"type": "noul", "noul": P(true)}; a choice answer is {"type": "choice", "choice", "confidence", "probabilities"}; a score answer is {"type": "score", "score", "confidence", "probabilities", "legend"}. usage is {"billing_units": <number of questions>, "input_tokens": <encoder tokens>, "output_tokens": 0}. The model page has a full
example with its measured output.
POST /v1/rank
| Field | Type | Notes |
|---|---|---|
model | string | Required. See above. |
context | string, object or array | The context; may be omitted. |
question | string | Appended after the context; may be omitted. |
answers | array of non-empty strings | Required, non-empty. The candidates. |
temperature | number | Optional, greater than 0 and at most 100; default 1. |
Response: {"model", "ranked": [{"rank", "candidate", "prob"}], "usage", "colabhive"}, best first;
the prob values sum to 1.
What ColabHive adds
colabhive:{"endpoint_id", "task_id", "head_sha256", "encoder_identity", "runtime_vendor"}— which endpoint and task served the request, which head scored it and on which encoder and GPU vendor. A client built for the upstream server ignores it.X-CLM-Latency-Ms: end-to-end time inside the API (queue, dispatch and encoding), in milliseconds.
To point a client written for the upstream server at ColabHive, set its base URL to
https://api.colabhive.com, send the Authorization: Bearer hive_… header, and add model to each
request body. Nothing else is promised to work unchanged; the Python SDK's
client.clm does all three and keeps the upstream question types
(Noul, Choice, Score).
Limits and errors
| Status | code | When |
|---|---|---|
| 413 | clm_too_many_candidates | More than 256 candidates (a noul question counts 2, a choice or score question one per option, /v1/rank one per answer). |
| 413 | too_many_questions | More than 64 questions. |
| 413 | too_many_tokens | More than 32768 encoder tokens after each text is truncated to the head's window. |
| 413 | clm_input_too_large | More input than one request may make the encoder read within the edge limit, or a JSON body over 12517376 bytes. |
| 400 | hf_repo_not_accepted, runtime_key_not_accepted, model_not_clm, or the model's own code | Invalid model or a request the model refuses (for example a chat-message state on a head trained on prose). |
| 404 | — | No CLM endpoint with that id or label that your account may call. |
| 409 | ambiguous_model | The label matches several CLM endpoints. |
| 502 | — | The model failed or returned an unexpected result. |
| 503 | wait reason, endpoint_lookup_unavailable, or a capacity code | Not ready or no capacity in time: wait Retry-After. |
| 504 | — | The request was running but could not finish within the edge limit. |
413 and the 400 refusals carry x-should-retry: false. These routes are not streamed: the API edge
closes a response that is silent for 300 s, and the gateway answers before that, exactly as for a
non-streamed chat completion — a model still loading gives 503 with Retry-After, and it keeps
loading after the cancellation. The Python SDK waits up to 300 s on these calls by default.
See Also
- Inference API — the native inference surface (sync/async, readiness, binary I/O)
- Actions API — discover callable models and their slugs
- Models API — import HF models, registry, capabilities
- CLM-8B — the scorer behind
/v1/systemoneand/v1/rank