Skip to main content

OpenAI-Compatible API

ColabHive exposes an OpenAI-compatible chat endpoint so you can point existing OpenAI clients at the platform with only a base-URL and API-key change.

POST https://api.colabhive.com/v1/chat/completions
GET https://api.colabhive.com/v1/models
GET https://api.colabhive.com/v1/models/{model}

Note the prefix: this surface is mounted at /v1 (the API root), not under /api/builder/v1.

The same root also serves the two routes of the CLM scorer, POST /v1/systemone and POST /v1/rank. They are not OpenAI routes; see CLM scorer routes.

Authenticate with your ColabHive API key (hive_…) — see Authentication. Errors are returned as OpenAI-style error objects ({"error": {"message": ..., "type": ...}}), so the official SDKs raise their normal exception types (AuthenticationError on 401, NotFoundError on 404).


Request​

{
"model": "qwen-2.5-7b-instruct",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "What is machine learning?"}
],
"max_tokens": 500,
"temperature": 0.7
}
FieldTypeNotes
modelstringRequired. A ColabHive endpoint UUID — the id of an entry in GET /v1/models. That is the canonical identifier and always routes to the right deployment. A name is also accepted, tried in this order: a catalog model name (resolved to an active base model), then the endpoint name, then the readable name that /v1/models returns. When two endpoints you can see carry the same label, the one in your account wins; if the label is still ambiguous the answer is 409 with code: "ambiguous_model" and the candidate ids. Unknown values return 404.
messagesarrayRequired, non-empty. OpenAI chat messages (role + content); tool role is accepted.
max_tokensintOptional.
temperaturenumberOptional.
streamboolOptional. See Streaming — note the current behavior.

Passthrough fields. Additional OpenAI fields are forwarded verbatim to the model server: tools, tool_choice, parallel_tool_calls, response_format, and other unknown fields. Tool calling requires the target model to have been started with tool-choice support; when available, the response includes tool_calls and finish_reason: "tool_calls".

Request size​

A request body may be up to 11.9 MiB (12517376 bytes) of compact UTF-8 JSON — well over a million tokens of ASCII context, so an agent carrying a 200k-token conversation fits many times over. A larger body gets 400 with the limit in the message:

{"error": {"message": "Input validation failed: Data exceeds maximum encoded JSON size (max: 12517376 bytes)", "type": "invalid_request_error"}}

The number is not arbitrary: the whole request travels inside one WebSocket frame to the GPU node that serves it, and the bound is that frame (12 MiB) minus the notification's envelope. Admitting more would not deliver more — it would turn an error you can act on into a failed dispatch. Compact the conversation against the model's max_model_len before you reach it.

Non-ASCII text is measured the same way: the budget counts UTF-8 bytes, so a character outside ASCII spends 2–4 of them, never one.


Response​

A standard OpenAI chat.completion object:

{
"id": "chatcmpl-8f2a...",
"object": "chat.completion",
"created": 1770000000,
"model": "qwen-2.5-7b-instruct",
"choices": [
{
"index": 0,
"message": {"role": "assistant", "content": "Machine learning is..."},
"finish_reason": "stop"
}
],
"usage": {"prompt_tokens": 0, "completion_tokens": 0, "total_tokens": 0}
}

When the underlying model server emits tool_calls, finish_reason, and token usage, those pass through unchanged. If the served model does not report token counts, usage fields are 0.

Without stream, the request is synchronous: nothing is sent until the whole completion is ready, and the API edge closes a response that stays silent for 300 s. The gateway answers before that limit. If the model is still not ready by then, the request is cancelled and is not billed, and the answer is HTTP 503 with a Retry-After header. Only a request the model completes at the very moment the cancellation reaches it counts as completed: the gateway returns that result when it still can.

{"error": {"message": "The model was not ready within the 300 s ...", "type": "api_error", "param": null, "code": "no_warm_replica"}}
  • code is why the request was still waiting, in the vocabulary of reason_code in Inference: no_warm_replica (the model has no ready replica yet; one is being started), resident_capacity_full (its replicas are busy), or another reason listed there.
  • Retry-After says when to retry, in seconds. While the model loads it is the rest of its measured load time, counted from when the load started, and at least 1. When the request waited for a busy replica it is a few seconds. It is 0 when the replica came up just as the limit was reached. The model keeps loading after the cancellation, so a retry after that delay finds it ready.
  • The OpenAI SDKs retry a 503 twice by default, but they wait Retry-After only when it is between 1 and 60 seconds. With a longer (or 0) value they use their own backoff, 0.5 s doubling up to 8 s, so they can use up their retries before a model that takes minutes to load is ready. Wait Retry-After yourself, or use stream: true.

If the model was ready but a non-streamed answer takes longer than the limit, the request is cancelled the same way and the answer is HTTP 504 with x-should-retry: false: send it with stream: true. Streamed requests are not cut at 300 s (see Streaming): the gateway keeps them alive while the model loads, so "stream": true is the way to wait for a model whose cold load takes minutes. They still have a ceiling, load and generation included: 20 minutes when model is an endpoint id or its name, and 1.25× the model's worst measured load time (at least 2 minutes) when it is a catalog model name. A client with its own shorter timeout — OpenCode's default is 5 minutes — gives up first.

Requests that set colabhive_execution (the Cohort preview) are not covered by this yet: they keep their own wait, up to the model's measured load time, and can still get the edge's own 504 from a model whose load takes longer than 300 s.


Streaming​

Passing "stream": true returns incremental, token-by-token SSE (Content-Type: text/event-stream): the deltas are forwarded from the model server as they are produced, ending with data: [DONE]. Use it exactly as you would with OpenAI.

stream = client.chat.completions.create(
model="5d21e32a-3bbb-4040-9c34-3b06c4415b84",
messages=[{"role": "user", "content": "Write a short paragraph about distributed computing."}],
max_tokens=800,
stream=True,
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="", flush=True)
Automatic fallback on older nodes

If the node serving your request runs a node-runtime older than 0.10.195, it cannot emit incremental deltas. In that case the gateway falls back to the previous behaviour — the full completion arrives as one chat.completion.chunk followed by data: [DONE]. The response is still valid SSE and the official SDKs consume it without changes, so a mixed-version fleet never breaks a client; you simply get the answer all at once instead of gradually.


Listing models​

GET /v1/models returns the standard OpenAI model list. The id of each entry is a ColabHive endpoint UUID — the exact value to pass as model in a chat request, so whatever a client lists is always something it can actually call.

{
"object": "list",
"data": [
{
"id": "5d21e32a-3bbb-4040-9c34-3b06c4415b84",
"object": "model",
"created": 1767654203,
"owned_by": "colabhive",
"name": "gpt-oss-20b",
"task_type": "text-generation",
"max_model_len": 131072,
"max_model_len_source": "resident"
}
]
}

name, task_type, model_config_id, tool_calling_configured and supports_tool_calls are ColabHive extensions — the official SDKs ignore unknown fields, while UIs that understand them can show a readable label instead of a raw UUID. name is also accepted as model, but the id is the identifier that never changes: prefer it in configuration files.

Context window​

max_model_len is the context window, in tokens, you should compact your conversation against. It is the window the replica actually runs, not the one the catalogue declares: a model that declares 262144 may be served on a board where only 32768 fits in the KV cache, and compacting against the declaration would fail the session halfway through.

max_model_len_source says where the number came from:

ValueMeaning
"resident"Measured: a replica is serving this endpoint right now and runs that window. When several replicas serve it, the smallest of their windows is published — you do not choose the replica.
"configured"No replica is serving yet; this is the window the catalogue declares, which is what a cold start will ask for.
nullUnknown. max_model_len is null too. Nothing is ever invented here — treat it as "no published window" rather than as a number.

What you see. Models owned by your account, plus the approved public catalogue. The list is scoped to models this surface can actually serve (chat-capable LLMs); trained tabular, forecasting and specialist endpoints are not listed here — reach those through the Inference API. CLM scorer endpoints are never listed and never accepted by /v1/chat/completions: call them through the CLM scorer routes.

Tool calling fields​

  • tool_calling_configured is true when exactly one tool-call parser is declared for the model and auto tool choice is enabled; otherwise it is null. It says the server can parse tool calls, not that the model uses them well.
  • supports_tool_calls is always null: the platform does not claim a capability from configuration. Check it with a real request — see Coding Agents: OpenCode setup.
  • model_config_id identifies the model configuration behind the endpoint. Two endpoints with the same model_config_id serve the same model.

Retrieve a model​

GET /v1/models/{model} retrieves a single entry. It accepts the endpoint UUID, or a label: the endpoint name or the readable name, with the same precedence as chat completions (your account's endpoints first) and the same 409 ambiguous_model when a label matches several. A catalog model name that is not also an endpoint label is not accepted here. Unknown or invisible values return 404 with an OpenAI error object.


Reasoning models​

Some models (for example gpt-oss-20b, Qwen3, DeepSeek-R1, GLM-4.7) reason before answering and return that trace in its own field. The user-facing text is always in the usual choices[0].message.content.

fieldwhat it holds
choices[0].message.contentthe answer
choices[0].message.reasoningthe reasoning trace. Canonical name — read this one.
choices[0].message.reasoning_contentthe same trace, kept populated as a legacy alias
Read reasoning, not reasoning_content

The engine renamed this field. A client that reads only the old name can get an empty string in silence while the new one is populated, and conclude the model does not reason when it does. ColabHive keeps both names populated so no existing client breaks, but new code should read reasoning and fall back:

message = response["choices"][0]["message"]
reasoning = message.get("reasoning") or message.get("reasoning_content") or ""

In a stream​

The trace arrives in its own delta key, interleaved with the answer's. Append them to separate buffers:

data: {"choices":[{"delta":{"reasoning":"Let me check the table..."}}]}
data: {"choices":[{"delta":{"content":"7 + 5 = 12"}}]}

A client that appends every delta into one string ends up with the trace glued to the answer. The non-streaming message ColabHive rebuilds from a stream carries both field names with the same value, so the streamed and non-streamed shapes agree.

include_reasoning​

Some builds omit the trace unless it is asked for. Send include_reasoning: true in the request body to be explicit; a build that always returns the trace ignores the flag, so sending it is safe either way.

Reasoning and tool_calls​

The tool-call parser only looks for calls in content. It never looks in the reasoning. If a thinking model emits its call inside the trace and the trace is not separated, tool_calls comes back empty and it looks like the model wrote prose instead of using its tools.

ColabHive enforces the invariant that prevents it — an endpoint serving tool calling always serves a reasoning parser — so on the hosted platform you should never see this. If you self-host and tool_calls is empty while content reads like a train of thought, that is the configuration, not the model. See Reasoning for how to confirm it.

Give reasoning models enough max_tokens

The reasoning trace consumes the same token budget as the answer. With a small max_tokens, the model can spend it all reasoning and return an empty content with finish_reason: "length". If you get blank answers from a reasoning model, raise max_tokens before anything else.


Examples​

OpenAI Python SDK​

from openai import OpenAI

client = OpenAI(
base_url="https://api.colabhive.com/v1",
api_key="hive_...",
)

# Discover what you can call — each `id` is usable as `model` verbatim.
for m in client.models.list():
print(m.id, m.name)

resp = client.chat.completions.create(
model="5d21e32a-3bbb-4040-9c34-3b06c4415b84",
messages=[{"role": "user", "content": "What is machine learning?"}],
max_tokens=500,
)
print(resp.choices[0].message.content)

cURL​

curl "https://api.colabhive.com/v1/models" \
-H "Authorization: Bearer hive_..."

curl -X POST "https://api.colabhive.com/v1/chat/completions" \
-H "Authorization: Bearer hive_..." \
-H "Content-Type: application/json" \
-d '{
"model": "5d21e32a-3bbb-4040-9c34-3b06c4415b84",
"messages": [{"role": "user", "content": "Hello!"}],
"max_tokens": 500
}'

CLM scorer routes: /v1/systemone and /v1/rank​

POST https://api.colabhive.com/v1/systemone
POST https://api.colabhive.com/v1/rank

These routes call a CLM scorer: a model that generates no text and scores the candidates you send. They keep the request and response format of the upstream CLM server (Contrastive-LM/CLM) and add one required field, model. Authenticate as for chat completions (Authorization: Bearer hive_…); the account comes from the key.

model​

A ColabHive CLM endpoint your account may call: its UUID, or its endpoint_name or display_name (the public one is clm-v0.1-8b; a head you trained is served by the endpoint you registered for it). Labels resolve with the same precedence as /v1/models: your account's endpoints first, endpoint_name before display_name; a label that still matches several CLM endpoints is 409 ambiguous_model with their ids. model is never:

  • a Hugging Face repo id (400 hf_repo_not_accepted);
  • a runtime key containing :trained: (400 runtime_key_not_accepted) — such a key is not bound to an account;
  • an endpoint that is not a CLM scorer (400 model_not_clm).

POST /v1/systemone​

FieldTypeNotes
modelstringRequired. See above.
statestring, object or arrayRequired. The context. A list of chat messages only for heads trained on chat-formatted states.
questionsobjectRequired, non-empty. {id: question}; each question is {"type": "noul" | "choice" | "score", "instructions", "criteria"}. At most 64 per request.
temperaturenumberOptional, greater than 0 and at most 100; default 1.
{
"model": "clm-v0.1-8b",
"state": "Customer: my invoice was charged twice and nobody answers the phone!",
"questions": {
"urgency": {"type": "noul", "instructions": "Is this urgent?"},
"department": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "Charges, invoices, refunds", "technical": "Bugs and outages"}}
}
}

Response: {"model", "answers": {id: answer}, "usage", "colabhive"}. A noul answer is {"type": "noul", "noul": P(true)}; a choice answer is {"type": "choice", "choice", "confidence", "probabilities"}; a score answer is {"type": "score", "score", "confidence", "probabilities", "legend"}. usage is {"billing_units": <number of questions>, "input_tokens": <encoder tokens>, "output_tokens": 0}. The model page has a full example with its measured output.

POST /v1/rank​

FieldTypeNotes
modelstringRequired. See above.
contextstring, object or arrayThe context; may be omitted.
questionstringAppended after the context; may be omitted.
answersarray of non-empty stringsRequired, non-empty. The candidates.
temperaturenumberOptional, greater than 0 and at most 100; default 1.

Response: {"model", "ranked": [{"rank", "candidate", "prob"}], "usage", "colabhive"}, best first; the prob values sum to 1.

What ColabHive adds​

  • colabhive: {"endpoint_id", "task_id", "head_sha256", "encoder_identity", "runtime_vendor"} — which endpoint and task served the request, which head scored it and on which encoder and GPU vendor. A client built for the upstream server ignores it.
  • X-CLM-Latency-Ms: end-to-end time inside the API (queue, dispatch and encoding), in milliseconds.

To point a client written for the upstream server at ColabHive, set its base URL to https://api.colabhive.com, send the Authorization: Bearer hive_… header, and add model to each request body. Nothing else is promised to work unchanged; the Python SDK's client.clm does all three and keeps the upstream question types (Noul, Choice, Score).

Limits and errors​

StatuscodeWhen
413clm_too_many_candidatesMore than 256 candidates (a noul question counts 2, a choice or score question one per option, /v1/rank one per answer).
413too_many_questionsMore than 64 questions.
413too_many_tokensMore than 32768 encoder tokens after each text is truncated to the head's window.
413clm_input_too_largeMore input than one request may make the encoder read within the edge limit, or a JSON body over 12517376 bytes.
400hf_repo_not_accepted, runtime_key_not_accepted, model_not_clm, or the model's own codeInvalid model or a request the model refuses (for example a chat-message state on a head trained on prose).
404—No CLM endpoint with that id or label that your account may call.
409ambiguous_modelThe label matches several CLM endpoints.
502—The model failed or returned an unexpected result.
503wait reason, endpoint_lookup_unavailable, or a capacity codeNot ready or no capacity in time: wait Retry-After.
504—The request was running but could not finish within the edge limit.

413 and the 400 refusals carry x-should-retry: false. These routes are not streamed: the API edge closes a response that is silent for 300 s, and the gateway answers before that, exactly as for a non-streamed chat completion — a model still loading gives 503 with Retry-After, and it keeps loading after the cancellation. The Python SDK waits up to 300 s on these calls by default.


See Also​

  • Inference API — the native inference surface (sync/async, readiness, binary I/O)
  • Actions API — discover callable models and their slugs
  • Models API — import HF models, registry, capabilities
  • CLM-8B — the scorer behind /v1/systemone and /v1/rank