Reasoning
A thinking model works before it answers. It emits a reasoning trace and then the answer, and the platform returns them as two separate fields, not one blob of text.
That separation is not cosmetic. If the trace is not separated, it lands inside the answer — and the machinery that reads the answer, including tool calling, stops working. This page says exactly what the contract is, in every layer, so you can tell a model problem from a configuration one.
The two fields
| field | what it holds |
|---|---|
choices[0].message.content | the answer. Always the user-facing text. |
choices[0].message.reasoning | the reasoning trace. Canonical name. |
choices[0].message.reasoning_content | the same trace, kept populated as a legacy alias. |
Read reasoning. The engine renamed this field, and its own documentation warns that a client
reading the old name can get an empty string in silence while the new one is populated — so a
client that only knows reasoning_content can conclude "this model does not reason" when it does.
ColabHive keeps both names populated so neither client breaks, but new code should read reasoning.
message = response["choices"][0]["message"]
answer = message["content"]
reasoning = message.get("reasoning") or message.get("reasoning_content") or ""
A model that does not reason simply has no reasoning field. That is an answer, not an error.
Streaming
In a stream the trace arrives in its own delta key, interleaved with the answer's:
data: {"choices":[{"delta":{"reasoning":"Let me check the table..."}}]}
data: {"choices":[{"delta":{"content":"7 + 5 = 12"}}]}
data: [DONE]
Read both keys on the delta and append them to separate buffers — a client that appends every delta into one string gets the trace glued to the answer, which is the same failure as not separating it at all.
answer, reasoning = [], []
for chunk in stream:
delta = chunk["choices"][0].get("delta") or {}
if delta.get("content"):
answer.append(delta["content"])
trace = delta.get("reasoning") or delta.get("reasoning_content")
if trace:
reasoning.append(trace)
When ColabHive stores or retries a streamed result it rebuilds a complete, non-streaming
ChatCompletion from those deltas, and that rebuilt message carries both field names with the
same value. So the shape you get from a stream and the shape you get from a plain request agree.
Reasoning and tool calling: the failure that looks like a bad model
The tool-call parser only looks for calls in content. It never looks in the reasoning.
So if the trace is not separated, a thinking model that decides to call write emits that call
inside its reasoning, the parser does not see it, and tool_calls comes back empty. From the
outside it looks exactly like the model ignored its tools and wrote prose instead.
This is not hypothetical. Two models with completely different tool-call templates — one emitting a JSON blob, the other XML with a raw-text body — produced the identical failure: minutes of work, zero tool calls, no files written, and a correct, detailed analysis sitting in the answer text. Two templates, one symptom, which is what rules out the model and the template and leaves the server configuration.
tool_calls is empty and content reads like thinking, it is configurationA thinking model whose reasoning is not separated cannot call tools. Before you change the prompt,
change the model, or conclude the model is weak at tool use, check that the reasoning is arriving in
its own field. If reasoning is absent while content reads like a train of thought, that is the
bug.
The platform enforces the invariant that makes this impossible to misconfigure: an endpoint that serves tool calling always serves a reasoning parser too. The parser is derived from the model's family, and a family the platform does not recognise gets no parser rather than a wrong one, because a wrong value stops the engine from starting at all.
Silent degradation, and how it is caught
The engine has one more way to lose the separation: if it cannot auto-initialise the reasoning token IDs for a model, it does not abort. It falls back to a pass-through parser that treats the whole output as content — the same symptom as having no parser, with no error on the request.
ColabHive detects that at replica readiness and records it as an error on the replica, so the
condition is visible in operations instead of being discovered as "the model stopped using tools".
If you are self-hosting, look for Auto-initialization of reasoning token IDs failed in the engine's
startup log: nothing about the requests themselves will tell you.
Token budget
The trace spends the same token budget as the answer. With a small max_tokens a thinking model
can spend all of it reasoning and return an empty content with finish_reason: "length".
If a reasoning model returns blank answers, raise max_tokens before changing anything else.
Which models reason
The model catalog marks it per model. As a rule the thinking families on ColabHive are Qwen3 and its distillations, DeepSeek-R1, GLM-4.5/4.7, gpt-oss, Nemotron-3 and Mistral's reasoning line; the non-thinking ones — Qwen2.5, Gemma-2, CodeLlama, Phi-3.5 — have no trace to separate and are unaffected by everything on this page.
See also Tool calling and OpenAI-compatible API.