Skip to main content

Reasoning

A thinking model works before it answers. It emits a reasoning trace and then the answer, and the platform returns them as two separate fields, not one blob of text.

That separation is not cosmetic. If the trace is not separated, it lands inside the answer — and the machinery that reads the answer, including tool calling, stops working. This page says exactly what the contract is, in every layer, so you can tell a model problem from a configuration one.

The two fields​

fieldwhat it holds
choices[0].message.contentthe answer. Always the user-facing text.
choices[0].message.reasoningthe reasoning trace. Canonical name.
choices[0].message.reasoning_contentthe same trace, kept populated as a legacy alias.

Read reasoning. The engine renamed this field, and its own documentation warns that a client reading the old name can get an empty string in silence while the new one is populated — so a client that only knows reasoning_content can conclude "this model does not reason" when it does. ColabHive keeps both names populated so neither client breaks, but new code should read reasoning.

message = response["choices"][0]["message"]
answer = message["content"]
reasoning = message.get("reasoning") or message.get("reasoning_content") or ""

A model that does not reason simply has no reasoning field. That is an answer, not an error.

Streaming​

In a stream the trace arrives in its own delta key, interleaved with the answer's:

data: {"choices":[{"delta":{"reasoning":"Let me check the table..."}}]}
data: {"choices":[{"delta":{"content":"7 + 5 = 12"}}]}
data: [DONE]

Read both keys on the delta and append them to separate buffers — a client that appends every delta into one string gets the trace glued to the answer, which is the same failure as not separating it at all.

answer, reasoning = [], []
for chunk in stream:
delta = chunk["choices"][0].get("delta") or {}
if delta.get("content"):
answer.append(delta["content"])
trace = delta.get("reasoning") or delta.get("reasoning_content")
if trace:
reasoning.append(trace)

When ColabHive stores or retries a streamed result it rebuilds a complete, non-streaming ChatCompletion from those deltas, and that rebuilt message carries both field names with the same value. So the shape you get from a stream and the shape you get from a plain request agree.

Reasoning and tool calling: the failure that looks like a bad model​

The tool-call parser only looks for calls in content. It never looks in the reasoning.

So if the trace is not separated, a thinking model that decides to call write emits that call inside its reasoning, the parser does not see it, and tool_calls comes back empty. From the outside it looks exactly like the model ignored its tools and wrote prose instead.

This is not hypothetical. Two models with completely different tool-call templates — one emitting a JSON blob, the other XML with a raw-text body — produced the identical failure: minutes of work, zero tool calls, no files written, and a correct, detailed analysis sitting in the answer text. Two templates, one symptom, which is what rules out the model and the template and leaves the server configuration.

If tool_calls is empty and content reads like thinking, it is configuration

A thinking model whose reasoning is not separated cannot call tools. Before you change the prompt, change the model, or conclude the model is weak at tool use, check that the reasoning is arriving in its own field. If reasoning is absent while content reads like a train of thought, that is the bug.

The platform enforces the invariant that makes this impossible to misconfigure: an endpoint that serves tool calling always serves a reasoning parser too. The parser is derived from the model's family, and a family the platform does not recognise gets no parser rather than a wrong one, because a wrong value stops the engine from starting at all.

Silent degradation, and how it is caught​

The engine has one more way to lose the separation: if it cannot auto-initialise the reasoning token IDs for a model, it does not abort. It falls back to a pass-through parser that treats the whole output as content — the same symptom as having no parser, with no error on the request.

ColabHive detects that at replica readiness and records it as an error on the replica, so the condition is visible in operations instead of being discovered as "the model stopped using tools". If you are self-hosting, look for Auto-initialization of reasoning token IDs failed in the engine's startup log: nothing about the requests themselves will tell you.

Token budget​

The trace spends the same token budget as the answer. With a small max_tokens a thinking model can spend all of it reasoning and return an empty content with finish_reason: "length".

If a reasoning model returns blank answers, raise max_tokens before changing anything else.

Which models reason​

The model catalog marks it per model. As a rule the thinking families on ColabHive are Qwen3 and its distillations, DeepSeek-R1, GLM-4.5/4.7, gpt-oss, Nemotron-3 and Mistral's reasoning line; the non-thinking ones — Qwen2.5, Gemma-2, CodeLlama, Phi-3.5 — have no trace to separate and are unaffected by everything on this page.

See also Tool calling and OpenAI-compatible API.