Skip to main content

Tool Calling (Function Calling)

ColabHive's OpenAI-compatible Chat Completions API supports tool calling (a.k.a. function calling): you pass a list of tools, and a capable model decides when to call one and returns a structured tool_calls response. Your application executes the tool and sends the result back.

This is what powers agentic clients (coding agents, deep-research agents, assistants) on top of ColabHive — the model does the reasoning and decides the calls; your code runs the tools, so sensitive data and side effects stay on your side.

Intel Arc native

Tool calling runs on ColabHive's Intel Arc (and NVIDIA) fleet — the same OpenAI tool-calling contract, served on Arc GPUs.

Endpoint​

POST https://api.colabhive.com/v1/chat/completions
Authorization: Bearer hive_...
Content-Type: application/json

Use any OpenAI-compatible SDK by pointing its base URL at https://api.colabhive.com/v1 and using your ColabHive API key.

Which models support it​

Tool calling is enabled per model (the model must ship a tool-aware chat template). Instruction-tuned models such as qwen-2.5-7b-instruct support it. See Available Models for the current list, or query the live catalog at GET /api/builder/v1/inference/models.

Example​

A single round trip where the model decides to call get_weather:

curl https://api.colabhive.com/v1/chat/completions \
-H "Authorization: Bearer $COLABHIVE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen-2.5-7b-instruct",
"messages": [
{"role": "user", "content": "What is the weather in Paris right now? Use the get_weather tool."}
],
"tools": [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current weather for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"]
}
}
}]
}'

Response (finish_reason: "tool_calls"):

{
"object": "chat.completion",
"model": "qwen-2.5-7b-instruct",
"choices": [{
"index": 0,
"message": {
"role": "assistant",
"content": null,
"tool_calls": [{
"id": "chatcmpl-tool-ac564de1",
"type": "function",
"function": {"name": "get_weather", "arguments": "{\"city\": \"Paris\"}"}
}]
},
"finish_reason": "tool_calls"
}]
}

Completing the loop​

Execute the tool in your application, then send the result back as a role: "tool" message referencing the tool_call_id:

from openai import OpenAI

client = OpenAI(base_url="https://api.colabhive.com/v1", api_key="hive_...")

messages = [{"role": "user", "content": "Weather in Paris? Use the tool."}]
resp = client.chat.completions.create(model="qwen-2.5-7b-instruct", messages=messages, tools=tools)

call = resp.choices[0].message.tool_calls[0]
messages.append(resp.choices[0].message) # the assistant's tool_calls
messages.append({ # your tool's result
"role": "tool",
"tool_call_id": call.id,
"content": '{"tempC": 18, "sky": "clear"}',
})
final = client.chat.completions.create(model="qwen-2.5-7b-instruct", messages=messages, tools=tools)
print(final.choices[0].message.content)

Thinking models​

If the model reasons before answering, its reasoning arrives in message.reasoning, separate from message.content. That separation is what makes tool calling work at all:

The tool-call parser only looks for calls in content. It never looks in the reasoning. So a thinking model that decides to call write inside its reasoning produces an empty tool_calls if the reasoning is not separated — and from the outside it looks exactly like the model ignored its tools and wrote prose.

ColabHive enforces the invariant that prevents it: an endpoint that serves tool calling always serves a reasoning parser, derived from the model's family. You do not configure anything.

Empty tool_calls with a thinking-looking content is configuration, not the model

Before you change the prompt, raise the temperature or switch models: check whether reasoning came back. If reasoning is absent and content reads like a train of thought, the reasoning is not being separated. On the hosted platform this should not happen; if you self-host, see Reasoning.

Two models with completely different tool-call templates — one a JSON blob, the other XML with a raw-text body — produced the identical failure this way: minutes of work, zero tool calls, no files written, with a correct and detailed analysis sitting in the answer text. Two templates and one symptom is what rules out the model.

And give them budget: the reasoning spends the same max_tokens as the answer, so a thinking model with a tight limit can spend it all reasoning and return an empty content with finish_reason: "length" — and no tool call.

Notes​

  • tool_choice — "auto" (default), "none", "required", or a specific function are supported and forwarded to the model.
  • Reasoning — read message.reasoning (the canonical field); reasoning_content is kept populated as a legacy alias. In a stream both arrive as their own delta keys. See Reasoning.
  • Streaming — stream: true forwards incremental token deltas from current node runtimes and ends with [DONE]. Older nodes fall back to one complete SSE chunk plus [DONE]; clients can detect the mode from the colabhive-stream-mode SSE comment.
  • Multi-turn — role: "tool" and assistant messages carrying tool_calls round-trip correctly.

Use it from an agent​

To run a coding agent such as OpenCode against ColabHive, see Coding Agents: colabhive agents init configures it and checks that the model returns tool calls. For SuperClaw, see Use ColabHive with SuperClaw / OpenCode.