Skip to main content

Agent & LLM Integration

ColabHive provides machine-readable resources designed for AI agents, LLM tool-use, and automated integrations.


Machine-Readable Resources​

ResourceURLPurpose
llm.txt/llm.txtConcise summary following llmstxt.org standard. Lists all endpoints, models, and capabilities in a format optimized for LLM context windows (~4K tokens).
llms-full.txt/llms-full.txtComplete API reference in a single markdown file (~12K tokens). Contains every endpoint with request/response schemas and examples.
OpenAPI 3.1/openapi.jsonFull OpenAPI specification for tool-use integrations (Cursor, Copilot, custom agents). Import directly into any OpenAPI-compatible tool.
Swagger UIapi.colabhive.com/api/builder/v1/docsInteractive API explorer with try-it-out functionality.
ReDocapi.colabhive.com/api/builder/v1/redocAlternative API documentation viewer.

The public URLs above describe the full Builder contract, including the webhook management routes. Webhook event schemas and receiver snippets are documented in Outbound Webhooks.


For AI Agent Developers​

Using ColabHive as a Tool​

If you're building an AI agent that needs to train or run ML models, ColabHive can be integrated as a tool:

1. Discovery — what models are available and ready?

GET /training/model-configs # Trainable models
GET /inference/models?include_readiness=true # All models + warm/cached/cold status
GET /endpoints?visibility=public&include_readiness=true # Public endpoints + readiness
GET /capabilities # Supported capability types

2. Training — train a model on user data:

POST /datasets/upload # Upload dataset
POST /training/runs # Start training
GET /training/runs/{id} # Poll progress
POST /training/runs/{id}/register-for-inference # Deploy

3. Inference — run predictions:

POST /endpoints/{id}/infer # Sync inference (sync: true) or async (sync: false)
GET /tasks/{task_id} # Poll async task results
POST /endpoints/{id}/upload-input # Upload binary files (image, audio)
PATCH /endpoints/{id}/scaling # Set scale-to-zero or a durable replica floor

scaling_mode: "on_demand" requires min_replicas: 0. Use scaling_mode: "minimum" with min_replicas >= 1 to keep capacity resident. The response includes current_replicas, the READY replicas measured now; when the platform cannot establish the count the key is left out, which means unknown, not zero.

Placement is the platform's job. There is no node parameter on inference, and an agent has no reason to pick one. required_node_ids exists to hard-pin the minimum replicas to named hardware when a contract demands it — the API rejects a list longer than the floor instead of silently weakening placement. It is a rare, deliberate choice: pinning removes the platform's ability to route around a busy or failing node, so it makes availability worse, never latency better.

Streaming & Explicit Async

POST /endpoints/{id}/infer:async is available and returns HTTP 202 with a task_id. Token-by-token streaming is served by /v1/chat/completions with "stream": true; there is no :stream route on the typed inference surface.

Optional terminal notifications​

Outbound webhooks notify an HTTPS receiver when inference tasks, Action/MCP invocations, training runs or merges finish when the extension is enabled for an account. They are not generic Actions, and polling remains the universal contract. Builder REST is authoritative; the Python SDK and MCP adapters preserve its roles, conditional updates, idempotency, and one-time secret rules. MCP tools require the false-by-default COLABHIVE_MCP_WEBHOOK_TOOLS_ENABLED gate; create and rotate additionally require COLABHIVE_ALLOW_WEBHOOK_SECRET_TOOLS.

OpenAPI Tool-Use Integration​

Import the OpenAPI spec into your agent framework:

LangChain:

from langchain_community.tools.openapi import OpenAPIToolkit
toolkit = OpenAPIToolkit.from_openapi_url("https://docs.colabhive.com/openapi.json")

Cursor / Copilot: Add to your project's .cursorrules or tool configuration:

API Spec: https://docs.colabhive.com/openapi.json

Custom Agent (direct HTTP):

import httpx

BASE = "https://api.colabhive.com/api/builder/v1"
HEADERS = {"X-API-Key": KEY, "X-Account-ID": ACCT}

# List available trainable models
models = httpx.get(f"{BASE}/training/model-configs", headers=HEADERS).json()

# Train
job = httpx.post(f"{BASE}/training/runs", headers=HEADERS, json={
"job_name": "auto-train",
"model_config_id": models["model_configs"][0]["model_id"],
"dataset_id": dataset_id,
"hyperparameters": {}
}).json()

# Poll
while True:
status = httpx.get(f"{BASE}/training/runs/{job['run_id']}", headers=HEADERS).json()
if status["status"] in ("completed", "failed"):
break
time.sleep(10)

# Deploy
endpoint = httpx.post(
f"{BASE}/training/runs/{job['run_id']}/register-for-inference",
headers=HEADERS,
json={"name": "auto-model", "description": "Auto-trained", "visibility": "account"}
).json()

# Infer
result = httpx.post(
f"{BASE}/endpoints/{endpoint['endpoint_id']}/infer",
headers=HEADERS,
json={"input": {"features": {"col1": 1.0}}}
).json()

Intelligent Model Routing (Readiness-Aware)​

Models have three readiness states that affect inference latency:

StateMeaningTypical LatencyBest For
warmLoaded in GPU VRAMWorkload-dependentSync inference without a model-load step
cachedOn disk, not loaded5-30sBackground tasks, async
coldNot on any node30-120sBatch jobs, async only

Why This Matters for Agents​

If your agent needs to call an LLM and multiple options are available (e.g., Qwen 7B and Mistral 7B are both chat models), choosing the warm one saves 30-120 seconds of cold start time.

Decision Tree​

1. Query: GET /inference/models?include_readiness=true&task_type=chat
2. Filter models by readiness:
├── warm models exist?
│ └── YES → Use warm model, sync mode, timeout=10s
├── cached models exist?
│ └── YES → Use cached model, sync mode, timeout=30s
└── All cold?
└── Use async mode, poll /tasks/{id}, notify user of delay

Readiness-Aware Agent Example​

import httpx

BASE = "https://api.colabhive.com/api/builder/v1"
HEADERS = {"X-API-Key": KEY, "X-Account-ID": ACCT, "Content-Type": "application/json"}

# Step 1: Discover models with readiness
models = httpx.get(
f"{BASE}/inference/models?include_readiness=true&task_type=chat",
headers=HEADERS
).json()

# Step 2: Sort by readiness — warm first, then cached, then cold
READINESS_RANK = {"warm": 0, "cached": 1, "cold": 2}
models_list = models if isinstance(models, list) else models.get("models", [])
models_list.sort(key=lambda m: READINESS_RANK.get(m.get("readiness", "cold"), 3))

best = models_list[0]
readiness = best.get("readiness", "cold")

# Step 3: Set timeout and mode based on readiness
if readiness == "warm":
sync, timeout = True, 10.0
elif readiness == "cached":
sync, timeout = True, 30.0
else:
sync, timeout = False, None # Async for cold models

# Step 4: Find the endpoint for this model
endpoints = httpx.get(
f"{BASE}/endpoints?visibility=public&task_type=chat&include_readiness=true",
headers=HEADERS
).json()
warm_endpoint = next(
(ep for ep in endpoints["endpoints"] if ep.get("readiness") == "warm"),
endpoints["endpoints"][0] # Fallback to first available
)

# Step 5: Run inference
payload = {
"input": {"messages": [{"role": "user", "content": "Hello!"}]},
"sync": sync,
}
if sync:
payload["sync_timeout_s"] = timeout

result = httpx.post(
f"{BASE}/endpoints/{warm_endpoint['endpoint_id']}/infer",
headers=HEADERS,
json=payload,
).json()

if result["status"] == "succeeded":
print(result["result"]) # Immediate result
else:
task_id = result["task_id"]
# Poll GET /tasks/{task_id} until completed
Readiness is Dynamic

Model readiness changes based on usage patterns — popular models stay warm automatically. If you consistently use a model, it will likely be warm for subsequent requests. The readiness data is cached for ~15 seconds, so it's safe to query frequently.


Input Schemas by Model Type​

Agents need to know the correct input format for each model type:

Model TypeInput SchemaExample
LLM (chat){"messages": [{"role": "user", "content": "..."}]}Chat completion
Embeddings{"text": "..."}Single text embedding
Rerank{"query": "...", "documents": ["..."]}Document reranking
Tabular (regression){"features": {"col1": 1.0, "col2": "val"}}Predict numeric value
Tabular (classification){"features": {"col1": 1.0, "col2": "val"}}Predict class
Forecasting{"values": [1.2, 3.4, ...], "horizon": 10}Time series forecast
Forecasting (multivariate){"series": {"close": [...], "volume": [...]}, "horizon": 5}Multi-column forecast
Image generation{"prompt": "...", "width": 1024, "height": 1024}Generate image
STT (speech-to-text)Upload audio via /upload-input, then {"input_url": "..."}Transcribe audio
OCRUpload image via /upload-input, then {"input_url": "..."}Extract text from image
Translation{"text": "...", "source_lang": "en", "target_lang": "es"}Translate text
Web fetch{"url": "https://..."}Fetch URL content
Web search{"query": "..."}Search the web
Moderation{"text": "..."}Content safety check

Authentication for Agents​

X-API-Key: hive_... # API key (required)
X-Account-ID: <uuid> # Optional; must match the API key's account when supplied
Content-Type: application/json # For JSON requests

Get your API key from console.colabhive.com. See Authentication for the full contract.


Rate Limits​

Rate limiting is enforced at the infrastructure layer (Nginx). When a limit is exceeded, requests receive HTTP 429 — back off and retry. The platform tracks a per-endpoint rate_limit_rpm, but it is not exposed to clients today.

Rate-limit contract

Plan limits are published at colabhive.com/pricing. Response headers (X-RateLimit-Limit, X-RateLimit-Remaining, X-RateLimit-Reset) are not emitted. Do not parse them; handle 429 by retrying with exponential backoff.