CLM-8B (Contrastive LM)
A scorer, not a generator: it answers typed questions about a state and ranks the candidates you give it. It never writes text.
Overview
- Endpoint:
clm-v0.1-8b(id43e0c9de-e380-4542-afae-7df190bb9c26), display name "CLM-8B (Contrastive LM, Qwen3-8B encoder)". Itstask_typeisclm. - What it is: a frozen Qwen/Qwen3-8B
encoder (revision
b968826d9c46dd6066d109eabc6255188de91218, bf16) reads the state and every candidate; the hidden state after the final norm, at the last real token (4096 values), goes through two small MLP heads — one for the state, one for the candidates — and is L2-normalized. Each candidate's score ismin(exp(logit_scale), 100) · cos(state, candidate) / temperature, and a softmax over one question's candidates is that question's answer. - Reference head: Contrastive-LM/CLM-v0.1-8B
at revision
e939398d4556fcd9400c76fa8c5a513202f42b0a, fileCLM_v0.1-8B.pt(75,557,149 bytes, sha256b2b4a8c9c2d39263eff78a351eb909a342ce9b3bf21a3f07c1d1bf15f1c4eda5). It is what a request uses unlessmodelnames the endpoint of a head you trained. - Routes:
POST /v1/systemoneandPOST /v1/rank, with the wire format of the upstream CLM server plus amodelfield, or the nativePOST /api/builder/v1/endpoints/{id}/inferwith{"system_one": {...}}or{"rank": {...}}asinput. - Trainable: the value is in domain heads. Train one with the
clm-head-infoncetemplate — see Fine-tune a CLM head. A trained head is served by this same model, one request at a time: no second copy of the encoder is loaded for it. - Not a chat model and not in Cohort: CLM endpoints are not listed by
GET /v1/models, are not accepted by/v1/chat/completions, and are refused by Cohort Latent Fabric, which runs only causal text-generation LLMs.
When to use
✅ Scoring a small, known set of options against a context: routing a ticket to a team, checking whether a proposed agent action is appropriate for the current state, picking the best of several candidate answers, or turning a rubric into a number. The answer is a probability distribution you can threshold, not text you have to parse.
✅ A domain you have labeled data for. Train a head on it; the reference head is a starting point, not a finished classifier (see Known limitations).
❌ Generating, summarizing or rewriting text: use an LLM.
❌ Retrieval over a large corpus: embed with Qwen3 Embedding 8B and rerank the top results with Qwen3 Reranker 8B. CLM scores at most 256 candidates per request.
❌ Text that is not English (see Known limitations).
Input contract
POST /v1/systemone
A state and a map of named, typed questions. Every question is answered against the same state.
{
"model": "clm-v0.1-8b",
"state": "Customer: my invoice was charged twice and nobody answers the phone!",
"questions": {
"urgency": {"type": "noul", "instructions": "Is this urgent?"},
"department": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "Charges, invoices, refunds", "technical": "Bugs and outages"}},
"frustration": {"type": "score", "instructions": "How frustrated is the customer?",
"criteria": ["Calm", "Frustrated", "Very angry"]}
}
}
Question type | criteria | Answer |
|---|---|---|
noul (yes/no) | optional {"true": "...", "false": "..."} descriptions | noul: the probability that the statement is true |
choice | {option: description}, at least one option; an empty description uses the key itself | choice (the most probable key), confidence, probabilities per key |
score | an ordered list of at least two levels | score (the expected level, from 0 to levels − 1), confidence, probabilities per level index, and legend (index → level text) |
stateis a string, an object or an array. Objects are read askey: valuelines and arrays as- itemlines, in the order you send them. Put the question ininstructions, not inside the state: the heads were trained on "context, blank line, question".- A list of chat messages (
[{"role": ..., "content": ...}]) is accepted asstateonly by a head trained on chat-formatted states (training taskclm). The reference head reads prose and refuses it. temperatureis optional, greater than 0 and at most 100, default1. Above 1 it flattens the distributions, below 1 it sharpens them.confidenceis the top probability minus the mean of the others, between 0 and 1.
Response — the numbers below are the output of this exact request on the public endpoint, measured on
{{MEASURED:clm.example.evidence_date}}:
{
"model": "clm-v0.1-8b",
"answers": {
"urgency": {"type": "noul", "noul": {{MEASURED:clm.example.urgency_noul}}},
"department": {"type": "choice", "choice": "{{MEASURED:clm.example.department_choice}}",
"confidence": {{MEASURED:clm.example.department_confidence}},
"probabilities": {"billing": {{MEASURED:clm.example.department_billing_prob}},
"technical": {{MEASURED:clm.example.department_technical_prob}}}},
"frustration": {"type": "score", "score": {{MEASURED:clm.example.frustration_score}},
"confidence": {{MEASURED:clm.example.frustration_confidence}},
"legend": {"0": "Calm", "1": "Frustrated", "2": "Very angry"},
"probabilities": {"0": {{MEASURED:clm.example.frustration_p0}},
"1": {{MEASURED:clm.example.frustration_p1}},
"2": {{MEASURED:clm.example.frustration_p2}}}}
},
"usage": {"billing_units": 3, "input_tokens": {{MEASURED:clm.example.input_tokens}}, "output_tokens": 0},
"colabhive": {
"endpoint_id": "43e0c9de-e380-4542-afae-7df190bb9c26",
"task_id": "…",
"head_sha256": "b2b4a8c9c2d39263eff78a351eb909a342ce9b3bf21a3f07c1d1bf15f1c4eda5",
"encoder_identity": "81376c49813a49f1f3fc1a0cb876a8b369f216e03edd9c669a26ed1a41f4942a",
"runtime_vendor": "{{MEASURED:clm.example.runtime_vendor}}"
}
}
POST /v1/rank
Free-form candidate answers for a context and a question, returned best first. It is a choice
question whose options are the answers.
{
"model": "clm-v0.1-8b",
"context": "Tides are caused mainly by",
"question": "Which answer completes the sentence?",
"answers": ["the gravity of the Moon", "the wind over the ocean"]
}
{
"model": "clm-v0.1-8b",
"ranked": [
{"rank": 1, "candidate": "{{MEASURED:clm.example.rank_first_candidate}}", "prob": {{MEASURED:clm.example.rank_first_prob}}},
{"rank": 2, "candidate": "{{MEASURED:clm.example.rank_second_candidate}}", "prob": {{MEASURED:clm.example.rank_second_prob}}}
],
"usage": {"billing_units": 1, "input_tokens": {{MEASURED:clm.example.rank_input_tokens}}, "output_tokens": 0},
"colabhive": {"endpoint_id": "43e0c9de-e380-4542-afae-7df190bb9c26", "task_id": "…",
"head_sha256": "b2b4a8c9c2d39263eff78a351eb909a342ce9b3bf21a3f07c1d1bf15f1c4eda5",
"encoder_identity": "81376c49813a49f1f3fc1a0cb876a8b369f216e03edd9c669a26ed1a41f4942a",
"runtime_vendor": "{{MEASURED:clm.example.runtime_vendor}}"}
}
What ColabHive adds to the upstream answer
modelis a ColabHive endpoint: its UUID, or itsendpoint_name/display_name. It is never a Hugging Face repo id (400 hf_repo_not_accepted) and never a runtime key containing:trained:(400 runtime_key_not_accepted: those keys are not bound to an account). To use a head you trained, pass the id of the endpoint you registered for it.- The
colabhiveblock names the endpoint and task that served the request, the sha256 of the head that scored it, the identity of the encoder it ran on, and the GPU vendor. A client written for the upstream server ignores it. - The
X-CLM-Latency-Msresponse header is the end-to-end time inside the API: queue, dispatch and encoding.
Limits
Per request, checked before anything is scored:
| Limit | Value | Over the limit |
|---|---|---|
Candidates (a noul question counts 2, a choice or score question one per option, /v1/rank one per answer) | 256 | 413 clm_too_many_candidates |
Questions in one /v1/systemone request | 64 | 413 too_many_questions |
| Encoder tokens, counted exactly after truncation | 32768 | 413 too_many_tokens |
| Encoder tokens the gateway lets one request cost, estimated as one token per UTF-8 byte of each text, so that it finishes within the API edge limit | {{MEASURED:clm.gateway_encoder_token_bound}} | 413 clm_input_too_large |
| JSON body | 12517376 bytes | 413 clm_input_too_large |
- Each text is cut to the head's window before it is encoded, and never errors for being long: 8,192
tokens for the reference head and heads trained with task
clm, 2,048 for heads trained with taskchoice. A state keeps its end (where the question is); a candidate keeps its beginning forclmheads and its end forchoiceheads. - The state is encoded once per distinct
instructions, and each distinct candidate text once, so questions that share instructions and options cost less than their count suggests. 413responses carryx-should-retry: false: split the request, do not resend it.
Errors
Errors use the OpenAI error object ({"error": {"message", "type", "param", "code"}}).
| Status | code | Meaning |
|---|---|---|
| 400 | model_not_clm | model names an endpoint that is not a CLM scorer. Use /v1/chat/completions for language models. |
| 400 | hf_repo_not_accepted, runtime_key_not_accepted | model is a Hugging Face repo id or a runtime key. |
| 400 | invalid_request, non_finite_logits, others | The model refused the request as invalid (for example a chat-message state on a prose head). Not retried. |
| 401 | — | Missing or invalid API key. |
| 404 | — | No CLM endpoint with that id or label that your account may call. |
| 409 | ambiguous_model | The label matches several CLM endpoints; the message lists their ids. |
| 413 | see Limits | Too many candidates, questions or tokens. |
| 502 | — | The model failed or returned an unexpected result. |
| 503 | the wait reason, endpoint_lookup_unavailable, or a capacity code | The model was not ready within the 300 s the API edge allows, or had no capacity; retry after Retry-After. The model keeps loading after the cancellation. |
| 504 | — | The request was running but could not finish within the edge limit: send fewer questions or shorter texts. |
Known limitations of the reference head
These are properties of the published reference head, reported upstream and reproduced before publication. A head trained on your data is how they are addressed.
scorequestions are not reliable zero-shot. The reference head tends to put its mass on an extreme level of the rubric (upstream issue #3). Train a head before you act on ascore.- Routing zero-shot is weak. Choosing among many intents it was not trained on (the upstream
Banking77 report, issue #13) is close to chance; the same task after fine-tuning is not. Treat
zero-shot
choiceover many options as a baseline to beat. - English only. Other languages score close to chance (upstream issue #22).
- Similar states look alike. States that differ only before the final question are encoded close
together (upstream issue #15): put the distinguishing facts in the state, and keep the question in
instructions.
Hardware and status
One replica holds one encoder on one GPU. Heads — the reference head or any head you trained — travel with each request and are cached next to the encoder in a fixed memory budget, so every head shares one copy of the encoder.
| Hardware | Inference | Head training | Measured GPU memory | Evidence date |
|---|---|---|---|---|
| NVIDIA RTX 3090 24 GB | {{MEASURED:clm.status.inference.nvidia_rtx3090}} | {{MEASURED:clm.status.training.nvidia_rtx3090}} | {{MEASURED:clm.vram_mb.nvidia_rtx3090}} MB | {{MEASURED:clm.evidence_date.nvidia_rtx3090}} |
| Intel Arc Pro B70 32 GB | {{MEASURED:clm.status.inference.intel_b70}} | {{MEASURED:clm.status.training.intel_b70}} | {{MEASURED:clm.vram_mb.intel_b70}} MB | {{MEASURED:clm.evidence_date.intel_b70}} |
| Intel, two 16 GB GPUs (layers split across both) | {{MEASURED:clm.status.inference.intel_2x16gb}} | not offered | {{MEASURED:clm.vram_mb.intel_2x16gb}} MB | {{MEASURED:clm.evidence_date.intel_2x16gb}} |
AMD Instinct {{MEASURED:clm.amd.board}} | {{MEASURED:clm.status.inference.amd_instinct}} | {{MEASURED:clm.status.training.amd_instinct}} | {{MEASURED:clm.vram_mb.amd_instinct}} MB | {{MEASURED:clm.evidence_date.amd_instinct}} |
| CPU | {{MEASURED:clm.status.inference.cpu}} | {{MEASURED:clm.status.training.cpu}} | — (RAM {{MEASURED:clm.ram_mb.cpu}} MB) | {{MEASURED:clm.evidence_date.cpu}} |
Every vendor serves the same encoder revision in bf16, nothing below it, and the same head gives the
same answers on each: a head trained on one vendor serves on the others. Before a vendor is listed as
ready its outputs were compared against an NVIDIA bf16 and a CPU fp32 reference on the same inputs.
End-to-end latency of the /v1/systemone example above on a warm replica, as the API sees it
(X-CLM-Latency-Ms):
| Hardware | p50 | p95 | Samples | Measured on |
|---|---|---|---|---|
| NVIDIA RTX 3090 | {{MEASURED:clm.latency_p50_ms.nvidia_rtx3090}} ms | {{MEASURED:clm.latency_p95_ms.nvidia_rtx3090}} ms | {{MEASURED:clm.latency_samples.nvidia_rtx3090}} | {{MEASURED:clm.evidence_date.nvidia_rtx3090}} |
| Intel Arc Pro B70 | {{MEASURED:clm.latency_p50_ms.intel_b70}} ms | {{MEASURED:clm.latency_p95_ms.intel_b70}} ms | {{MEASURED:clm.latency_samples.intel_b70}} | {{MEASURED:clm.evidence_date.intel_b70}} |
| Intel, two 16 GB GPUs | {{MEASURED:clm.latency_p50_ms.intel_2x16gb}} ms | {{MEASURED:clm.latency_p95_ms.intel_2x16gb}} ms | {{MEASURED:clm.latency_samples.intel_2x16gb}} | {{MEASURED:clm.evidence_date.intel_2x16gb}} |
| AMD Instinct | {{MEASURED:clm.latency_p50_ms.amd_instinct}} ms | {{MEASURED:clm.latency_p95_ms.amd_instinct}} ms | {{MEASURED:clm.latency_samples.amd_instinct}} | {{MEASURED:clm.evidence_date.amd_instinct}} |
| CPU | {{MEASURED:clm.latency_p50_ms.cpu}} ms | {{MEASURED:clm.latency_p95_ms.cpu}} ms | {{MEASURED:clm.latency_samples.cpu}} | {{MEASURED:clm.evidence_date.cpu}} |
The first request with a head that is not yet on the node also downloads it (75 MB for the reference
head): {{MEASURED:clm.head_cold_start_p50_ms}} ms p50, {{MEASURED:clm.head_cold_start_p95_ms}} ms
p95, measured on {{MEASURED:clm.evidence_date.head_cold_start}}. A replica that is not loaded
answers 503 with Retry-After while the encoder loads (see Errors). These are
measurements, not guarantees; measure on your own nodes and inputs.
Training a domain head — hyperparameters
Model ID: clm-head-infonce
The template trains the two heads over the frozen encoder, in one job on one GPU. The full walkthrough, with the dataset formats, is in Fine-tune a CLM head.
| Parameter | Default | Range / options | Notes |
|---|---|---|---|
task | "clm" | "clm", "choice" | clm: state/action pairs (contrastive). choice: multiple-choice questions in the System One format. |
loss | "infonce" | "infonce", "softce" | softce requires task choice. |
targets | "soft" | "soft", "hard" | Task choice: soft keeps the gold distribution, hard its most probable label. |
epochs | 20 | 1 to 200 | Maximum epochs of the head training; early stopping may end sooner. |
patience | 5 | 1 to 200 | Epochs without a better validation score before training stops. |
seed | 1234 | 0 to 2147483647 | Seed of the data split and of the head training. |
batch | null | 8 to 8192 | null = task default: 2048 for clm, 256 for choice. |
lr | null | greater than 0, at most 1.0 | null = task default: 2e-3 · √(1024/width) · √(batch/1024) for clm, 5e-4 for choice. |
max_tokens | null | 16 to 8192 | Encoder tokens kept per text. null = task limit: 8192 for clm, 2048 for choice. A value above the task's limit fails before the GPU is used. |
batch_tokens | null | 16 to 8192 | Padded tokens per encoder pass while the dataset is embedded. null = max_tokens. |
max_total_tokens | 50000000 | 1 to 200000000 | Upper bound on the encoder tokens read from the dataset; a larger dataset fails before encoding starts. |
holdout_frac | 0.1 | 0.01 to 0.49 | Share held out for validation (and, for choice, for test). |
Licenses
- Encoder weights: Qwen/Qwen3-8B, Apache-2.0 (Alibaba Cloud). Reference head: Contrastive-LM/CLM-v0.1-8B, Apache-2.0. See Model licenses.
- The serving and training code (
clm_runtime) includes modified code from Contrastive-LM/CLM at commitbb42c6cand is Apache-2.0, with its ownLICENSEandNOTICE. The ColabHive node runtime, SDK and MCP server are MIT; the orchestrator is under the Business Source License 1.1. - A head you train inherits the licenses of the encoder and of the head it started from.
Cost
CLM has no per-model price. Requests served on your own nodes are not charged; on capacity ColabHive operates they are metered like any other inference, per accelerator-hour — see Pricing.
Quick start
curl -X POST "https://api.colabhive.com/v1/rank" \
-H "Authorization: Bearer $COLABHIVE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "clm-v0.1-8b",
"context": "Tides are caused mainly by",
"question": "Which answer completes the sentence?",
"answers": ["the gravity of the Moon", "the wind over the ocean"]}'
from colabhive import ColabHive
from colabhive.clm import Choice, Noul, Score
client = ColabHive(api_key="hive_...", account_id="YOUR_ACCOUNT_ID")
r = client.clm.system_one(
model="clm-v0.1-8b",
state="Customer: my invoice was charged twice and nobody answers the phone!",
questions={
"urgency": Noul(instructions="Is this urgent?"),
"department": Choice(instructions="Which team should handle this?",
criteria={"billing": "Charges, invoices, refunds",
"technical": "Bugs and outages"}),
"frustration": Score(instructions="How frustrated is the customer?",
criteria=["Calm", "Frustrated", "Very angry"]),
},
)
print(r.answers["department"].choice, r.answers["urgency"].noul, r.latency_ms)
client.clm is in the Python SDK from colabhive 0.10.0 — see the SDK reference.
Authors: José Luis Minich, Maximiliano Lucius.