Skip to main content

CLM-8B (Contrastive LM)

A scorer, not a generator: it answers typed questions about a state and ranks the candidates you give it. It never writes text.

Overview​

  • Endpoint: clm-v0.1-8b (id 43e0c9de-e380-4542-afae-7df190bb9c26), display name "CLM-8B (Contrastive LM, Qwen3-8B encoder)". Its task_type is clm.
  • What it is: a frozen Qwen/Qwen3-8B encoder (revision b968826d9c46dd6066d109eabc6255188de91218, bf16) reads the state and every candidate; the hidden state after the final norm, at the last real token (4096 values), goes through two small MLP heads — one for the state, one for the candidates — and is L2-normalized. Each candidate's score is min(exp(logit_scale), 100) · cos(state, candidate) / temperature, and a softmax over one question's candidates is that question's answer.
  • Reference head: Contrastive-LM/CLM-v0.1-8B at revision e939398d4556fcd9400c76fa8c5a513202f42b0a, file CLM_v0.1-8B.pt (75,557,149 bytes, sha256 b2b4a8c9c2d39263eff78a351eb909a342ce9b3bf21a3f07c1d1bf15f1c4eda5). It is what a request uses unless model names the endpoint of a head you trained.
  • Routes: POST /v1/systemone and POST /v1/rank, with the wire format of the upstream CLM server plus a model field, or the native POST /api/builder/v1/endpoints/{id}/infer with {"system_one": {...}} or {"rank": {...}} as input.
  • Trainable: the value is in domain heads. Train one with the clm-head-infonce template — see Fine-tune a CLM head. A trained head is served by this same model, one request at a time: no second copy of the encoder is loaded for it.
  • Not a chat model and not in Cohort: CLM endpoints are not listed by GET /v1/models, are not accepted by /v1/chat/completions, and are refused by Cohort Latent Fabric, which runs only causal text-generation LLMs.

When to use​

✅ Scoring a small, known set of options against a context: routing a ticket to a team, checking whether a proposed agent action is appropriate for the current state, picking the best of several candidate answers, or turning a rubric into a number. The answer is a probability distribution you can threshold, not text you have to parse.

✅ A domain you have labeled data for. Train a head on it; the reference head is a starting point, not a finished classifier (see Known limitations).

❌ Generating, summarizing or rewriting text: use an LLM.

❌ Retrieval over a large corpus: embed with Qwen3 Embedding 8B and rerank the top results with Qwen3 Reranker 8B. CLM scores at most 256 candidates per request.

❌ Text that is not English (see Known limitations).

Input contract​

POST /v1/systemone​

A state and a map of named, typed questions. Every question is answered against the same state.

{
"model": "clm-v0.1-8b",
"state": "Customer: my invoice was charged twice and nobody answers the phone!",
"questions": {
"urgency": {"type": "noul", "instructions": "Is this urgent?"},
"department": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "Charges, invoices, refunds", "technical": "Bugs and outages"}},
"frustration": {"type": "score", "instructions": "How frustrated is the customer?",
"criteria": ["Calm", "Frustrated", "Very angry"]}
}
}
Question typecriteriaAnswer
noul (yes/no)optional {"true": "...", "false": "..."} descriptionsnoul: the probability that the statement is true
choice{option: description}, at least one option; an empty description uses the key itselfchoice (the most probable key), confidence, probabilities per key
scorean ordered list of at least two levelsscore (the expected level, from 0 to levels − 1), confidence, probabilities per level index, and legend (index → level text)
  • state is a string, an object or an array. Objects are read as key: value lines and arrays as - item lines, in the order you send them. Put the question in instructions, not inside the state: the heads were trained on "context, blank line, question".
  • A list of chat messages ([{"role": ..., "content": ...}]) is accepted as state only by a head trained on chat-formatted states (training task clm). The reference head reads prose and refuses it.
  • temperature is optional, greater than 0 and at most 100, default 1. Above 1 it flattens the distributions, below 1 it sharpens them.
  • confidence is the top probability minus the mean of the others, between 0 and 1.

Response — the numbers below are the output of this exact request on the public endpoint, measured on {{MEASURED:clm.example.evidence_date}}:

{
"model": "clm-v0.1-8b",
"answers": {
"urgency": {"type": "noul", "noul": {{MEASURED:clm.example.urgency_noul}}},
"department": {"type": "choice", "choice": "{{MEASURED:clm.example.department_choice}}",
"confidence": {{MEASURED:clm.example.department_confidence}},
"probabilities": {"billing": {{MEASURED:clm.example.department_billing_prob}},
"technical": {{MEASURED:clm.example.department_technical_prob}}}},
"frustration": {"type": "score", "score": {{MEASURED:clm.example.frustration_score}},
"confidence": {{MEASURED:clm.example.frustration_confidence}},
"legend": {"0": "Calm", "1": "Frustrated", "2": "Very angry"},
"probabilities": {"0": {{MEASURED:clm.example.frustration_p0}},
"1": {{MEASURED:clm.example.frustration_p1}},
"2": {{MEASURED:clm.example.frustration_p2}}}}
},
"usage": {"billing_units": 3, "input_tokens": {{MEASURED:clm.example.input_tokens}}, "output_tokens": 0},
"colabhive": {
"endpoint_id": "43e0c9de-e380-4542-afae-7df190bb9c26",
"task_id": "…",
"head_sha256": "b2b4a8c9c2d39263eff78a351eb909a342ce9b3bf21a3f07c1d1bf15f1c4eda5",
"encoder_identity": "81376c49813a49f1f3fc1a0cb876a8b369f216e03edd9c669a26ed1a41f4942a",
"runtime_vendor": "{{MEASURED:clm.example.runtime_vendor}}"
}
}

POST /v1/rank​

Free-form candidate answers for a context and a question, returned best first. It is a choice question whose options are the answers.

{
"model": "clm-v0.1-8b",
"context": "Tides are caused mainly by",
"question": "Which answer completes the sentence?",
"answers": ["the gravity of the Moon", "the wind over the ocean"]
}
{
"model": "clm-v0.1-8b",
"ranked": [
{"rank": 1, "candidate": "{{MEASURED:clm.example.rank_first_candidate}}", "prob": {{MEASURED:clm.example.rank_first_prob}}},
{"rank": 2, "candidate": "{{MEASURED:clm.example.rank_second_candidate}}", "prob": {{MEASURED:clm.example.rank_second_prob}}}
],
"usage": {"billing_units": 1, "input_tokens": {{MEASURED:clm.example.rank_input_tokens}}, "output_tokens": 0},
"colabhive": {"endpoint_id": "43e0c9de-e380-4542-afae-7df190bb9c26", "task_id": "…",
"head_sha256": "b2b4a8c9c2d39263eff78a351eb909a342ce9b3bf21a3f07c1d1bf15f1c4eda5",
"encoder_identity": "81376c49813a49f1f3fc1a0cb876a8b369f216e03edd9c669a26ed1a41f4942a",
"runtime_vendor": "{{MEASURED:clm.example.runtime_vendor}}"}
}

What ColabHive adds to the upstream answer​

  • model is a ColabHive endpoint: its UUID, or its endpoint_name / display_name. It is never a Hugging Face repo id (400 hf_repo_not_accepted) and never a runtime key containing :trained: (400 runtime_key_not_accepted: those keys are not bound to an account). To use a head you trained, pass the id of the endpoint you registered for it.
  • The colabhive block names the endpoint and task that served the request, the sha256 of the head that scored it, the identity of the encoder it ran on, and the GPU vendor. A client written for the upstream server ignores it.
  • The X-CLM-Latency-Ms response header is the end-to-end time inside the API: queue, dispatch and encoding.

Limits​

Per request, checked before anything is scored:

LimitValueOver the limit
Candidates (a noul question counts 2, a choice or score question one per option, /v1/rank one per answer)256413 clm_too_many_candidates
Questions in one /v1/systemone request64413 too_many_questions
Encoder tokens, counted exactly after truncation32768413 too_many_tokens
Encoder tokens the gateway lets one request cost, estimated as one token per UTF-8 byte of each text, so that it finishes within the API edge limit{{MEASURED:clm.gateway_encoder_token_bound}}413 clm_input_too_large
JSON body12517376 bytes413 clm_input_too_large
  • Each text is cut to the head's window before it is encoded, and never errors for being long: 8,192 tokens for the reference head and heads trained with task clm, 2,048 for heads trained with task choice. A state keeps its end (where the question is); a candidate keeps its beginning for clm heads and its end for choice heads.
  • The state is encoded once per distinct instructions, and each distinct candidate text once, so questions that share instructions and options cost less than their count suggests.
  • 413 responses carry x-should-retry: false: split the request, do not resend it.

Errors​

Errors use the OpenAI error object ({"error": {"message", "type", "param", "code"}}).

StatuscodeMeaning
400model_not_clmmodel names an endpoint that is not a CLM scorer. Use /v1/chat/completions for language models.
400hf_repo_not_accepted, runtime_key_not_acceptedmodel is a Hugging Face repo id or a runtime key.
400invalid_request, non_finite_logits, othersThe model refused the request as invalid (for example a chat-message state on a prose head). Not retried.
401—Missing or invalid API key.
404—No CLM endpoint with that id or label that your account may call.
409ambiguous_modelThe label matches several CLM endpoints; the message lists their ids.
413see LimitsToo many candidates, questions or tokens.
502—The model failed or returned an unexpected result.
503the wait reason, endpoint_lookup_unavailable, or a capacity codeThe model was not ready within the 300 s the API edge allows, or had no capacity; retry after Retry-After. The model keeps loading after the cancellation.
504—The request was running but could not finish within the edge limit: send fewer questions or shorter texts.

Known limitations of the reference head​

These are properties of the published reference head, reported upstream and reproduced before publication. A head trained on your data is how they are addressed.

  • score questions are not reliable zero-shot. The reference head tends to put its mass on an extreme level of the rubric (upstream issue #3). Train a head before you act on a score.
  • Routing zero-shot is weak. Choosing among many intents it was not trained on (the upstream Banking77 report, issue #13) is close to chance; the same task after fine-tuning is not. Treat zero-shot choice over many options as a baseline to beat.
  • English only. Other languages score close to chance (upstream issue #22).
  • Similar states look alike. States that differ only before the final question are encoded close together (upstream issue #15): put the distinguishing facts in the state, and keep the question in instructions.

Hardware and status​

One replica holds one encoder on one GPU. Heads — the reference head or any head you trained — travel with each request and are cached next to the encoder in a fixed memory budget, so every head shares one copy of the encoder.

HardwareInferenceHead trainingMeasured GPU memoryEvidence date
NVIDIA RTX 3090 24 GB{{MEASURED:clm.status.inference.nvidia_rtx3090}}{{MEASURED:clm.status.training.nvidia_rtx3090}}{{MEASURED:clm.vram_mb.nvidia_rtx3090}} MB{{MEASURED:clm.evidence_date.nvidia_rtx3090}}
Intel Arc Pro B70 32 GB{{MEASURED:clm.status.inference.intel_b70}}{{MEASURED:clm.status.training.intel_b70}}{{MEASURED:clm.vram_mb.intel_b70}} MB{{MEASURED:clm.evidence_date.intel_b70}}
Intel, two 16 GB GPUs (layers split across both){{MEASURED:clm.status.inference.intel_2x16gb}}not offered{{MEASURED:clm.vram_mb.intel_2x16gb}} MB{{MEASURED:clm.evidence_date.intel_2x16gb}}
AMD Instinct {{MEASURED:clm.amd.board}}{{MEASURED:clm.status.inference.amd_instinct}}{{MEASURED:clm.status.training.amd_instinct}}{{MEASURED:clm.vram_mb.amd_instinct}} MB{{MEASURED:clm.evidence_date.amd_instinct}}
CPU{{MEASURED:clm.status.inference.cpu}}{{MEASURED:clm.status.training.cpu}}— (RAM {{MEASURED:clm.ram_mb.cpu}} MB){{MEASURED:clm.evidence_date.cpu}}

Every vendor serves the same encoder revision in bf16, nothing below it, and the same head gives the same answers on each: a head trained on one vendor serves on the others. Before a vendor is listed as ready its outputs were compared against an NVIDIA bf16 and a CPU fp32 reference on the same inputs.

End-to-end latency of the /v1/systemone example above on a warm replica, as the API sees it (X-CLM-Latency-Ms):

Hardwarep50p95SamplesMeasured on
NVIDIA RTX 3090{{MEASURED:clm.latency_p50_ms.nvidia_rtx3090}} ms{{MEASURED:clm.latency_p95_ms.nvidia_rtx3090}} ms{{MEASURED:clm.latency_samples.nvidia_rtx3090}}{{MEASURED:clm.evidence_date.nvidia_rtx3090}}
Intel Arc Pro B70{{MEASURED:clm.latency_p50_ms.intel_b70}} ms{{MEASURED:clm.latency_p95_ms.intel_b70}} ms{{MEASURED:clm.latency_samples.intel_b70}}{{MEASURED:clm.evidence_date.intel_b70}}
Intel, two 16 GB GPUs{{MEASURED:clm.latency_p50_ms.intel_2x16gb}} ms{{MEASURED:clm.latency_p95_ms.intel_2x16gb}} ms{{MEASURED:clm.latency_samples.intel_2x16gb}}{{MEASURED:clm.evidence_date.intel_2x16gb}}
AMD Instinct{{MEASURED:clm.latency_p50_ms.amd_instinct}} ms{{MEASURED:clm.latency_p95_ms.amd_instinct}} ms{{MEASURED:clm.latency_samples.amd_instinct}}{{MEASURED:clm.evidence_date.amd_instinct}}
CPU{{MEASURED:clm.latency_p50_ms.cpu}} ms{{MEASURED:clm.latency_p95_ms.cpu}} ms{{MEASURED:clm.latency_samples.cpu}}{{MEASURED:clm.evidence_date.cpu}}

The first request with a head that is not yet on the node also downloads it (75 MB for the reference head): {{MEASURED:clm.head_cold_start_p50_ms}} ms p50, {{MEASURED:clm.head_cold_start_p95_ms}} ms p95, measured on {{MEASURED:clm.evidence_date.head_cold_start}}. A replica that is not loaded answers 503 with Retry-After while the encoder loads (see Errors). These are measurements, not guarantees; measure on your own nodes and inputs.

Training a domain head — hyperparameters​

Model ID: clm-head-infonce

The template trains the two heads over the frozen encoder, in one job on one GPU. The full walkthrough, with the dataset formats, is in Fine-tune a CLM head.

ParameterDefaultRange / optionsNotes
task"clm""clm", "choice"clm: state/action pairs (contrastive). choice: multiple-choice questions in the System One format.
loss"infonce""infonce", "softce"softce requires task choice.
targets"soft""soft", "hard"Task choice: soft keeps the gold distribution, hard its most probable label.
epochs201 to 200Maximum epochs of the head training; early stopping may end sooner.
patience51 to 200Epochs without a better validation score before training stops.
seed12340 to 2147483647Seed of the data split and of the head training.
batchnull8 to 8192null = task default: 2048 for clm, 256 for choice.
lrnullgreater than 0, at most 1.0null = task default: 2e-3 · √(1024/width) · √(batch/1024) for clm, 5e-4 for choice.
max_tokensnull16 to 8192Encoder tokens kept per text. null = task limit: 8192 for clm, 2048 for choice. A value above the task's limit fails before the GPU is used.
batch_tokensnull16 to 8192Padded tokens per encoder pass while the dataset is embedded. null = max_tokens.
max_total_tokens500000001 to 200000000Upper bound on the encoder tokens read from the dataset; a larger dataset fails before encoding starts.
holdout_frac0.10.01 to 0.49Share held out for validation (and, for choice, for test).

Licenses​

  • Encoder weights: Qwen/Qwen3-8B, Apache-2.0 (Alibaba Cloud). Reference head: Contrastive-LM/CLM-v0.1-8B, Apache-2.0. See Model licenses.
  • The serving and training code (clm_runtime) includes modified code from Contrastive-LM/CLM at commit bb42c6c and is Apache-2.0, with its own LICENSE and NOTICE. The ColabHive node runtime, SDK and MCP server are MIT; the orchestrator is under the Business Source License 1.1.
  • A head you train inherits the licenses of the encoder and of the head it started from.

Cost​

CLM has no per-model price. Requests served on your own nodes are not charged; on capacity ColabHive operates they are metered like any other inference, per accelerator-hour — see Pricing.

Quick start​

curl -X POST "https://api.colabhive.com/v1/rank" \
-H "Authorization: Bearer $COLABHIVE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "clm-v0.1-8b",
"context": "Tides are caused mainly by",
"question": "Which answer completes the sentence?",
"answers": ["the gravity of the Moon", "the wind over the ocean"]}'
from colabhive import ColabHive
from colabhive.clm import Choice, Noul, Score

client = ColabHive(api_key="hive_...", account_id="YOUR_ACCOUNT_ID")
r = client.clm.system_one(
model="clm-v0.1-8b",
state="Customer: my invoice was charged twice and nobody answers the phone!",
questions={
"urgency": Noul(instructions="Is this urgent?"),
"department": Choice(instructions="Which team should handle this?",
criteria={"billing": "Charges, invoices, refunds",
"technical": "Bugs and outages"}),
"frustration": Score(instructions="How frustrated is the customer?",
criteria=["Calm", "Frustrated", "Very angry"]),
},
)
print(r.answers["department"].choice, r.answers["urgency"].noul, r.latency_ms)

client.clm is in the Python SDK from colabhive 0.10.0 — see the SDK reference.


Authors: José Luis Minich, Maximiliano Lucius.