Fine-tune a CLM head
CLM-8B scores candidates with a frozen Qwen3-8B encoder and
two small heads. The reference head is general; a head trained on your own labeled data is what makes
the scores useful in your domain. Training changes only the heads — the encoder is never modified — so
a trained head is a file of about 75 MB, and serving it loads nothing new: it rides with each request
to the same clm-v0.1-8b replicas.
Template: clm-head-infonce (framework clm). One job, on one GPU, in two phases:
- Embed. The encoder reads every text in the dataset once and stores the vectors on the node. This phase is most of the time and is resumable: a job re-dispatched to the same node shortly after an interruption continues from the last completed shard of 10,000 rows.
- Train the heads. The encoder is released and the two heads are trained on the stored vectors, with early stopping on a validation split.
The job writes artifacts/clm_head.pt and artifacts/clm_head_manifest.json. The manifest binds the
head to the exact encoder it was trained on (repository, revision, dtype, pooling, tokenizer and chat
template); a head is refused by any other encoder, so it can never score silently on the wrong one.
Choose the task
| Task | Your data | What the head learns | Longest text |
|---|---|---|---|
clm (default) | trajectories: a state and the action taken in it | to score the action that fits a state above the others (contrastive, InfoNCE) | 8,192 tokens |
choice | multiple-choice questions in the /v1/systemone format, with the right answer | to answer noul, choice and score questions like yours | 2,048 tokens |
Use choice when you will call the head with fixed questions (routing, triage, rubric scoring). Use
clm when the candidates are open-ended actions, such as the next step of an agent.
Dataset formats
Upload one JSONL file (one JSON object per line) as a dataset — see
Preparing Datasets and the Datasets API. A dataset
with more than one .jsonl file is refused.
Task clm
{"state": "Ticket: VPN drops every 10 minutes since the update.", "action": "Ask for the client version and the OS", "task_id": "t-1841", "step_idx": 0}
{"state": "Ticket: VPN drops every 10 minutes since the update.\nUser: 5.2.1 on macOS 15", "action": "Link the known-issue article for 5.2.1 and offer the 5.2.2 build", "task_id": "t-1841", "step_idx": 1}
{"state": [{"role": "user", "content": "Refund my last order"}], "action": "Look up the order and check the refund window", "task_id": "t-2207", "step_idx": 0}
| Field | Required | Meaning |
|---|---|---|
state | yes | The context: text, or a list of chat messages (role + content). A head trained on chat messages expects chat messages at serving time. |
action | yes | The action taken in that state: a non-empty string. |
task_id | yes | Groups the steps of one task. The validation split is by task_id, so at least two distinct values are needed. |
step_idx | yes | The step within the task. Rows with the same task_id and step_idx are never used as negatives of each other. |
trajectory_id, reward | no | Kept with the row; not used by the loss. |
Task choice
Each row is one /v1/systemone request plus its answers. gold maps each question id to its right
label, and optionally to a full probabilities distribution (soft targets).
{"id": "q-001", "state": "My card was charged twice for the same order.", "questions": {"intent": {"type": "choice", "instructions": "What does the customer need?", "criteria": {"refund": "Money back for a charge", "card_lost": "Block a lost or stolen card", "transfer": "Send money to someone"}}}, "gold": {"intent": {"label": "refund"}}}
{"id": "q-002", "state": "I can't find my card anywhere since yesterday.", "questions": {"intent": {"type": "choice", "instructions": "What does the customer need?", "criteria": {"refund": "Money back for a charge", "card_lost": "Block a lost or stolen card", "transfer": "Send money to someone"}}}, "gold": {"intent": {"label": "card_lost", "probabilities": {"card_lost": 0.9, "refund": 0.1}}}}
| Field | Required | Meaning |
|---|---|---|
id | yes | Row id. Rows are split into train, validation and test by id, so a row never appears in two splits; at least three distinct ids are needed. |
state | yes | Text, an object or an array (as in /v1/systemone). Chat messages are refused: choice heads read prose. |
questions | yes | {question_id: {type, instructions, criteria}}, exactly as in /v1/systemone. |
gold | yes | {question_id: {label, probabilities?}}. label is an option key for choice, true/false for noul, and the level index ("0", "1", …) for score. A question without a gold label is not trained on. |
state, questions and gold may also be JSON-encoded strings.
Hyperparameters
Model ID: clm-head-infonce
This is the template's strict contract: unknown keys and out-of-range values are refused when the run is created. The same table is on the model page.
| Parameter | Default | Range / options | Notes |
|---|---|---|---|
task | "clm" | "clm", "choice" | clm: state/action pairs (contrastive). choice: multiple-choice questions in the System One format. |
loss | "infonce" | "infonce", "softce" | softce requires task choice. |
targets | "soft" | "soft", "hard" | Task choice: soft keeps the gold distribution, hard its most probable label. |
epochs | 20 | 1 to 200 | Maximum epochs of the head training; early stopping may end sooner. |
patience | 5 | 1 to 200 | Epochs without a better validation score before training stops. |
seed | 1234 | 0 to 2147483647 | Seed of the data split and of the head training. |
batch | null | 8 to 8192 | null = task default: 2048 for clm, 256 for choice. |
lr | null | greater than 0, at most 1.0 | null = task default: 2e-3 · √(1024/width) · √(batch/1024) for clm, 5e-4 for choice. |
max_tokens | null | 16 to 8192 | Encoder tokens kept per text. null = task limit: 8192 for clm, 2048 for choice. A value above the task's limit fails before the GPU is used. |
batch_tokens | null | 16 to 8192 | Padded tokens per encoder pass while the dataset is embedded. null = max_tokens. |
max_total_tokens | 50000000 | 1 to 200000000 | Upper bound on the encoder tokens read from the dataset; a larger dataset fails before encoding starts. |
holdout_frac | 0.1 | 0.01 to 0.49 | Share held out for validation (and, for choice, for test). |
null means "the task's default": leave the key out, or send null, and the job picks the value
for the task you chose.
Run it
from colabhive import ColabHive
client = ColabHive(api_key="hive_...", account_id="YOUR_ACCOUNT_ID")
dataset = client.datasets.upload(name="support-intents", file="./intents.jsonl")
job = client.training.create(
model="clm-head-infonce",
dataset_id=dataset.id,
job_name="support-intents-head",
hyperparameters={"task": "choice", "loss": "softce", "targets": "soft"},
)
job.wait()
print(job.status, job.get_metrics())
Metrics reported while the job runs:
- task
clm:lossduring training;val_loss,val_top1andwithin_task_top1(the selection metric) on the validation tasks; - task
choice:accandsoft_ceon validation, andtest_accandtest_soft_ceon the held-out test ids at the end.
Before the GPU is used, the job validates the whole dataset and the configuration and fails with a
named error if something is wrong: dataset_invalid (with the line number), dataset_too_small,
max_total_tokens_exceeded, choice_max_tokens_exceeded, choice_chat_state_unsupported. A
non-finite loss fails the job instead of producing a head.
Serve the head
Register the finished run as an endpoint, then pass that endpoint as model:
endpoint = client.training.register_for_inference(
run_id=job.id,
name="support-intents-head",
description="CLM head for support ticket intents",
visibility="account",
)
r = client.clm.system_one(
model=endpoint.endpoint_id, # or its name, "support-intents-head"
state="I lost my card on the train this morning.",
questions={"intent": {"type": "choice", "instructions": "What does the customer need?",
"criteria": {"refund": "Money back for a charge",
"card_lost": "Block a lost or stolen card",
"transfer": "Send money to someone"}}},
)
print(r.answers["intent"].choice, r.colabhive.head_sha256)
The request runs on a clm-v0.1-8b replica with your head; colabhive.head_sha256 in the answer
names the head that scored it. The endpoint follows the account rules of any trained model: other
accounts cannot call it unless you publish it. Never pass a runtime key containing :trained: as
model: it is refused, because it is not bound to your account.
Start from a head you already trained
Pass the earlier run (or its model version) as base. The new job starts from that head instead of the
reference head:
from colabhive import ArtifactRef
job = client.training.create(
model="clm-head-infonce",
dataset_id=new_dataset.id,
base=ArtifactRef.job(previous_job.id), # or ArtifactRef.model_version(version_id)
hyperparameters={"task": "choice"},
)
base must be a trained CLM head over the same encoder: anything else is refused with 422
(clm_base_not_a_head, clm_base_encoder_mismatch) before a job exists. Without base, training
starts from the reference head.
Hardware
The job needs one GPU that holds the encoder in bf16 with room for its activations. It never splits the encoder across GPUs, and nothing below bf16 is used.
| Hardware | Head training | Reference run: choice on a fixed Banking77 subset | Encoder throughput while embedding | Measured on |
|---|---|---|---|---|
| NVIDIA RTX 3090 24 GB | {{MEASURED:clm.status.training.nvidia_rtx3090}} | {{MEASURED:clm.training_time.banking77_choice.nvidia_rtx3090}} | {{MEASURED:clm.encoder_tokens_per_s.nvidia_rtx3090}} tokens/s | {{MEASURED:clm.evidence_date.training.nvidia_rtx3090}} |
| Intel Arc Pro B70 32 GB | {{MEASURED:clm.status.training.intel_b70}} | {{MEASURED:clm.training_time.banking77_choice.intel_b70}} | {{MEASURED:clm.encoder_tokens_per_s.intel_b70}} tokens/s | {{MEASURED:clm.evidence_date.training.intel_b70}} |
AMD Instinct {{MEASURED:clm.amd.board}} | {{MEASURED:clm.status.training.amd_instinct}} | {{MEASURED:clm.training_time.banking77_choice.amd_instinct}} | {{MEASURED:clm.encoder_tokens_per_s.amd_instinct}} tokens/s | {{MEASURED:clm.evidence_date.training.amd_instinct}} |
| Intel, two 16 GB GPUs | not offered | — | — | — |
| CPU | {{MEASURED:clm.status.training.cpu}} | {{MEASURED:clm.training_time.banking77_choice.cpu}} | {{MEASURED:clm.encoder_tokens_per_s.cpu}} tokens/s | {{MEASURED:clm.evidence_date.training.cpu}} |
The reference run has {{MEASURED:clm.training.banking77_rows}} rows. The embedding phase scales with
the tokens in your dataset: divide them by the throughput above for a first estimate. The job also
reserves {{MEASURED:clm.training.cpu_cores}} CPU cores and {{MEASURED:clm.training.ram_gb}} GB of
RAM on the node. These are measurements of the reference run, not guarantees.
Training on your own nodes is not charged; on capacity ColabHive operates it is billed per accelerator-hour — see Pricing.
Related
- CLM-8B model page — contract, limits, known limitations
- OpenAI-compatible API: CLM scorer routes
- Training API · Merge & Retrain API
- Register a Trained Model
Authors: José Luis Minich, Maximiliano Lucius.