Skip to main content

Phi-3.5 Mini

Compact, fast, edge-friendly small language model.

Overview​

  • model_name: phi-3.5-mini (curated base)
  • Source: microsoft/phi-3.5-mini-instruct
  • Scale: ~3.8B parameters
  • Served on: GPU via vLLM
  • Public endpoint: phi-3.5-mini-public (task type chat, billed per request in USD)

When to use​

✅ Low-latency assistants, high-volume / cost-sensitive workloads, simple Q&A, extraction, and classification.

❌ For deeper reasoning, code, or multilingual work, step up to a 7B+ model (Mistral 7B, Qwen 2.5, Llama 3.1).

Live specs​

Context window, VRAM footprint, price, and readiness are served from the live catalog — this page does not hardcode them:

curl "https://api.colabhive.com/api/builder/v1/endpoints?visibility=public&search=phi-3.5"

Quick start​

from colabhive import ColabHive

client = ColabHive(api_key="hive_...", account_id="YOUR_ACCOUNT_ID")

result = client.endpoints.infer(
endpoint_id="phi-3.5-mini-public", # SDK resolves the name to a UUID
input_data={"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "What is the capital of France?"},
]},
max_tokens=100,
temperature=0.5,
)

# Sync by default; a "queued" status means poll GET /tasks/{task_id}
print(result["result"] if result.get("status") != "queued"
else client.endpoints.get_task(result["task_id"]))

Tips​

  • Keep prompts direct; low temperature (0.3–0.7) for factual answers.
  • Best for short outputs; use a larger model for long-form generation.

Next steps​


Authors: José Luis Minich, Maximiliano Lucius.