Skip to main content

Self-host ColabHive Community

ColabHive Community v2.3 is the self-hosted control plane: it enrolls your nodes and handles placement, routing, model lifecycle and observability across your CPU/GPU hardware. You operate the deployment, its security, backups, upgrades and availability.

This page takes you from the download to a first inference on your own node: requirements, download, start, one hostname in front, your account, the starter catalog, the node install and a test request.

License​

  • The control plane (packages/orchestrator, ColabHive Orchestrator 1.0.1) is licensed under the Business Source License 1.1 (licensor: Aureus Technologies LLC). The license text is packages/orchestrator/LICENSE inside the bundle.
  • No license fee for labs, researchers and developers — the license's Additional Use Grant:
    • developing, evaluating, testing, benchmarking, demonstrating, learning about or researching software, models or hardware, including in continuous integration;
    • personal, non-commercial use by an individual on hardware they own or control;
    • research and teaching by academic institutions and non-profit research organizations.
  • Other production use needs a Pro, Team or Enterprise license. Describe your deployment on colabhive.com/get-started.
  • Each version changes to the MIT License four years after it is published.
  • Everything else in the bundle — node runtime, SDK and the other services — is MIT. NOTICE maps each package to its license, and licenses/THIRD_PARTY_LICENSES.md lists the dependencies the build installs and the service images Compose pulls, each under its own license.

Requirements​

Control plane host​

The bundle states no minimum CPU, RAM or disk size beyond what the table says. This is what it runs and uses:

AreaWhat the bundle uses
Operating system and toolsLinux with Docker Engine and Docker Compose v2. Image builds need BuildKit (docker buildx): the Builder image depends on its Dockerfile.dockerignore, which only BuildKit applies. Python 3 and OpenSSL, which scripts/setup-env.sh uses to generate local credentials.
Containersdocker compose up -d --build starts 21 services. It builds nine application images from the bundle's source and pulls the rest: CockroachDB, MinIO, Redis, nginx, Prometheus, Grafana, Alertmanager and two exporters. Services behind the memory*, webhooks and retired-scaffolds profiles stay off.
MemoryCockroachDB is the only container with a memory limit: 4 GiB (mem_limit: 4g, started with --cache=1GiB --max-sql-memory=1.5GiB). The other services have no limit, so size the host for them on top of that.
DiskDocker volumes for CockroachDB, MinIO, Redis, Prometheus (30 days of retention), Grafana and Alertmanager, plus the build cache of the nine images.
NetworkEvery published port binds to 127.0.0.1: the API gateway on 8080, the orchestrator on 8001, the node WebSocket gateway on 8013, the Builder gateway on 8014 and the training scheduler on 8003. No other machine reaches the control plane until you put a reverse proxy in front of it (one hostname in front).

Node hosts​

scripts/install-node.sh prepares a node and enrolls it. On every node it needs or installs:

  • Linux with apt-get, dnf or yum. The vendor steps below are automated for Ubuntu and Debian, and for Fedora and RHEL-family distributions where the table says so.
  • sudo and systemd, for the colabhive-node service. Run the installer as root or as a user who can use sudo; the service runs as that user.
  • Python 3.10 or newer with pip and venv. The installer stops if python3 is missing, and the node runtime package requires 3.10. On apt it installs python3-pip and python3-venv when pip is missing.
  • curl and pciutils, installed when missing. Without lspci every GPU check is skipped and the node enrolls as CPU-only.
  • Docker Engine, installed from the distribution when missing (docker.io on apt, docker on dnf and yum), enabled at boot, with the installing user added to the docker group.

What it detects and sets up for each accelerator vendor:

VendorDetected whenKernel driverRuntime and container accessWhat the installer checks
NVIDIAlspci lists an NVIDIA deviceWhen nvidia-smi is missing it installs one: nvidia-driver-550, then 545, then 535, then ubuntu-drivers install on Ubuntu; nvidia-driver on Debian; akmod-nvidia on Fedora and the RHEL family. Other distributions: install it yourself. Reboot and run the installer again after a driver install.NVIDIA Container Toolkit from NVIDIA's repository (apt and dnf/yum), configured for Docker with nvidia-ctk runtime configure --runtime=docker.nvidia-smi runs inside nvidia/cuda:12.1.0-base-ubuntu22.04 with --gpus all.
Intel (discrete Arc: Alchemist, Battlemage)a display device from vendor 8086 bound to the xe driver, or named Arc, DG2, BMG or Battlemage, or with a Battlemage PCI ID (8086:e2xx). Integrated UHD, Iris and HD Graphics are not treated as accelerators.The in-kernel xe/i915 driver. The installer installs no kernel driver.Level Zero and OpenCL runtime (intel-opencl-icd, libze1, libze-intel-gpu1, clinfo, and xpu-smi when available) from Intel's graphics repository on Ubuntu, distribution packages on Debian, intel-opencl and intel-level-zero on Fedora and the RHEL family. The installing user joins render and video; containers get /dev/dri. PyTorch XPU goes into the node's virtualenv (about 1–2 GB) to read VRAM.clinfo lists an Intel device; a container sees /dev/dri/renderD*; torch.xpu.is_available() is true.
AMD (Radeon, Instinct)a vendor 1002 device in the display or accelerator class; once amdgpu is loaded, only a dedicated compute GPU in the KFD topology counts, not an APU's integrated GPUamdgpu must already be loaded and bound to the card. The installer never installs a kernel driver; it points you to amdgpu-dkms from repo.radeon.com.Only rocm-smi and rocminfo, from AMD's ROCm repository on Ubuntu; never the full ROCm stack. Other distributions: install rocm-smi yourself. The installing user joins render and video; containers get /dev/kfd and /dev/dri.rocm-smi --showproductname lists a GPU; /dev/kfd and /dev/dri/renderD* exist; a container sees both.
CPU onlyno GPU found——The node enrolls with 0 GPUs.

The installer checks no minimum driver, kernel or ROCm version: it checks that each stack can see the GPU, and prints what to fix when it cannot. On a node with a GPU it holds the running kernel and the installed kernel metapackages (apt-mark hold) and excludes them from unattended upgrades, because a kernel update can stop GPU compute while the node still lists its GPUs. On dnf and yum it only warns you to use dnf versionlock. COLABHIVE_PIN_KERNEL=0 skips this; CPU-only nodes are never pinned.

A node opens these connections itself:

  • your control plane, for enrollment, heartbeats and tasks, and its task WebSocket;
  • console.colabhive.com/install/artifacts, where the installer downloads the node runtime package (MIT) and its checksum by default. --base-url points it at your own copy with the same layout (<base-url>/latest/manifest.json, the wheel and its .sha256);
  • the package repositories above, download.pytorch.org on Intel nodes, Docker Hub for the check containers, and the registries of the inference images your control plane dispatches: with the starter catalog, registry.colabhive.com, pulled by version tag and checked against its SHA-256 digest;
  • huggingface.co, for the model weights, at the revision the catalog pins.

Download and verify​

The bundle is generated from a strict source allowlist. It contains no .env, repository history, runtime data, backups, planning material, caches or secret paths. Its build gate scans for high-confidence secret patterns, verifies every file against the embedded SHA-256 manifest, and runs docker compose config --quiet after extraction.

curl -fLO https://colabhive.com/downloads/community/v2.3/colabhive-community-v2.3.tar.gz
curl -fLO https://colabhive.com/downloads/community/v2.3/colabhive-community-v2.3.tar.gz.sha256
sha256sum -c colabhive-community-v2.3.tar.gz.sha256
tar -xzf colabhive-community-v2.3.tar.gz
cd colabhive-community-v2.3
sha256sum -c MANIFEST.sha256

The published archive's SHA-256 is 42a15d6c3ccdc2de6472473675c1bf57d640ad60d135700266c0d0e14b504035. Treat a checksum failure as a hard stop: do not run an archive that does not match both checks.

Start the control plane​

For the uses the Additional Use Grant covers — developing, evaluating, testing, benchmarking, demonstrating, learning about or researching software, models or hardware (including CI); personal, non-commercial use on hardware you own or control; and research and teaching at academic and non-profit research organizations — on a trusted network:

./scripts/setup-env.sh development
docker compose config --quiet
docker compose up -d cockroachdb
./infra/cockroachdb/apply-all-schemas.sh
docker compose up -d --build

Development mode creates local-only credentials. Do not expose that configuration to the public internet. Check startup with docker compose ps and ./scripts/health-check.sh.

Serve the console and the API from one hostname​

Nodes on other machines, and the browser that runs the console, reach the control plane through one hostname of yours (cp.example.org on this page; on a trusted network it can be an internal name) that a reverse proxy on the control-plane host routes to the loopback ports. Serve the console and the API from that same hostname: the API gateway adds no CORS headers and the sign-in and orchestrator services send none, so a console on another origin cannot sign in or call the API.

The in-Compose API gateway (8080) carries enrollment, heartbeats, task claims and the console's API calls. The proxy sends the node's task WebSocket, its response streams and its lease renewals to the WebSocket gateway and the orchestrator directly, and the OpenAI-compatible API (/v1/) to the Builder gateway. It refuses the two orchestrator routes that only the Builder gateway calls (/api/v1/inference/custom and /api/v1/inference/batch), as the hosted service does. With nginx:

server {
listen 443 ssl;
server_name cp.example.org;
ssl_certificate /etc/ssl/certs/cp.example.org.pem;
ssl_certificate_key /etc/ssl/private/cp.example.org.key;
# The gateway's own limit. nginx defaults to 1 MB, which cuts uploads and long
# response streams from nodes.
client_max_body_size 10G;

# Node task channel: WebSocket gateway.
location /ws/tasks {
proxy_pass http://127.0.0.1:8013;
proxy_http_version 1.1;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection "upgrade";
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_read_timeout 7d;
proxy_send_timeout 7d;
}

# Response streams from nodes, and the rest of /api/v1/inference/ (an account's
# Hugging Face credential, cluster state): orchestrator, unbuffered both ways or the
# deltas arrive together at the end, marked as outside traffic the way the gateway
# marks what it forwards. ^~ keeps the catch-all regex below from taking them.
location ^~ /api/v1/inference/ {
proxy_pass http://127.0.0.1:8001;
proxy_http_version 1.1;
proxy_request_buffering off;
proxy_buffering off;
proxy_read_timeout 300s;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Edge-Public "1";
}

# Called only by the Builder gateway, inside Compose: never from outside.
location = /api/v1/inference/custom { return 404; }
location ^~ /api/v1/inference/batch { return 404; }

# Lease renewals from nodes: orchestrator.
location ~ ^/api/v1/tasks/[^/]+/renew_lease$ {
proxy_pass http://127.0.0.1:8001;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Edge-Public "1";
}

# OpenAI-compatible API: Builder gateway, unbuffered for streamed answers. The
# gateway answers a request that does not stream within 300 s.
location /v1/ {
proxy_pass http://127.0.0.1:8014;
proxy_http_version 1.1;
proxy_buffering off;
proxy_read_timeout 300s;
proxy_send_timeout 300s;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
proxy_set_header X-Edge-Public "1";
}

# Every other API route, including enrollment and heartbeats: API gateway.
location ~ ^/(api|inference|training|storage)/ {
proxy_pass http://127.0.0.1:8080;
proxy_http_version 1.1;
proxy_read_timeout 300s;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
}

# Console.
location / {
proxy_pass http://127.0.0.1:3000;
proxy_set_header Host $host;
proxy_set_header X-Forwarded-Proto $scheme;
}
}

Create your account​

An account is created the first time a user signs in to the console, and sign-in is with Google. Compose does not start the console: it is packages/frontend, a Next.js app you build once.

  1. In Google Cloud, create an OAuth client of type Web application with the redirect URI https://cp.example.org/auth/callback.

  2. In .env, set GOOGLE_CLIENT_ID and GOOGLE_CLIENT_SECRET from that client, FRONTEND_URL=https://cp.example.org and GOOGLE_REDIRECT_URI=https://cp.example.org/auth/callback. If you generated .env in production mode (setup-env.sh production), also set SESSION_COOKIE_DOMAIN=cp.example.org: in production mode the session cookie otherwise carries the hosted service's domain, .colabhive.com, and the browser drops it, so sign-in never completes. In development mode the cookie is bound to the hostname that served it and nothing needs setting. Apply them with docker compose up -d user-management.

  3. Build and start the console with your hostname and client ID compiled in. Without NEXT_PUBLIC_API_BASE_URL, the image calls https://api.colabhive.com instead of your control plane.

    docker build -f packages/frontend/Dockerfile \
    --build-arg NEXT_PUBLIC_API_BASE_URL=https://cp.example.org \
    --build-arg NEXT_PUBLIC_GOOGLE_CLIENT_ID=<your-client-id> \
    -t colabhive-console .
    docker run -d --name colabhive-console --restart unless-stopped \
    -p 127.0.0.1:3000:3000 colabhive-console
  4. Open https://cp.example.org and sign in. Your first sign-in creates your user and a personal account that you own.

Enable your account​

New accounts start on starter, the plan for ColabHive-operated capacity, which allows no own nodes: asking for an enrollment token answers Active BYO node limit reached for the account plan. On your own control plane, move your account to internal with the bundled script, using the address you signed in with:

./scripts/community/bootstrap-operator.sh --email you@example.org

internal means an account of whoever operates this control plane: no seat, API-key, node or request limits, and never charged by its own metering. The script changes one thing, that account's plan_tier, and only after it finds exactly one account owned by that address. It prints the account before and after, and running it again changes nothing. Keep the account ID it prints for the next step. --dry-run shows what it would change.

Load the starter catalog​

The schema creates the catalog empty: no image, runtime or model to dispatch. Load the starter catalog, with the account ID the bootstrap printed:

./scripts/community/seed-catalog.sh --account-id <your-account-id>

It adds:

  • the versioned inference images, each pinned to its registry.colabhive.com tag and SHA-256 digest;
  • the runtimes for each accelerator: vLLM, Transformers and the specialist server on NVIDIA, Intel (XPU) and AMD (ROCm), and Transformers and llama.cpp on CPU;
  • three open models with permissive licenses, as public entries of your catalog, validated on NVIDIA, Intel and AMD:
ModelHugging Face repository (license)What it doesWhere it runs
specialist-moderateunitary/toxic-bert (Apache-2.0)Text moderationA CPU-only node, or a GPU
specialist-embeddingssentence-transformers/all-MiniLM-L6-v2 (Apache-2.0)384-dimension text embeddingsA GPU, 512 MB of VRAM
phi-3.5-minimicrosoft/phi-3.5-mini-instruct (MIT)Chat with tool calling (vLLM, hermes parser)A GPU with about 7.3 GB of VRAM free

--account-id also creates, in your account, one inference endpoint per model, on demand: none keeps a replica loaded. Chat through /v1/chat/completions works without one; the specialists are called through theirs, and /v1/models lists the chat model only once it has one. The script prints the row counts before and after, the vendors each model can run on, and the endpoint IDs.

It only adds rows, in one transaction: a row that already exists is never changed, so running it again changes nothing. --dry-run shows the counts and changes nothing. The file it loads, scripts/community/seed-catalog.sql, lists every row.

Install a node​

Always pass --api-url

install-node.sh enrolls against https://api.colabhive.com unless you pass --api-url. The command the console shows next to a new token downloads the installer from console.colabhive.com and has no --api-url either. Use only the token from it, with the bundle's installer and your own URL.

Enrollment is once per machine: running the installer again with a new token registers a second node.

1. Get an enrollment token. On the control-plane host, from the bundle directory, with the account ID the bootstrap printed:

OPS_KEY="$(grep -E '^ORCHESTRATOR_OPS_KEY=' .env | cut -d= -f2-)"
curl -fsS -X POST http://127.0.0.1:8080/api/v1/nodes/enrollment-tokens \
-H "X-API-Key: $OPS_KEY" -H 'Content-Type: application/json' \
-d '{"account_id": "<your-account-id>", "note": "first node"}'

The token in the response (chive_enroll_…) works once and expires after 15 minutes. The console issues the same token at /nodes/register. A 409 that says Active BYO node limit reached means the account is still on starter: run the bootstrap above.

2. Run the installer from the bundle on the node. Copy scripts/install-node.sh from the bundle directory to the node first (for example with scp), then:

sudo bash install-node.sh --token chive_enroll_... --api-url https://cp.example.org

Run it as root, as above, or as a user who can use sudo: it asks for the password once, and the colabhive-node service then runs as that user. It installs what the node requirements list, enrolls the node against https://cp.example.org/api/v1/agent/enroll, and starts the service with that URL.

The node opens its task WebSocket to GATEWAY_URL. With --api-url the installer sets it from that URL, https becoming wss and http becoming ws on the same host and port (wss://cp.example.org here), in the service unit. To use another address, put GATEWAY_URL= in /etc/colabhive/node.env before you run the installer: it keeps that file, and the value there wins. The file also holds the passphrase that encrypts the node's key, which the installer creates on the first run and reuses after that.

3. Check that it enrolled. On the node:

sudo cat /root/.colabhive/node.json # node_id, account_id and "api_url": your URL
# (as another user: ~/.colabhive/node.json)
systemctl cat colabhive-node | grep -E 'ORCHESTRATOR_URL|GATEWAY_URL'
sudo journalctl -u colabhive-node -f # "Attempting WebSocket connection to wss://cp.example.org",
# then "WebSocket connected"

On the control plane, the node belongs to your account and its heartbeat, sent every 15 seconds, keeps last_heartbeat current:

docker compose exec cockroachdb ./cockroach sql --insecure --database=colabhive --execute \
"SELECT node_id, status, node_runtime_version, last_heartbeat FROM nodes
WHERE account_id = '<your-account-id>' ORDER BY joined_at DESC"

Serve your first model​

A model loads on a node the first time it is asked for, and stays loaded while it is used. Three things come first.

1. A Hugging Face token for your account. Your node downloads the model weights from Hugging Face with a token of the account that owns it. The starter models are not gated: a read token of any Hugging Face account is enough. On the control-plane host:

OPS_KEY="$(grep -E '^ORCHESTRATOR_OPS_KEY=' .env | cut -d= -f2-)"
curl -fsS -X PUT http://127.0.0.1:8001/api/v1/inference/credentials/huggingface \
-H "X-API-Key: $OPS_KEY" -H "X-Account-ID: <your-account-id>" \
-H 'Content-Type: application/json' -d '{"token": "hf_..."}'

The orchestrator checks the token with Hugging Face and stores it encrypted; the answer is its status ("configured": true), never the token. Without one the model never loads on your node: the orchestrator logs HF warmup blocked ... hf_node_owner_token_required and the request waits until it times out.

2. An API key. In the console, open https://cp.example.org/builder/settings/api-keys and create one. It starts with hive_live_ and is shown once.

3. A test request. On a node with an NVIDIA, Intel or AMD GPU, chat with phi-3.5-mini:

export COLABHIVE_API_KEY=hive_live_...
curl -N https://cp.example.org/v1/chat/completions \
-H "Authorization: Bearer $COLABHIVE_API_KEY" -H 'Content-Type: application/json' \
-d '{"model": "phi-3.5-mini", "stream": true,
"messages": [{"role": "user", "content": "Say hello in five words."}]}'

The first request starts the model on your node: the node pulls the vLLM image, several GB, and downloads the weights. With "stream": true the connection stays open while it loads, for up to 20 minutes, and the answer streams in once the model is up. A request without stream waits at most 300 seconds and answers 504 if the model is still loading; send it again once it is up. After that, answers start within seconds. Any OpenAI client works the same way, with base URL https://cp.example.org/v1, your key, and phi-3.5-mini as the model; tools and tool_choice are passed through.

On a CPU-only node, call specialist-moderate through the endpoint the catalog step created (its ID is in that step's output, or in the console under Endpoints):

curl -fsS https://cp.example.org/api/builder/v1/endpoints/<endpoint-id>/infer \
-H "Authorization: Bearer $COLABHIVE_API_KEY" -H 'Content-Type: application/json' \
-d '{"input": {"text": "You are wonderful."}, "sync": true, "sync_timeout_s": 300}'

While the model is not loaded yet, the answer carries a task_id and a status that is not yet succeeded, instead of the result. Ask for the task until it is succeeded; its result holds is_toxic, toxicity_score and one score per label:

curl -fsS https://cp.example.org/api/builder/v1/tasks/<task-id> \
-H "Authorization: Bearer $COLABHIVE_API_KEY"

specialist-embeddings takes {"input": {"texts": ["first text", "second text"]}} on a GPU node the same way, through its own endpoint.

Configuration for a hardened deployment​

Generate the production template, replace every placeholder through your secrets manager, enable TLS, restrict CORS and network ingress, set SESSION_COOKIE_DOMAIN to your hostname (why), then validate Compose before starting:

./scripts/setup-env.sh production
./scripts/generate-secrets.sh
# Move generated values to your secrets manager and replace every <PLACEHOLDER> in .env.
if grep -nE '<[^>]+>' .env; then echo 'unresolved placeholders' >&2; exit 1; fi
docker compose config --quiet
docker compose up -d --build

Running it in production needs the commercial license above. Before accepting workloads, configure backups, restore drills, monitoring receivers, storage retention and an upgrade window.

Reproduce the release gate​

The bundle and its checksum come from the repository with:

python3 scripts/generate_community_bundle.py

The generator uses fixed archive metadata and sorted inputs, so unchanged source produces the same archive bytes. It fails closed on missing required inputs, symlinks, forbidden paths, sensitive patterns, manifest drift, unsafe tar members or an invalid extracted Compose model.

For help, email support@colabhive.com. Report security issues privately to security@colabhive.com; never attach credentials or an .env file.