Private Agentic Infrastructure
ColabHive gives infrastructure and AI teams one maintained path to run inference, fine-tuning and model operations across the heterogeneous CPU and GPU infrastructure they control.
This page explains the whole model in one place: what it is for, what a private cluster is made of, how work gets placed, where capacity comes from when the cluster runs out, and the explicit boundary of each operating path.
1. The problem this solves
Teams building agentic systems keep hitting the same wall. The agent itself is cheap to write; the infrastructure underneath it is not:
- The work is not one shape. An agentic workflow is orchestration, document ingestion, retrieval, ranking, classification, tool calls and generation. Only some of that belongs on a GPU, but the usual answer is to put all of it on one.
- The hardware is not uniform. Real organisations own a mix: some NVIDIA cards, maybe an Intel or AMD box, and a lot of CPU capacity that sits idle. Most platforms assume a homogeneous fleet and ignore everything that doesn't match.
- The data has a boundary. Sensitive documents, financial data and customer records are exactly what agentic systems are most useful on — and exactly what cannot be shipped to a third-party inference API by default.
- Capacity is spiky. Private capacity is sized for the normal case, so the abnormal case either fails, queues forever, or forces a permanent over-purchase.
ColabHive's position is that these are one problem, not four. A maintained control plane should know about every enrolled processor, place each task according to capacity and runtime compatibility, and provide a supported path across heterogeneous hardware. Temporary external capacity is a separate Enterprise add-on using the customer's provider account.
2. What a private cluster is
A Private Cluster is a scoped deployment of customer-controlled nodes running the ColabHive node runtime. It is the default topology:
| Node kind | Typical contents | Runs |
|---|---|---|
| GPU node | NVIDIA (CUDA), Intel Arc (XPU), AMD (ROCm) cards | LLM and generative inference, fine-tuning |
| CPU node | Server or workstation CPUs, large RAM | orchestration, retrieval, classical ML, data services |
| Mixed node | CPUs plus one or more GPUs of any vendor | both lanes on the same machine |
Nodes are not required to be identical, co-located, or connected by special interconnects. There is no NVLink or InfiniBand requirement, because expert models train and serve as independent jobs rather than as one tightly-synchronised cluster-wide computation.
Models, datasets, endpoints, API keys and trained artifacts are scoped to an account. Accounts do not see each other's private resources, and an endpoint published as private is not callable from another account. This API and resource scoping does not, by itself, select eligible nodes or guarantee isolated placement for that account.
3. How nodes are enrolled
Enrollment is a one-time operation per machine. The short version:
- An operator generates an enrollment token for the account.
- The installer runs on the machine, installs the node runtime, and registers it against that token.
- The node detects its own hardware — CPU cores and topology, memory, GPUs and their memory, disk and network — and reports it to the control plane.
- The node opens a persistent connection and begins sending heartbeats. Capabilities are refreshed from those heartbeats, so a node that changes (new card, more RAM) converges without re-enrolling.
- From that point the scheduler treats it like any other node: it is eligible for whatever workloads its processors, memory and vendor backends can actually satisfy.
Container images are always referenced by an explicit version — never a floating latest tag — so a
node's runtime and model images are reproducible and auditable.
Enrollment is not idempotent: running the installer again with a fresh token registers a second node rather than updating the first. To update an already-enrolled node, update its runtime rather than re-enrolling it.
4. Two execution lanes: CPU and GPU
ColabHive treats CPU and GPU as complementary rails of the same agentic system, not as a primary and a fallback.
The CPU agentic lane
Work that is latency-sensitive, concurrency-heavy, memory-bound or simply not matrix-multiplication:
- agent orchestration
- document ingestion and transformation
- parsing, auxiliary OCR and preprocessing
- retrieval and context preparation
- embeddings and reranking, where the model and latency/quality requirements allow it
- classification and small models
- classical machine learning
- high-concurrency services
- compression, encryption and data movement
- hyperparameter tuning and jobs that do not need a dedicated GPU
Model merging also lives here: merging is weight arithmetic, so it runs on the CPU lane in bf16 and never consumes GPU quota.
On Intel systems, the lane has more to work with than the cores alone — AMX for bf16 and int8 matrix work, and the on-die QAT, IAA and DSA engines for cryptography, compression and data movement. On our reference node these engines are present and enumerated by the platform. Consuming them from the runtime is an active engineering track, and we publish measured results rather than projected speedups — so no acceleration figures are claimed here yet.
The GPU generative lane
Work that is bandwidth-bound or too large for a CPU to serve at a useful rate:
- large language models
- generative inference
- multimodal models
- long contexts
- fine-tuning and QLoRA
- models that need high memory bandwidth
- tensor parallel, and workloads that exceed a single GPU
Vendor is a property of the node, not of the workload. The same model definition can be served through CUDA, ROCm or XPU backends; which architectures are supported on which backend is data-driven, not hard-coded. See the Compatibility Matrix.
5. How placement is decided
Placement is a property of the control plane with an enforced eligibility boundary. When work arrives, the control plane first resolves the most restrictive node policy declared at account, workload and task level. It then asks: which eligible available node can run this task with the required runtime and capacity? Account scoping alone does not select a node; the explicit node-eligibility policy does.
The inputs it uses:
- Capacity and fit — does the target node have enough GPU memory (or RAM, for CPU work) for this model, given what is already resident on it?
- Vendor compatibility — is there a backend and image for this model architecture on that node's vendor?
- Warm state — is the model already resident somewhere? Serving from a warm replica avoids a cold start, so warm nodes are strongly preferred.
- Load — current concurrency and queue depth per node and per GPU.
- Availability — nodes that are offline, cordoned, draining or in maintenance are excluded.
Account scoping controls which resources a caller may see or invoke; it is not an account-aware placement control today. Global operator settings separately control whether Burst automation may rent temporary external capacity.
Two consequences worth stating plainly. First, a task is queued rather than failed when nothing fits right now: the admission path answers with an accepted-and-queued response plus an ETA instead of a surprise error. Second, the scheduler does not ask which vendor owns the workload — it asks which nodes can satisfy it, and picks among those.
6. The three capacity paths
One control plane, three places capacity can come from, each with a different owner and contract.
Path 1 — Private Cluster · default topology
Customer-controlled Intel, NVIDIA, AMD and CPU nodes in a scoped deployment. Account-scoped API resources do not imply an end-to-end placement-isolation guarantee.
Path 2 — Share Hive · account opt-in, unselected by default
No capacity is shared unless the account owner enables it. This choice does not imply universal isolation, residency or contributor-earnings guarantees.
Path 3 — Private Cloud Burst · Enterprise add-on
The customer connects its DigitalOcean account and pays the provider directly. ColabHive can provision, enroll, drain and release eligible temporary nodes, and adds no fee on the provider charges. See Private Cloud Burst.
Three sources of capacity
It is worth separating these, because they are often conflated:
| Path | What it is | Status |
|---|---|---|
| DigitalOcean | Provider for the customer-owned Cloud Burst path | Enterprise add-on |
| Share Hive | Account-level opt-in, unselected by default; no capacity is shared until enabled by the account owner | Account-controlled |
| ColabHive-operated nodes | Capacity used by hosted Starter, Pro and Team plans | Hosted |
7. Boundaries and controls
API account scoping and workload placement are separate concerns. The current controls are:
| Control | What it means | Status |
|---|---|---|
| Private Cluster topology | Customer-controlled nodes in a scoped deployment; this does not turn API account scoping into a universal placement-isolation guarantee | Default path |
| Account-scoped resources | Accounts scope datasets, trained models, endpoints, training runs, inference tasks and keys; other accounts cannot see or call those private resources. Public trained models remain globally readable | Product control; complete trained-model boundary requires Builder 1.1.7+ / Model Registry 0.1.1+ |
| Account/workload/task node eligibility | own_hardware_only, own_account_only or any_node; the most restrictive declaration wins and no-capacity fails closed | Product control |
| Capacity- and runtime-aware placement | Placement uses fit, vendor compatibility, warm state, load and node availability inside the deployment boundary | Product control |
| Auditable model lineage | Base model, datasets, jobs, adapters and merge method are recorded per version and queryable as a chain | Product control |
| Burst controls | Node limit, configured spend fields, rental pacing, spend tracking and kill switch; no universal capacity guarantee | Enterprise add-on |
| Drain and release lifecycle | Temporary capacity is cordoned, drained make-before-break, then released | Enterprise add-on |
| Residency and provider requirements | Defined in the applicable deployment agreement; no universal guarantee is published | Contract boundary |
| Workload-level permissions | Most Builder routes still treat API-key scopes as account-wide authority. Cohort run/cancel/health is the documented exception and enforces inference:execute for restrictive keys | API boundary |
We deliberately avoid absolute security language. ColabHive has not been through a formal zero-trust implementation or a third-party compliance audit, and claims neither. If your requirements need per-workload residency or granular scope enforcement, put those requirements in the deployment agreement before building against them.
8. What happens during a burst
Cloud Burst is the path that introduces temporary external capacity, so it is worth reading end to end.
- Deficit is observed. A capacity gap has to persist for several consecutive minutes. A momentary spike never rents anything.
- Local plays are exhausted first. The planner tries rebalancing models across GPUs, evicting low-value replicas, and making room by preemption. Only demand that physically cannot fit counts.
- Corroboration. Independent signals — memory pressure or an ageing backlog — must confirm the deficit before money is spent.
- One node is rented. A cloud GPU is provisioned, the node runtime is installed, and the node enrolls into the hive. In the first automated production cycle this took about three minutes.
- It is scheduled like any other node, according to capacity, runtime compatibility, warm state, load and availability.
- Reabsorption is attempted continuously. As soon as the workload fits back onto owned hardware — or the node goes idle for a sustained window — a make-before-break drain recreates the needed replicas on physical GPUs first, lets in-flight requests finish, and only then destroys the cloud node.
- Training jobs are never interrupted. A burst node running a training job is not destroyed mid-run; the drain waits for the job, however long it takes.
Throughout, global operator controls and a manual kill switch govern Burst activity. The configured node, hourly and monthly fields can stop new rentals at their limits. No enabled daily cap or default contractual daily ceiling is claimed; when Burst is disabled or a configured control prevents a rental, excess demand queues with an ETA.
The first end-to-end automated cycle rented a DigitalOcean NVIDIA H100, served 224 production inference requests, then drained and destroyed the node cleanly at a total cost of $4.82.
9. How agents, models and tools interact
ColabHive is not a model catalogue with an API in front of it. Through the
MCP server, an agent discovers the public capabilities and private resources
visible to its account and calls them as ordinary tools. Most API-key scopes remain descriptive;
restrictive Cohort run/cancel/health calls are the enforced inference:execute exception:
- expert models — your trained and merged specialists
- embeddings and rerankers — retrieval building blocks
- forecasting — time-series and volatility models
- classification — text and tabular classifiers
- document services — OCR, transcription, translation, moderation
- model operations —
training.mergeandtraining.retrain_on - account-scoped runtime resources — private dataset, trained-model, endpoint, task and training visibility follows the calling account; placement eligibility is a separate boundary and public trained models remain globally readable
Merge and retrain are live MCP-accessible operations and preserve lineage. A resulting version may
carry the candidate lifecycle label, but that label must not be presented as proof of a mandatory
human approval gate: mandatory automated evaluation gates and an enforced promote/rollback workflow
are outside the public contract. Placement is capacity- and runtime-aware within the deployment boundary.
How models improve over time — train, merge, evaluate, promote, serve, retrain on top, with full lineage and rollback — is covered in The Model Flywheel. The division of labour is worth stating once:
The infrastructure decides where work runs. The Model Flywheel governs how each expert improves over time.
10. Product capabilities and explicit boundaries
Product capabilities
- Heterogeneous scheduling across NVIDIA (CUDA), Intel (XPU) and CPU nodes under one control plane
- Node enrollment, hardware detection, capacity reporting and warm/cold-aware placement
- Private Cluster as the default topology plus enforced account/workload/task node eligibility; neither control is a universal country, provider, jurisdiction or physical-isolation guarantee
- A catalog spanning LLMs, specialists, generative, forecasting and tabular ML, plus Hugging Face import
- Merge and retrain-on-top as platform operations, with per-version lineage
- MCP server exposing catalog tools plus merge and retrain operations
- REST API, OpenAI-compatible surface and Python SDK
Enterprise add-on
- Private Cloud Burst — temporary capacity in the customer's DigitalOcean account, with global controls and a manual kill switch; paid to the provider, with no ColabHive fee
Validated in proof of concept
- Xeon + 6× Intel Arc Pro B70 as a complete heterogeneous node — multi-GPU LLM inference (including a 62 GB mixture-of-experts model served tensor-parallel across four cards), embeddings, reranking, multimodal and coding model families, small financial workloads inside agentic pipelines, and co-scheduling with other private nodes
Outside the public contract
- Automated evaluation gates: regression detection, forgetting checks, promote/rollback
- Consuming Intel AMX and the on-die QAT / IAA / DSA engines from the CPU lane
- Per-workload residency, provider restrictions and enforced per-key scopes
- Universal residency/provider enforcement and granular per-key scope enforcement
Not part of the offer
- Additional cloud providers for burst capacity
- An economic layer for capacity contributors and model publishers
- Agent-initiated retraining with explicit, enforced evaluation and promotion controls
Where to go next
- Platform Overview — the components and how a request flows through them
- Elastic Cloud Burst — the burst tier in detail
- The Model Flywheel — how experts improve over time
- Inference Lifecycle — cold start, warm models, latency
- MCP — connect an agent to the hive
Authors: J.L. Minich, M. Lucius