Skip to main content

Cohorts and Latent Fabric

An inference cohort is an opt-in execution made of a durable root task and a small, bounded graph of member tasks. Members can exchange an evaluator-approved intermediate representation inside the private data plane. Public clients see only lifecycle events, the terminal answer, aggregate timing/cost, member count and an explicit fallback reason.

The production lane is intentionally limited to causal/autoregressive text- generation LLMs. Cohort is not enabled for encoders such as BERT, embedding or reranking models, the CLM-8B scorer and its trained heads, diffusion models, or other generative media runtimes: a CLM endpoint, or the endpoint of a trained CLM head, is refused at admission.

The first product lane is exactly one producer and one consumer. A SemanticBridge exports a bounded hidden-state artifact, task-scoped memory stages it, and an authenticated tensor transport delivers it to the consumer bridge. Every step is fenced by tenant, cohort, revision and physical attempt lease. Memory is released and zeroized after acknowledgement or recovery.

Member dispatch is also fenced at the node boundary. A Redis publish is not treated as delivery: the gateway acknowledges only after the target WebSocket accepts the notification. If the node is disconnected or the acknowledgement is absent, the Orchestrator releases its exact lease with compare-and-swap and requeues the member. Reconnect therefore resumes the same durable Cohort without waiting for lease expiry or silently losing the task.

The safe route is selected from observed capability. A direct device path is not assumed: when a topology lacks specific evidence, the provider uses host staging and TLS 1.3 mutual authentication. Hardware brands and runtime names are methodological evidence, not part of the public contract or a commercial support promise.

Classic remains authoritative when execution is omitted. fallback=single allows the same root to return to Classic after a bounded Cohort failure; other fallback policies fail explicitly if unavailable.

Research evaluations are not hidden product defaults. The sealed R4-CEM v6 study uses textual sender-to-reader messages over the Classic API path: 600 items, 11 conditions and two reader backbones. It does not attribute its measurements to the production hidden-state Cohort path.

R4-M is separate from R4-CEM and from this product memory fabric. It evaluates a query-blind textual-memory policy on 78 LongMemEval items — 72 answerable knowledge-update questions and 6 abstention (false-premise) questions — with 3,120 scheduled item-runs and 624 PM/PM-shuffled reader runs. A validity audit (2026-09-07) showed the aggregate over those 78 mixes the two kinds: the policy scored 0/72 on answerable updates and 6/6 on abstention. R4-M is reported as a diagnostic result and its preprint was withdrawn on 2026-09-12. The evaluation harness constructs that memory explicitly; neither execution.mode=cohort nor the Python SDK enables it. The studies were not crossed and therefore do not identify a memory-by-multi-agent interaction. Status and scope are dated on the public research page.