Running agents in parallel
Batch work — a migration applied file by file, tests written module by module — is where parallel agent sessions pay off. This page is what one replica does with several sessions at once, measured.
One replica serves a batch at once
A replica runs concurrent requests in the same batch, so up to its batch size a parallel session does not wait behind the others. Past that size, requests wait for a free slot.
Measured on 2026-09-21 against one gpt-oss-20b replica, with identical short coding prompts and
256 output tokens each:
| Parallel requests | Wall time | Median latency | Slowest | Total throughput |
|---|---|---|---|---|
| 1 | 2.7 s | 2.7 s | 2.7 s | 93 tokens/s |
| 4 | 3.7 s | 3.7 s | 3.7 s | 275 tokens/s |
| 8 | 4.4 s | 4.0 s | 4.4 s | 461 tokens/s |
| 12 | 7.3 s | 5.3 s | 7.2 s | 423 tokens/s |
| 16 | 8.6 s | 5.7 s | 8.6 s | 477 tokens/s |
Up to 8, adding sessions multiplies throughput and barely moves latency. Past 8 the answers split in two groups: at 12, eight came back in about 5 seconds and four in about 7; at 16, eight in about 4 seconds and eight in about 8. The replica admitted 8 concurrent requests and the rest waited for a slot, so throughput stopped growing.
The batch size is per replica and per model, and it is not published by the API. Treat these numbers as the shape of the curve, not as a guarantee for another model, another prompt length or another board.
Sizing a batch of tasks
- Start around 8 parallel sessions per model and measure. More sessions do not finish sooner; they queue.
- Mind the request limit of your key. Sessions sharing one key share its limit per minute — 60 on Starter — and several agents each sending a request every few seconds pass it quickly. See Known limits.
- Keep the shared part of the prompt identical across sessions. Sessions that start with the same system prompt and instructions reuse the replica's prefix cache — see Known limits.
- Start on a warm model. A cold one makes every session wait for the same load — see Known limits.
More than one replica
max_replicas on an endpoint caps how many replicas it can run; whether another one starts depends
on free capacity at that moment. On an endpoint in your own account you set it with
PATCH /endpoints/{endpoint_id}/scaling (operator role), together with min_replicas to keep replicas
loaded. Shared catalog endpoints are scaled by the platform — see
Known limits.