Scale-out architecture

One endpoint.
Up to sixteen brains.

AgentBox nodes link over a low-latency mesh and present to your apps as a single OpenAI-compatible endpoint. Add boxes to run more agents in parallel, push more tokens per second, hold larger models in aggregate memory — without changing a line of integration code.

Cluster configurator

Aggregate capacity, based on Pro nodes (32GB · 6 TOPS each)

2 parallel nodes

Aggregate memory
64 GB
Aggregate NPU
12 TOPS
Parallel agents
8
Throughput (rel.)

Run two 7B models side by side, or one model with double the concurrent request capacity.

Reserve from
$3,196 $1,598
50% off launch
Reserve cluster →

Throughput shown is relative and illustrative; real numbers depend on model, quantization and context length.

How clustering works

Three ways to spend more nodes

More parallel agents

Pin different models or agent roles to different nodes — a router, a coder, a reranker, a vision sidecar — all answering at once behind one gateway.

Higher throughput

Load-balance one model across nodes for near-linear gains in requests per second — ideal for batch RAG, embeddings and high-concurrency assistants.

Larger models

Pool aggregate memory across the mesh to host models too big for a single node, with shards coordinated automatically.

A stack of four AgentBox nodes interconnected in a cluster with glowing teal cabling

Stack them on a shelf, rack them, or rail-mount them in a cabinet. Every node runs the same image and joins the fabric with one command.

Build your fabric

Start with one Pro and grow as your workload does. Cluster reservations lock founding-batch pricing per node.

Reserve your nodes →
Questions

Clustering FAQ

From 1 to 16. Nodes join a low-latency mesh and present to your apps as a single OpenAI-compatible endpoint.
No. The cluster looks like one endpoint, so you can scale from one node to sixteen without changing a line of integration code.
Yes. Aggregate memory is pooled across the mesh to host larger models — like Kimi-Dev 72B — with shards coordinated automatically.

More questions? See the full FAQ or talk to us.