One endpoint.
Up to sixteen brains.
AgentBox nodes link over a low-latency mesh and present to your apps as a single OpenAI-compatible endpoint. Add boxes to run more agents in parallel, push more tokens per second, hold larger models in aggregate memory — without changing a line of integration code.
Cluster configurator
Aggregate capacity, based on Pro nodes (32GB · 6 TOPS each)
2 parallel nodes
Run two 7B models side by side, or one model with double the concurrent request capacity.
Throughput shown is relative and illustrative; real numbers depend on model, quantization and context length.
Three ways to spend more nodes
More parallel agents
Pin different models or agent roles to different nodes — a router, a coder, a reranker, a vision sidecar — all answering at once behind one gateway.
Higher throughput
Load-balance one model across nodes for near-linear gains in requests per second — ideal for batch RAG, embeddings and high-concurrency assistants.
Larger models
Pool aggregate memory across the mesh to host models too big for a single node, with shards coordinated automatically.
Stack them on a shelf, rack them, or rail-mount them in a cabinet. Every node runs the same image and joins the fabric with one command.
Build your fabric
Start with one Pro and grow as your workload does. Cluster reservations lock founding-batch pricing per node.
Reserve your nodes →Clustering FAQ
More questions? See the full FAQ or talk to us.