Launch edition  Two new boxes  Reservations open

Cloud-grade AI.
Your hardware. Your data.

AgentBox is a rugged, palm-sized appliance that runs 3B–7B language models natively on a 6 TOPS NPU. No tokens. No metering. No data leaving the building — and it scales from 1 to 16 nodes when you need more.

Fully refundable reservation  Ships Q4 2026  Founding-batch pricing locked

AgentBox edge AI appliance — compact dark aluminum chassis with a glowing teal light seam, floating on black

Built on a proven, second-sourced edge stack

┷ Rockchip RK3588┷ 6 TOPS NPU┷ Ubuntu 22.04 LTS┷ Ollama┷ llama.cpp┷ NVMe storage┷ Docker / containerd┷ A/B OTA updates ┷ Rockchip RK3588┷ 6 TOPS NPU┷ Ubuntu 22.04 LTS┷ Ollama┷ llama.cpp┷ NVMe storage┷ Docker / containerd┷ A/B OTA updates
$0
Per-token API fees
6 TOPS
On-device NPU
100%
Local / air-gappable
1→16
Clusterable nodes
Why AgentBox

The cloud is rented. Metal is owned.

Three principles drive every design decision — no black boxes, no surprise bills, no data exhaust.

Sovereign isolation

Compute belongs to the owner. Run high-utility models entirely inside your network perimeter — air-gap it if you want. Nothing phones home.

Predictable economics

One capital purchase replaces a metered cloud bill that scales with usage. The more you run it, the more it saves. Capped, not metered.

Frictionless stacks

Ships with Ubuntu 22.04 LTS, Ollama, llama.cpp and a container runtime pre-wired. Your first token is five minutes away, not five days.

Engineered, not assembled

Precision metal,
tuned for sustained inference.

CNC aluminum that doubles as a heat spreader. Active cooling sized for hours of load, not a boot demo. Field-serviceable NVMe and wireless. This is hardware built to run hot models, quietly, for years.

8-core
ARM CPU (RK3588)
up to 32GB
LPDDR memory
256GB
NVMe storage
2.5GbE
networking
Full specification
Macro close-up of AgentBox black anodized aluminum heatsink fins glowing teal between the fins
Macro of the RK3588 system-on-module carrier board with NVMe slot and gold contacts
Macro of AgentBox rear I/O panel — Gigabit Ethernet, HDMI and USB ports backlit in teal
Three-quarter angle render of the AgentBox appliance showing the heatsink top and accent LED
Scale-out architecture

Start with one.
Grow to sixteen.

AgentBox nodes link over a low-latency mesh and present as a single endpoint. Run bigger models, more agents in parallel, and higher throughput — just by adding boxes. No re-architecture.

2 nodes4 nodes8 nodes16 nodes
See cluster scaling
An array of AgentBox nodes in a parallel compute cluster with teal status lights
Do the math

Most teams break even in under 6 months

Shift steady inference off metered cloud APIs and onto a box you own. Our calculator shows your payback period against GPT-class and open-model API pricing.

Calculate your savings
Questions

Straight answers

3B-class models run comfortably on every SKU. The Pro (16–32GB) targets 7B quantized models as its primary workload, plus embeddings, rerankers and RAG stacks. For larger models, cluster multiple nodes.
For steady, repeated inference it usually is. A metered API bills every token forever; AgentBox is a one-time purchase. The break-even depends on your volume — try the ROI calculator with your own numbers.
It boots into an appliance mode with a first-run wizard for network, registration and model install. Ollama and an OpenAI-compatible API are pre-wired — point your existing app at the box and go.
Yes — nodes link over a low-latency mesh and present as one endpoint. Scale from 1 to 16 boxes for more parallel agents, higher throughput or larger models. See the Clusters page.
Founding batch ships Q4 2026. A reservation holds your place and founding-batch pricing, and is fully refundable until your unit ships.

See the full FAQ for models, pricing, privacy and setup.

Own your intelligence.

Reserve an AgentBox today. Lock founding-batch pricing, hold your place in the Q4 2026 run, and stop renting your own data back from the cloud.