New · Launch edition

Local coding models,
on hardware you own.

The AgentBox Developer Edition runs open coding LLMs entirely on local metal — autocomplete, refactors, test generation and agentic workflows that never ship a line of your source to a third party.

50% off launch $1,798 $899 reserve · founding batch

Ships Q4 2026 · Sold at cost to launch the edition · Fully refundable reservation

AgentBox Developer Edition carrier board with NVMe storage
The lineup

Pick your coding model. Run it locally.

Every model below is open-weight and pre-tuned for the RK3588 NPU. Swap between them from the appliance console — no re-imaging, no cloud round-trip.

CodeGemma

Google

Google's Gemma-family coding model. Strong fill-in-the-middle completion and instruction following at 2B and 7B — a great everyday default on a single node.

Runs on: single node

Qwen2.5-Coder

Alibaba

Top-tier open coding accuracy for its size. The 7B variant fits a Pro-class node and leads most local HumanEval comparisons.

Runs on: single node

DeepSeek-Coder

V2-Lite

A lean mixture-of-experts coder with broad language coverage and long-context repo reasoning, quantized to run comfortably on-device.

Runs on: single node

StarCoder2

BigCode

Permissively licensed and trained on a transparent code corpus — a dependable base for completion and fine-tuning on your own repos.

Runs on: single node

Kimi-Dev

Cluster

Moonshot AI's 72B software-engineering model. Too large for one box — pool aggregate memory across a clustered AgentBox to host it, and let the mesh coordinate the shards.

Runs on: multi-node cluster (see Clusters)

Your model

BYO-weights

Bring any GGUF or Ollama-compatible model, or fine-tune an open base on your private codebase. The box is yours — so is what runs on it.

Runs on: single node or cluster

Ollama llama.cpp OpenAI-compatible API VS Code / JetBrains bridge GGUF quantization

Code never leaves

No proprietary source, secrets or customer data sent to a hosted model. Air-gap it and your IP stays inside the perimeter — by construction.

Zero per-token billing

Autocomplete fires thousands of times a day. On metered APIs that compounds forever; on a box you own it's free after the one-time purchase.

Drop into your editor

Exposes an OpenAI-compatible endpoint, so VS Code, JetBrains and your CI agents point at the box and just work. First completion in five minutes.

Developer Edition spec

Tuned for sustained inference loops.

Built on the Pro chassis with developer defaults pre-wired: model registry, editor bridge, and a container runtime ready for your RAG and agent stacks.

16–32GB
LPDDR memory
512GB
NVMe (model cache)
6 TOPS
on-device NPU
1→16
clusterable nodes
Compare all hardware
AgentBox · Developer Edition
50% off launch
$1,798 $899

reserve · founding batch · fully refundable

Reserve Developer Edition →
  • CodeGemma · Qwen2.5-Coder · DeepSeek-Coder · StarCoder2
  • Kimi-Dev 72B across a cluster
  • OpenAI-compatible API + editor bridge
  • Ollama · llama.cpp · containerd pre-wired
  • Fully refundable until your unit ships in Q4 2026
Questions

Developer Edition FAQ

On a single node: CodeGemma (Google Gemma), Qwen2.5-Coder, DeepSeek-Coder-V2-Lite and StarCoder2. You can also bring any GGUF or Ollama-compatible model, or fine-tune an open base on your own code.
Yes — Kimi-Dev is a 72B model, too large for one box, so it runs across a clustered AgentBox that pools aggregate memory across nodes. See the Clusters page for sizing.
No. Inference runs entirely on the box — no proprietary source, secrets or customer data is sent to a hosted model. You can air-gap it and keep your IP inside the perimeter.
Yes. It exposes an OpenAI-compatible endpoint, so editors, CLI tools and CI agents that speak that API point at the box and just work.
It ships Q4 2026. Reservations are fully refundable until your unit ships, and founding-batch units are sold at cost to launch the edition.

More questions? See the full FAQ or talk to us.