Cluster Configurator
Size an on-premise MindRouter inference cluster. Drag the radar handles to describe your workload, including the intelligence level you need relative to today's open-weight frontier, and get a hardware build, price range, and power estimate. From a single NVIDIA DGX Spark to a cluster of HGX B300s.
Planning tool only. We sell nothing and endorse no brand.
MindRouter is free, open-source software. The vendors, hardware, and models named here
(NVIDIA, Supermicro, and others) are reference examples used to estimate cost and power.
Comparable equipment and models from other vendors and brands will work just as well.
Sized down to a single DGX Spark? Our companion DGX Stack repo is a ready-to-run vLLM + OCR deployment for that machine, with optional MindRouter integration.
Workload Profile — drag the handles
Dimensions influence each other — raising demand lifts the budget floor; dragging
budget down trades away intelligence, then throughput and concurrency.
Recommended Build
—
Estimated Hardware Cost
—
Sizing assumptions
Models are served quantized (NVFP4 primary; FP8 where VRAM allows) on vLLM-class
serving stacks. A model must fit within one server's pooled GPU memory (DGX Spark may pair
two linked units). Concurrency assumes ~35% of connected chat users generate
at once and API traffic peaks at ~6× its weekly average. Power assumes each GPU
averages ~50% of its maximum draw during inference serving (band shown: 40–65%), plus
platform overhead. Prices are USD street-price ranges compiled from vendor and
reseller listings (see the data date above), plus a 3–10% integration allowance. Supermicro
is used as the reference server vendor. Estimates only — not a quote.