Cluster Configurator

Size an on-premise MindRouter inference cluster. Drag the radar handles to describe your workload, including the intelligence level you need relative to today's open-weight frontier, and get a hardware build, price range, and power estimate. From a single NVIDIA DGX Spark to a cluster of HGX B300s.

Planning tool only. We sell nothing and endorse no brand. MindRouter is free, open-source software. The vendors, hardware, and models named here (NVIDIA, Supermicro, and others) are reference examples used to estimate cost and power. Comparable equipment and models from other vendors and brands will work just as well.
Pricing data updated · USD street prices · estimates, not quotes

Sized down to a single DGX Spark? Our companion DGX Stack repo is a ready-to-run vLLM + OCR deployment for that machine, with optional MindRouter integration.

Workload Profile — drag the handles
Dimensions influence each other — raising demand lifts the budget floor; dragging budget down trades away intelligence, then throughput and concurrency.
Fine-tune exact values
Recommended Build
Estimated Hardware Cost
Sizing assumptions
Models are served quantized (NVFP4 primary; FP8 where VRAM allows) on vLLM-class serving stacks. A model must fit within one server's pooled GPU memory (DGX Spark may pair two linked units). Concurrency assumes ~35% of connected chat users generate at once and API traffic peaks at ~6× its weekly average. Power assumes each GPU averages ~50% of its maximum draw during inference serving (band shown: 40–65%), plus platform overhead. Prices are USD street-price ranges compiled from vendor and reseller listings (see the data date above), plus a 3–10% integration allowance. Supermicro is used as the reference server vendor. Estimates only — not a quote.