Your private AI cluster,reachable from anywhere.
Turn the machines you already own into one OpenAI-compatible inference cluster. LAN discovery, cloud relay, and distributed model sharding — all in one app.
Need help? Install guide · GitHub
Built for real hardware
Everything you need to build a private inference network
Zero-config networking
Nodes connect outbound to the relay. No firewalls, no port forwarding, no networking headaches.
Distributed inference
Split a model across machines — every node's GPU and CPU contributes to each token.
OpenAI-compatible
Point Cursor, Open WebUI, or curl at your cluster. Drops in anywhere the OpenAI SDK works.
Local-first
Models run on hardware you control. The relay only brokers metadata and token streams.
Auto-updating
One click updates the app in place — no reinstall, no browser. New features land automatically.
Real-time monitoring
Live system pill shows node count, VRAM usage, GPU type, token throughput, and active requests.
Up and running in 3 steps
Download & launch
Install on macOS or Windows. The setup wizard auto-detects Ollama, LM Studio, and your GPUs.
Connect your machines
Generate an API key on the dashboard. Paste it into each machine. They auto-connect on every boot.
Chat or distribute
Use the Playground for load-balanced routing, or distribute a model across machines for combined compute.