AJRouter is a local Windows control plane for llama.cpp and audio.cpp. Start and stop workloads from a loopback dashboard, route models behind an OpenAI-compatible text API, monitor GPU usage and logs, and coordinate text, speech and music workloads. It bundles the runtimes and installer; you bring your own GGUF models. No cloud account, subscription or telemetry.
Hi, I’m AJ. AJRouter grew out of my Windows local AI setup becoming too many manual steps: starting services, switching GGUF profiles, watching VRAM, and keeping text and audio workloads from colliding. I built a control plane around llama.cpp and audio.cpp so the setup is repeatable and coding agents can use a stable local API. On my RTX 5060 Ti 16 GB, one surprise was that pushing MTP draft depth higher made a five-task test slower. I documented the build, measurements and their limits here: https://ajthe.dev/case-studies/p...
This is a paid Windows download with the runtimes included; you supply your own GGUF models. If you run local models, what part of operating the stack costs you the most time?