QikLM is web-based AI orchestration that runs on your local hardware. It connects directly to your llama.cpp or vLLM engines. There is nothing embedded. No Docker. No Python. No npm. Just one small binary. Features: Unified API: Run local models with cloud APIs like DeepSeek or GLM. Auto-Engine Boot: Auto-boot on prompt with no preloading. Chat Workspace: Full chat UI with web search, GPU monitoring, and LAN access. Integrations: Claude Code, Codex, and more. Privacy-first. Self-hosted. Qik.
No reviews yetBe the first to leave a review for QikLM
Maker
📌
Running local AI or large language models is a great way to have privacy and sovereignty. However managing the models and inference is currently a chore. There are currently other apps that try to make this easier but they fall short in some areas. I was consumed with frustration of having to sacrifice raw power for ease of use. Most setups required separate components like chat webUI's running in docker and all these different configurations just to have a setup. I wanted an all in one system that made it easy to work with local and external commercial models and run the inference engines directly so I could control the version of the engine to use. When the engine is bundled with software, it's not possible to try the latest features until that software is finally updated and that can take weeks or longer depending on their update scheduling. QikLM solves most of the friction of running your own local setup and running the inference binary of your choice with easy model management and first class agentic integrations.