Launched this week

Rondine
Run the right local LLM for your hardware
11 followers
Run the right local LLM for your hardware
11 followers
Rondine is an open-source control plane for local LLMs. It detects RAM and VRAM, recommends models that fit, applies hardware-tuned settings, downloads weights, and starts an OpenAI-compatible server. It supports Apple Silicon, NVIDIA GPUs, and DGX Spark through llama.cpp, MLX-LM, and vLLM. Instead of creating another inference engine, Rondine coordinates proven runtimes and shows every launch plan before execution.




Running local models is becoming much more interesting. Curious where you've seen the strongest pull so far—privacy, latency, cost, or something you didn't expect?