Launching today

Ferrum
Run local LLMs with one Rust binary
4 followers
Run local LLMs with one Rust binary
4 followers
Ferrum is an open-source, Rust-native runtime for running and serving local language models on Apple Silicon Metal and NVIDIA CUDA. One binary gives you an interactive CLI plus OpenAI-compatible Chat Completions and Responses APIs, without requiring Python, PyTorch, or vLLM at runtime. Choose a model explicitly, run it locally, or expose the same model over HTTP.
