Launching today
Ferrum

Ferrum

Run local LLMs with one Rust binary

4 followers

Ferrum is an open-source, Rust-native runtime for running and serving local language models on Apple Silicon Metal and NVIDIA CUDA. One binary gives you an interactive CLI plus OpenAI-compatible Chat Completions and Responses APIs, without requiring Python, PyTorch, or vLLM at runtime. Choose a model explicitly, run it locally, or expose the same model over HTTP.

Ferrum makers

Here are the founders, developers, designers and product people who worked on Ferrum