Launching today
Ferrum

Ferrum

Run local LLMs with one Rust binary

4 followers

Ferrum is an open-source, Rust-native runtime for running and serving local language models on Apple Silicon Metal and NVIDIA CUDA. One binary gives you an interactive CLI plus OpenAI-compatible Chat Completions and Responses APIs, without requiring Python, PyTorch, or vLLM at runtime. Choose a model explicitly, run it locally, or expose the same model over HTTP.
Ferrum gallery image
Ferrum gallery image
Ferrum gallery image
Free
Launch Team
Skydive
SkydiveAI coworkers that actually get work done
Promoted