Launching today

ds4
Local inference engine for DeepSeek V4, GLM, Qwen
1 follower
Local inference engine for DeepSeek V4, GLM, Qwen
1 follower
DwarfStar 4 (ds4) is a narrow C inference engine from Salvatore Sanfilippo (antirez) for running DeepSeek V4/V4.1 Flash, GLM 5.x and Qwen3.8 Flash Next locally on high-memory Mac, CUDA and ROCm machines. It uses asymmetric quantization on routed experts and a KV cache that persists to SSD and survives restarts. Ships a CLI, OpenAI/Anthropic-compatible local APIs, and a native coding agent sharing the same model state. MIT licensed. Benchmark: 34.4 t/s generation on M5 Max 128GB.


Free
Launch Team

Framer AI AgentsDesign and publish professional sites with AI
Promoted