Launching today

OptiQ
Quantize, serve and code with local models on your Mac
2 followers
Quantize, serve and code with local models on your Mac
2 followers
OptiQ is the local-LLM stack for Apple Silicon. It measures each layer's sensitivity and gives sensitive layers more bits, so its mixed-precision quants beat a flat 4-bit quant at the same size. One MLX-native engine runs a CLI and OpenAI/Anthropic-compatible server, a local web workbench (Lab) and a terminal coding agent (Code). On Gemma-4-12B (M3 Max) it decodes 1.7x faster on prose and up to 2.9x on code edits, identical output. 200+ quants on Hugging Face. No PyTorch, no cloud.




Hey Product Hunt 👋
OptiQ started from a simple frustration: local models on a Mac usually mean one flat 4-bit quant, whether a layer can take it or not.
OptiQ measures each layer's sensitivity (a KL-divergence pass) and spends the bit budget where it matters. Sensitive layers stay at 8-bit, the rest drop to 4-bit, at the same average size. The gap is biggest on small models: +4.4 Capability Score on MiniCPM5-1B, +3.2 on Spark-X2.5-4B and +1.9 on Qwen3.5-4B over uniform 4-bit.
It's one MLX-native install with three tools: a server that speaks both the OpenAI and Anthropic APIs, a local web workbench (optiq lab) to chat, quantize, fine-tune and compare models, and a terminal coding agent (optiq code) for whatever model you're serving. On Gemma-4-12B on an M3 Max we measured 1.7x faster decode on prose and up to 2.9x on code edits, with identical output (mlx-optiq.com/compare).
200+ quants are on Hugging Face and load with stock mlx_lm.load. `pip install mlx-optiq`
What we'd most like to hear: which model you'd want quantized next, and what breaks on your Mac.