Launched this week
oMLX

oMLX

Mac LLM server that cuts agent wait times from 90s to 5s

69 followers

oMLX turns your Mac into a full LLM inference server, run from the menu bar. It serves text, vision, OCR, embedding and reranker models with continuous batching, plus a RAM+SSD tiered KV cache that survives restarts, so Claude Code and Cursor respond in about 5s instead of 90s. OpenAI and Anthropic compatible APIs drop straight in. Native Swift, not Electron. Apache 2.0, open source.
oMLX gallery image
oMLX gallery image
oMLX gallery image
oMLX gallery image
oMLX gallery image
Free
Launch Team
Wispr Flow: Dictation That Works Everywhere
Wispr Flow: Dictation That Works EverywhereStop typing. Start speaking. 4x faster.
Promoted