Under the hood: how Linda runs a computer-use agent locally on your Mac with MLX
A few people asked what runs under the hood, so here is the tech core and the decisions behind it.
Local first, on MLX
Linda runs open models directly on Apple Silicon through MLX: GLM, Qwen, Gemma and Granite. It works on macOS 26, on any Mac from M1 to M5. No cloud by default, so your screen and your data stay on your Mac. If you want a frontier model for a hard task, you can plug in your own API key.
Why several model families
No single small model wins at everything. Chat, vision on screenshots and reliable tool calls stress different strengths, so Linda lets you pick and switch instead of locking you to one model.
Laya, the decision layer
Local models get slow and confused when you feed them long histories. Laya sits between the conversation and the model. It compacts the history and decides what context the model needs for the next step. Fewer tokens in means faster answers and fewer mistakes on a laptop GPU.
Memory and computer use
Linda keeps memory in a vector database (Qdrant), so it recalls past context without stuffing everything into the prompt. On top of that, Linda sees your screen and acts: it clicks, types and moves across apps. You can also drive it by voice.
What I'm weighing next, and I'd love your input
Teach by demonstration: you show the steps once, Linda repeats them with new data.
Automations: run tasks on a schedule, a cron, or when an event happens.
Which one would you use first? Which task on your Mac would you hand over today? And if you run local models on MLX, which ones work best for you?

Replies
Be the first to reply
Have a question or a thought to share? Add a comment above to start the conversation.