MonkeysCode is an agentic IDE and a standalone desktop Code Agent, powered by Capuchin, our own AI model. Capuchin scores 62.1% on SWE-bench Pro against Claude Opus at 64.3%, and serves 92.7% of production requests.
Run agents on your machine, a remote dev box over SSH, a cloud container, or a peer machine. 50+ permission-gated tools. Checkpoints on every file. Tests run before anything lands.
Claude, Gemini and ChatGPT included in every plan, or bring a local model. Free 30 days, no card.
This is the 2nd launch from MonkeysCode. View more
MonkeysCode Code Agent
Launching today
Run AI coding agents local, over SSH, or in the cloud
Most AI coding tools assume your code is on the laptop in front of you. The MonkeysCode Code Agent runs agents on your local machine, a remote dev box over SSH, a cloud container, or a peer machine. Nobody else does SSH.
It's powered by Capuchin, our own model: 62.1% on SWE-bench Pro vs Claude Opus at 64.3%. 50+ permission-gated tools, checkpoints on every file, tests run before anything lands.
Same engine as our IDE, included in every plan. Free 30 days, no card.
Hey Product Hunt 👋
We launched the MonkeysCode IDE here a while back. This is the thing people asked for most: a way to run agents without the editor open.
The Code Agent is a desktop app. Start a run, walk away, come back to a reviewed diff with tests green.
The part I'm most pleased with: it runs where your code actually is. Four backends, meaning local, SSH to a remote dev box, a cloud container, or a peer machine over WebSocket. Every other agent tool assumes you're sitting at the machine with the code on it, and that's not how most teams work. There's a dev box. A build server. A staging environment.
Two smaller things worth mentioning:
Checkpoints don't use git stash. Every file gets snapshotted before the agent's first change. We built our own store rather than using git stash push -u, because that would destroy your own uncommitted work alongside the agent's.
Diagnostics only report what the agent broke. It captures a baseline when the run starts, so if your codebase already has 400 warnings, you don't get 400 warnings.
We also published benchmarks for Capuchin, our own model: 62.1% on SWE-bench Pro against Claude Opus at 64.3%. We're 2.2 points behind, measured on our own production endpoint with the methodology published.
Everything's in existing plans at no extra charge. Free 30 days, no card.
Happy to answer anything about the execution backends, the sandbox model, or the economics of serving your own model instead of reselling an API.
Jorge
Report
No reviews yetBe the first to leave a review for MonkeysCode