NeoSmith Code - Small Models. Big Intelligence

by
NeoSmith AI is proud to launch small, cost-efficient coding models that drops into whichever AI coding agent you already use. You get frontier model quality at 60% lower cost. Your IDE, your agent, your workflows shall be all unchanged.

Add a comment

Replies

Best
Maker
📌
Rising and unpredictable expenditures on AI coding Agents have become a significant financial burden for companies of all sizes. AI spend is outgrowing headcount and eating into corporate R&D spends. Uber, Microsoft cancelled their licenses. Introducing Neosmith Models that helps organizations achieve frontier-model quality at a fraction of the cost through below features: - Tiered AI Models for different tasks and complexity: We offer a choice of model tiers (Lite, Standard, Pro) behind a single endpoint, specially tuned for various SW development tasks. You can choose the model tier depending on the complexity of the tasks. - Frontier Accuracy: Due to our unique approach, our models' accuracy is equivalent to or better than the best frontier models (we measure on the same or higher public benchmarks that most frontier labs use). - Meet the developer where they are. Our model endpoints integrate seamlessly with popular Coding IDEs and agents, including Claude Code, VS Code, JetBrains, Codex, OpenCode, Cline, Continue. So you don't need to change any developer behaviour, your tools. Also all your automated workflows and agents built in these coding tools will work seamlessly. - Enterprise Grade Security, Governance and Controls: We provide full enterprise-grade security and data controls. We do not use your data to train our models.

NeoSmith Maestro


Hi Product Hunt 👋

I'm Udit, co-founder of NeoSmith. Before this I ran AI/ML platform engineering at Palo Alto Networks and was a founding engineer at Ola Krutrim, and the same conversation kept happening: a team adopts an AI coding assistant, developer velocity goes up, and eight months later someone in finance is staring at a six-figure monthly inference bill asking what exactly is being bought.

The uncomfortable part is that most of that bill is spent on tasks that don't need a frontier model. Renaming a variable, writing a test stub, applying a lint fix, summarising a diff — these get billed at the same rate as designing a distributed system.

Maestro is a model layer that sits underneath your existing coding agent. Instead of one large model doing everything, it orchestrates a bouquet of small ones — a cheap driver, a rescue model, a mid-size triage model — and escalates to a frontier tier only when execution-verified signals say it's genuinely needed. In practice that's under 0.5 frontier calls per problem.


Where it lands:


  • 92.2% pass@1 on LiveCodeBench v6 — +2.4 over Claude Fable 5, +4.4 over Opus 4.8

  • 74.2% pass@1 on SWE-bench Pro, one standard agent scaffold — ahead of Sakana Fugu (73.7) and Opus 4.8 (69.2)

  • ~$1.10 per 100 turns — roughly 15× cheaper than Opus, 30× cheaper than Fable


You point your tool at our endpoint and nothing else changes. OpenAI- and Anthropic-API-compatible, so Claude Code, Cursor, Cline and Continue all work unmodified.


We published the whole evidence bundle, not just the claims. Trajectories, generated patches, the grader, per-instance PASS/FAIL against gold, and per-call routing telemetry — enough to re-derive every number above yourself: →


Including the parts that don't flatter us. SWE-bench Pro is scaffold-dependent, and we beat Claude Fable 5 at pass@2, and we say so in the report rather than burying it.