HarnessRouter - Bring the world's best AI agents into your app, with one API

Building an AI agent backend yourself takes months: sandboxes, orchestration, retries, cost controls. HarnessRouter runs it for you. One API in, finished work out: code, files, videos, games. Trusted by top medical research institutions, leading healthcare companies, and cutting-edge startups in multiple domains.

Add a comment

Replies

Best

Hey Product Hunters,

We're the HarnessRouter team.

Building an AI agent backend yourself takes months: a sandbox per run, agent runtime, tool orchestration, files and artifacts, sessions and streaming, retries and timeouts, permissions, cost controls. And the upgrades, fixes, and maintenance never stop.

So we built HarnessRouter: bring the world's best AI agents into your app, with one API. Codex, Claude Code, Hermes, and more, running as your product's backend.


Here's what people are already building with HarnessRouter as their backend:

🎬 A cutting-edge AI video marketing startup uses HarnessRouter to drive their content generation loop, producing human-level viral video content.
🏥 A top medical research institution uses HarnessRouter to build an academic brain that connects all their operational data into a unified agent plane.
✅ A leading healthcare compliance company uses HarnessRouter to bake human domain expertise into agent skills, saving their customers hundreds of hours of labor.

What HarnessRouter does: your app sends a task through one API, HarnessRouter runs it using Codex, Claude Code, Hermes, and soon more, then sends back finished work: code, files, videos, games.

How you use it:
1. Drop our AGENTS.md into Cursor or Claude Code
2. Describe your idea, for example: "Build a founder launch-video app: users drop a logo and an idea, get a launch video."
3. Add your API key. Ship.

What happens behind the scenes: each task runs inside its own sandbox, traced step by step, so you can see exactly what the agent did. What comes back is structured and renderable, not just a chat reply: reviewable diffs, generated files, images, confirmations of real tool actions. And you're not locked into one harness. Claude Code today, Codex tomorrow, swappable with one line of config, and flexible enough to run thousands of harnesses concurrently.

Free during a 7-day trial. Would love for you to try it and tell us what your app should build. Reading every comment today. 🙏

Amazing launch congrats👏   does each sandbox have direct outbound internet access for tool scraping or can we configure strict egress rules?

 Thanks, Priya! 🙌 Yes, sandboxes have outbound internet access today (agents can fetch/scrape/install as needed), each isolated in its own VM. Configurable egress policies (allowlists/deny-by-default) are on the roadmap, and I would love to hear your use case, feel free to DM me!

 Great to hear.. Egress controls will be a huge win for enterprise use cases. sure i'll dm

   Sounds great, we'd love to collaborate with you!

Congratulations on your launch,  Curious about the handoff between agents. If I build with Claude Code today and switch to Codex later, how much context and setup carries over without rebuilding the workflow?

Thank you  , great question.

The parts we want to make portable are the tools, skills, and instructions, where the workflow and evaluation criteria live.

Think of it like Lego: tools, skills, and instructions are the reusable blocks. The harness is the base you can swap, test, and rebuild on.

If you start with Claude Code and later switch to Codex, you should not have to rebuild the whole workflow. You can keep the portable layer, run it against the new harness, compare performance, and adjust only the parts where that harness behaves differently.

That way, switching harnesses becomes an iterative optimization process, not a full rewrite. And we provide the observability and tracing tools to help you see what changed and adjust with confidence.

   let's visualize it.

   Here is an example how we provide tracing:

 Great question, Habib. We have a feature called fork a harness, which copies the configuration of the harness into a different base harness (codex -> claude code, hermes -> codex, etc) in one click. From that fork point, the builder or the builder's coding agent can iterate the harness configuration to bridge the harness specific setup differences. We call it vibe harness engineering.

Epic launch! One question I had while looking through it: what is the most common mistake teams make when they try to build multi-agent infrastructure on their own before using something like HarnessRouter?

 Great question, Mohit. I think teams, especially highly technical ones, tend to build their own agent harness infrastructure in-house, and that decision itself is surprisingly risky. We made the same mistake, until earlier this year we realized agent harnesses are getting commoditized fast. Competing with frontier labs, well-funded startups, and open source communities by hand-crafting our own harness just wasn't a wise use of our time.

So we chose to leverage rather than build. That freed us up to focus on what actually differentiates us, and it's let us build great applications both for ourselves (more launches coming soon!) and for our customers. HarnessRouter is us packaging up those learnings and handing the same leverage to the builder community.

That's a really interesting shift in thinking. The line "we chose to leverage rather than build" stood out to me. It feels like AI is changing where startups should spend their structuring time. Do you think, in a few years, the competitive advantage won't be who builds the best infrastructure, but who understands the customer problem best?

 inspirational discussion, Mohit. Don't get me wrong, I still believe the best infrastructure will be a competitive advantage. That's why we spend a lot of time designing the our harness service infrastructure in an elastic serverless way that is cost efficient, so we can best serve our customers. But the infrastructure advangate alone is not enough as a moat, that's why we send forward-deployed engineers / founders to customer site, work side by side with our customer, and make sure our product can really solve their problem. Yes, it's like Palantir, in an AI-native way of operation

   I love the “leverage rather than build” framing. I also fully agree that whoever owns the insight into the problem and the path to the solution will have the competitive advantage. Everything we are building toward today points to one vision: product follows your words. If you can articulate the problem clearly, a productized solution should be delivered almost instantly.

   I’ve personally built an agent harness to run my projects, so I know many ways to get it wrong.

You end up designing orchestration across tools, MCP, models, sandboxes, files, permissions, usage caps, retries, streaming, and observability. Each piece becomes almost its own product, and then you still have to think carefully about how all those pieces interact.

That’s why I think this is the better path for most teams: instead of building and maintaining the harness itself, stand on the shoulders of existing harnesses and customize the parts that actually encode your domain: instructions, tools, skills, evals, and product workflow.

Fewer infrastructure mistakes, less maintenance, and faster, better outcomes.

   yes, Mark and I had a dinner exactly 5 weeks ago, that's where this all started. We have independently pursued the same mountain from two different directions, and when we met in the middle, that's all this ideas emerge and HarnessRouter becomes a product.

Thanks for the thoughtful reply. I actually like the idea of pairing strong infrastructure with deep customer understanding instead of treating them as separate things. The Palantir comparison was especially interesting. It made me wonder where you think that customer knowledge eventually "lives." Is it mostly in the founders' heads through conversations, or do you wantedly try to turn those learnings into a repeatable system that improves the product over time? Curious how your team approaches that.

 This is such an inspiring question, Mohit! It touches on the deeper philosophy of how we think about the relationship between humans and AI in the future of knowledge work.

We believe customer knowledge will eventually live in a versioned, temporal knowledge graph—what we call a contextual graph. Retrieval should be not only domain-aware but also time-aware, enabling the system to model how knowledge evolves and supersedes previous knowledge over time.

The system becomes a digital twin of an enterprise’s collective human knowledge, though biases and hidden nuances cannot always be fully captured digitally. No system can represent reality with 100% accuracy. Instead, it must dynamically balance accuracy, quality, latency, and cost against the complexity of the real world.

The most important role of a service provider is to navigate that balance and deliver a system that satisfies the constraints set by the customer.

That actually clicked for me. It feels like the infrastructure is becoming real important thing, while the real differentiation shifts toward the quality of the domain knowledge and workflows you build on top of it. I'm curious, do you think that advantage eventually increases through proprietary data and customer feedback, or it is more about continuously refining the workflow itself?

 You are right on the point.

To answer your question: I think the advantage comes from both:

1. Deep knowledge of your customers’ problems.

2. Operational insight gained while solving those problems repeatedly.

The first is strengthened by proprietary data and customer feedback. The second comes from continuously refining the workflow: what to automate, where humans stay in the loop, which tools to use, and how to evaluate quality.

The real moat is when those two compound together: customer understanding turns into better workflows, and better workflows generate more insight from every run.

One API across agents solves the problem every team hits by month three, getting locked into a single vendor's agent right when a better one ships. Smart wedge. The place this gets sticky is evals and observability. Whoever owns routing is best positioned to tell me which agent wins for which task, and per-task performance data turns you from a router into the layer teams cannot rip out.

 you absolutely nailed it, Shivangi! Yes, just like there are a lot of model routers / token routers vendors in the market, the market need a vendor neutral platform and aggregator for the harness layer. Harness + Model = Agent. With the M harnesses * N models options, any business have the flexibility to avoid vendor lock-in for the best of their benefit.

 Hi Shivangi, you’re exactly right. Builders don’t want to be locked into a single agent, especially when the best system may combine different agents, each optimized for a specific task.

We’re already seeing labs use HarnessRouter’s infrastructure to evaluate models and agents across different tasks. Over time, this performance data can power smarter routing and become an increasingly valuable intelligence layer.

neat landing page and easy to understand explanation. tho, one question comes to mind: Since different harnesses output varying artifact formats, how does HarnessRouter enforce a consistent schema when returning structured results back to the host app?

 thank you for your insightful question! We built a OpenAI responsive API compatible contract for the harness response, and cannicalized all harness repsonse to this format. So that any AI app built with OpenAI responsive API can do a one line swap to upgrade their one-shot chat completion backend to an agentic intelligence backend.

   I love this conversation! This is exactly why we’re building a unified response contract, along with the router.

Different harnesses can produce very different artifacts, but the host app should receive them through one consistent interface: messages, files, artifacts, tool events, status, traces, and final outputs.

That contract is what lets multiple harnesses work behind the same app interface.

Building sandbox infrastructure from scratch is an absolute sinkhole for dev time.. Congrats 🙌.
quick question are there any plans to open-source parts of the harness or SDK down the road for local self-hosting?

   We love this question! Some parts of it are already on their way to being open source. We plan to build with the community and contribute as much as we can.

 Thanks, Vikram! 🙏 And yeah — that sinkhole is very real, that's kind of the whole bet behind building this.

Open-sourcing parts of it is something we've talked about internally. The parts that'd make the most sense to open up are probably the runner/SDK side (how a harness actually executes a turn) rather than the routing/billing/multi-tenant plumbing, which is more specific to running this as a hosted service.

No promises on dates, but if self-hosting is something you'd actually use, worth telling me more about the use case (compliance? cost? just wanting to hack on it?). That's the kind of signal that'd push it up the list!

💡 Bright idea
Not being locked into one harness is the claim I'd want to poke at, since swapping Claude Code for Codex with one line of config sounds simple, but different harnesses often disagree in how they interpret the same instructions, how they handle tool permissions, and what kind of output they naturally produce. If someone builds their AGENTS.md and their whole prompting approach tuned around Claude Code's habits, does switching the config line actually give equivalent results with Codex, or does it quietly need a rewrite of the instructions to get the same quality out of a different underlying agent. Also, on the healthcare compliance use case specifically, running sandboxed agents against real operational and patient adjacent data raises the obvious question of where that data physically lives during a run and after it finishes. Is anything retained on your infrastructure once a task completes, or does everything get destroyed with the sandbox, since that is usually the first thing a healthcare compliance team will ask before they trust a third party backend with their data.

 Great questions, both are things we learned the hard way rather than designed on a whiteboard.

On harness portability, you're right that a bare prompt tuned to one agent doesn't transfer for free, and we don't claim swap-and-forget. What makes switching practical for us is that the harness definition stays deliberately thin, and the domain expertise lives in two layers that are genuinely portable: skills (explicit, procedural instructions paired with deterministic scripts the agent invokes: recipes, not model-idiom prompting) and typed tools (every tool exposes a strict input schema; the schema is the contract, and both runtimes consume it the same way). The system prompt describes the job, not the model's habits.

The honest second half: every harness config ships with an eval suite, real domain cases scored against rubrics, with human-expert alignment checks on the grader. Switching the underlying agent means re-running the suite and fixing what it surfaces, which is usually tool-boundary discipline rather than instruction rewrites, where runtimes differ in how carefully they fill tool arguments, we add validation guards at the tool layer instead of forking the instructions. Today we run some verticals on one runtime and some on the other, chosen by eval results, not loyalty. So "one line of config" is the mechanism; the eval suite is what makes it safe to actually pull the lever.

On data residency: this is the first question every healthcare compliance team asks, and the answer is structural, not policy: agent runs execute in sandboxes inside the same private network where the data already lives, there is no third-party backend the documents travel to. Each run gets an isolated workspace; reference material is mounted read-only; the workspace is destroyed when the task completes. What persists is exactly what a compliance team wants persisted: the audit trail, what was asked, which tools ran, what was produced, not copies of source documents. Customer data is never used for model training. Happy to walk through the architecture in detail under NDA, that conversation is usually shorter than people expect because the answer to "where does our data go" is "nowhere."

  Great question and great answer. One thing I’d add: we think about routing at the level of the whole execution setup.

For a specific task, the relevant unit is [harness + model] plus [instructions + tools + skills]. Over time, the performance data for each setup: success rate, cost per task, latency, reliability, and output quality, becomes the asset.

Rather than comparing Codex vs. Claude in isolation, we help teams compare complete setups and route each task to the one that performs best.

Love the infrastructure angle.

Curious—do you think the long-term moat is access to many models, or helping developers choose the right one automatically?

Great question Ultimately, we see it working this way:

Different tasks are routed to different agent harnesses, models, skills, and tools. The routing is dynamic and adapts as agent harnesses and models evolve.

 I like that direction. If routing becomes the core abstraction, it feels like developers may eventually stop choosing models directly and instead optimize for outcomes.

I'd enjoy continuing this conversation and sharing a few thoughts on where I think this category is heading. If you're open to it, what's the best email to reach you?

 absolutely and I totally agree with you on aiming for the optimized results rather than worrying about selecting the models. Please reach out to us at Happy to have more conversations with you!

 inspirational question, Aryan! I in-vision a dynamic routing to the best harness + model combination adaptively will be a moat, and the orchestration layer stays on top of frontier model labs, which is to stay

 I like that framing. If the orchestration layer becomes the durable abstraction while models keep changing underneath, it creates a very different kind of moat.

I'd enjoy continuing this conversation and sharing a few thoughts on where I think this category is heading. If you're open to it, what's the best email to reach you?

 Hi Aryan, super excited for that! You can reach out to us at

 Thanks! I’ve just sent it over.

Looking forward to hearing your thoughts whenever you have a chance.

Congrats on the launch Richard!

 thank you for your support, Samir!

 Thanks for supporting our team Samir!

Congratulations!

 thank you for your support, Huisong!

  Thanks for supporting our team!

the eval-suite-gates-the-switch answer to Brandon above is a genuinely satisfying answer, more concrete than most vendor-lock-in claims I see here. question on that - are those eval suites something your team builds and maintains centrally per vertical, or can a customer plug in their own rubric and test cases for their specific use case before they trust a harness swap in production?

 great question, Gal. For enterprise customers with deep domain expertise, we send forward-deployed engineer to customer site and work with customer side by side, distill the human expertise into both the agent skills, AND the evaluation sets. To iterate, we build tools for human experts to easily rank and label different versions of agent response, and use that signals to gate go/no-go on harness swap

   We also built the observability and tracing layer to help teams decide whether a harness swap is actually production-ready, instead of treating it as a blind config change.

   that combo makes sense, thanks both. follow-up - is the tracing/observability layer something we'd get self-serve out of the box, or is it bundled with the forward-deployed engagement Richard described above? asking because the eval-suite value is clear when you've got an FDE sitting with you building the rubric, but a smaller team without that kind of engagement would still want some way to see "here's why this swap passed or failed" before flipping it in prod.

12
Next