Merge - AI-native code review assessments

Merge is a online code review assessment that helps engineering teams assess engineering judgement. With Merge, candidates review a PR, just like on the job. Then, our AI agent addresses PR comments in realtime, simulating a real engineer. At the end, we assess bug coverage, communication, PR quality, and token use efficiency.

Add a comment

Replies

Best

Hey ProductHunt 👋!

We’re 5 founders who’ve collectively done over 250 interviews - everywhere from startups to FAANG+ to quant shops. Most of the interviews we’ve done were Leetcode based or tested skills that weren’t used on the job.

Meanwhile, at each of our companies, though, PR counts have nearly tripled. Most of us haven’t manually edited a line of code in a year. Our teams are putting more and more emphasis on code and architecture reviews, yet hiring processes haven’t changed whatsoever.

Even the AI-assisted ones we’ve done still assess code output as the primary evaluation metric. As agents develop, we truly believe this will not be the most challenging ability for an engineer to have.

We’ve seen that the real difficulty with using AI is not just reviewing code your AI generates; it's reviewing code that another engineer’s AI has generated, having little context yourself.

That’s why we’re launching Merge today: to help hiring teams assess engineering judgement. Here’s how it works:

1. Candidates are shown a small codebase to understand and a PR to review and comment on.

2. An AI agent addresses each PR comment via a code change or reply, simulating a real engineer.

3. Candidates can repeat until 5 revisions are used up or time runs out.

Throughout this process, we assess the following:

1. Coverage - How many bugs or vulnerabilities did the candidate identify and address?

2. Communication - Was the candidate efficient and constructive with their feedback?

3. Efficiency - How many revisions and tokens did the review take?


We’re the first platform that can show you exactly how efficient a candidate is with token use, LLM costs, and PR revisions — all of which are becoming exceedingly important in industry positions.

If you’re interested in the next-generation of engineering hiring, book a demo with us! Feel free to ask any questions below as well.

 Hey. On the hiring side, take-homes and live coding both broke for us once candidates started running the same AI we do. Curious how you're scoring: the final diff, or the reasoning during the review itself? Mostly wondering how you keep it from just rewarding whoever has the better model open in another tab.