A few hours into our launch, Lenz is now sitting at #3 Product of the Day on Product Hunt.
For a small, bootstrapped team, this is a pretty special milestone. We re incredibly grateful to everyone who has checked out the launch, tried the product, asked thoughtful questions, and shared feedback along the way.
Congrats on the Launch @kostaj This hits close to me, spent Years in Security Testing and the same problem shows up there: one Models blind spot becomes everyones blind spot if you don't cross check.
Lenz
@bdennis11907 Very much our experience too. Even just organizing an adversarial debate with the same model still helps reduce blind spots somewhat.
Lancepilot
Congrats team..idea of having a jury review the evidence is quite interesting.
Lenz
Thanks, @istiakahmad! Most of the standalone ideas implemented in Lenz (the jury review, the adversarial debate, the separate research step, the use of different models, the preliminary framing step, etc.) are well-researched and known to improve accuracy. We tested and fine-tuned a specific ensemble and packaged it as an easy-to-use API so everyone can use it when verification matters.
EverTutor AI
Love the idea and positioning! Excited to see where it goes huge congrats on the launch
Lenz
@suryansh_tiwari2 Thanks!
Congrats on launch number two. Publishing the data and prompts for anyone to pick apart is a good look.
Paint the Cameras Dead
This feels especially important right now. We are living in a world where AI is becoming one of the first places people go to check what is true, yet your research shows that frontier models disagree on 63% of real-world fact checks.
That is such a powerful reminder that a confident AI answer is not the same thing as a verified answer.
Really love what you’re building with Lenz. The need for trustworthy verification is only going to grow as AI becomes more embedded in how we work, learn, and make decisions.
At this point, I’m just happy to have an AI that occasionally says, “Let me check that” instead of confidently making things up 😂
Lenz
@bogomep Indeed. In our experience, the models' self-reported level of confidence in their answers is mostly noise. They aren't trained to self-assess. Even worse, their incentives (to keep the user happy and engaged) push them to show high confidence when asked. On a 1-10 scale, in 76% of answers, the models report confidence of 9 or 10. And the correlation between confidence level and how much they agree with each other is not too strong either: https://lenz.io/research/llm-disagreement#confidence