While building Arcytic, we noticed something unexpected.
The biggest AI failures weren t obvious hallucinations.
They were confident answers that were subtly wrong and hard to catch.
When we compared the same question across multiple models, two patterns kept showing up: