I spent six months building an app around that idea and I'm still not sure it's true.
The setup: four models from different providers answer the same question independently, without seeing each other's replies. A moderator then compares them and reports where they agree, where they conflict, and what to do next.
Sometimes the disagreement is the whole point one model catches a risk the others missed. Other times you get four versions of the same answer and the extra complexity bought you nothing.
So I'm curious about your experience: where did multiple independent answers genuinely win over one, and where was it just overhead?