the multi-expert setup actually holds up once you push on it, which is rare for this kind of demo. I spent a while going back and forth with the makers on how disagreement between experts gets handled in the merge step, and their answer (an editor/moderator pass rather than just averaging outputs) matches what I saw when I tried a task with genuinely conflicting expert takes. the thinking-process view where you can watch the debate happen is the single best trust signal I've seen in a tool like this, way better than just trusting a summary