We build agent products and MCP servers are becoming a real part of that surface area, so a dedicated testing/eval layer is exactly the kind of tool that should have existed already. Being able to run evals and CI/CD gates against a server before shipping it, instead of manually poking at it in Claude or ChatGPT and hoping nothing regressed, is the big win here. The multi-client angle (testing across ChatGPT, Claude, Copilot) is smart since MCP behavior isn't always consistent client to client.
MCPJam
Hey Gal! If you go directly into the platform, you'll be able to test this " just want a quick sanity check against a single client" in <5 sec! Let me know if this isn't the case - check quickly at app.mcpjam.com (you don't even have to sign-in!)