Same task, different destination: DungeonQ's real MCP comparison

An unchanged file does not tell you whether an agent read it. Our new DungeonQ walkthrough checks the destination and the source's business-read journal as well.

Astra max received the same artificial document-review task through real MCP in two operator-provisioned routes. It chose its own calls to inspect the document, save a note, and read that note through a scoped ticket.

Across two successful paired runs:

• DIRECT reached the artificial source: 2 business reads in one run, 3 in the other.

• DIVERT worked in the persistent synthetic world: 0 source business reads in both runs, with different document bytes.

• All four sessions completed the five tool categories and retrieved the same note they had saved. The artificial target remained unchanged.

Separate scripted checks show a useful world-only Wrong Ticket being rejected by the origin, and a saved note surviving an actual runtime restart. The operator retains observations; any bounded adaptation still needs separate authority.

This is the mechanism behind DungeonQ's defensive deception runtime: give designated suspicious AI-agent sessions somewhere consistent to work while people or authorized defender agents prepare a response. The new evidence checks routing and authority; it does not measure how much response time is gained.

The boundary matters: these were artificial-file workflows with routes chosen in advance, not autonomous attacks. We do not claim Astra was fooled or that production systems were protected. One earlier setup attempt made no runtime call and remains INCONCLUSIVE.

The 2:04 film combines actual operation footage with a clearly labeled recorded-results summary:

The Apache-2.0 host example, installation steps and evidence summary are public. Its transport and boundary tests need no model API key or subscription:

19 views

Add a comment

Replies

Best

Thanks for putting the code and install steps on GitHub, and for making the transport tests work without a model API key. That makes it easy to try. Congrats on the walkthrough, Ranopha!

Reading the comparison, I’d be interested in an interrupted write: the synthetic world saves a note, but the client never receives confirmation. Can the operator’s journal distinguish that from a call that never committed? That distinction mattered in Turnfeed’s retry handling—an agent’s “failed” response wasn’t enough evidence to retry a public write. I haven’t run DungeonQ; this is a suggestion for a future walkthrough.

 Thanks, Theodor — agreed: a failed agent response is not enough evidence to retry a write. The earlier walkthrough did not make that distinction clear enough.

I’ve now added read-only write-status inspection in DungeonQ 0.14.0. A verified journal entry returns COMMITTED with the original saved result. An exact authenticated no-effect refusal can return NOT_COMMITTED at its checkpoint; missing or ambiguous evidence remains UNKNOWN. Reusing an identity with different input returns CONFLICT. None of these automatically authorizes a retry.

The new walkthrough actually drops the facade connection after saving the note. The MCP consumer recovers through a read-only lookup, independently reads the note back, and finds the same commit after restart — one write, zero automatic write retries. The original UNKNOWN transport history is retained, so that deliberately faulted run still fails overall acceptance.

The film, raw results, source and reproduction instructions are together here:

This is a scripted test with artificial resources, not a new model evaluation or a guarantee for external public writes. Your interrupted-call example was a useful gap to close.