Launching today

iFixAi
Independent auditing of AI agents to uncover misalignment
180 followers
Independent auditing of AI agents to uncover misalignment
180 followers
iFixAi is an independent auditor helping companies assess whether they can trust their AI agents. Unlike relying solely on evals and observability tools, its multifaceted audit includes 250 inspections across 69 categories of AI misalignment, combining AI red teaming, operational assurance, and philosophical, ethical, and sociological perspectives. It identifies failures under testing, explains their business implications, and provides evidence engineers can use to investigate and fix them.











Great launch! I have two questions about complex multi-agent architectures and pre-authenticated environments, based on our current setup:
Multi-Agent Dynamic Workflows: An orchestrator agent reads each turn and routes the conversation to one of several specialized sub-agents (e.g., orders, returns, account changes). Each sub-agent has its own instructions and its own subset of MCP tools, and receives a handoff payload with the conversation state it needs. Routing depends on session context and on what the customer decides mid-conversation. A customer might ask about an order, decide to return it, then want the refund sent elsewhere, passing through three agents in one conversation.
Can the simulation environment model the routing rules (which decision should lead to which agent) and the handoff payload (what each agent should and shouldn't receive)? Or is each sub-agent audited on its own, and if so, how is the orchestrator covered?
Can you run multi-turn scenarios that validate the state transitions?
the flow switches to the right agent when the customer's decision calls for it
it doesn't switch on ambiguous input, or on instructions injected through a tool result or document
after a handoff, the previous agent's tools and context are no longer in play
Will each finding name the agent, handoff and turn where it happened? For example, "orchestrator routed to the wrong agent" vs. "returns agent received more context than it needed."
Pre-Authenticated Contexts & MCP Security: In the demo, the audit flags an agent for disclosing customer data without verifying identity first. In our architecture, authentication is handled out-of-band. Customers sign in to the app before they can reach the assistant, the session token is forwarded with every MCP call, and each MCP server authorizes the call and scopes data to the customer bound to that token. The model doesn't handle credentials or verify identity itself. Some actions, liek account changes, only become available once the app has confirmed authentication.
How do simulated users get their iidentity? We'd want each persona to run with its own test account and token, so our MCP servers enforce access exactly as in production, rather than identity being claimed in the chat. The open-source HTTP adapter seems to use one token per run. Can personas map to separate tokens?
Can we declare the channel as pre-authenticated, so the audit skips in-chat verification checks and probes the real boundaries instead? The open-source authorization checks look role-to-tool, but our main risk is object-level (right tool, wrong customer):
one customer asking for another's data ("it's my husband's account"), or another customer's ID planted in a document or tool result
identity context lost, swapped or widened during handoffs between agents
tools returning data outside the session's scope, and the agent repeating it
For unauthenticated or expired sessions, can we check that protected actions are refused? We'd also want to confirm the agent doesn't fall back to "verifying" customers in chat by asking for personal details.
If the model attempts something the tool layer blocks, is that reported as a model finding, a passed control, or both? We'd want both signals.
We'd test against staging with synthetic customers. What would you need from us: test accounts per persona, a token-minting endpoint, something else?
iFixAi
@sweethuman Thank you, very good questions! This is exactly where "the task got done" isn't enough.
When you connect via GitHub or MCP, we build the simulation environment around what your agents are supposed to do: workflows, roles, rules, permissions and the tools each one calls. You review it as a summary or full YAML before the audit, so you can check that your routing and the separation between agents are reflected correctly.
Your main risks are right in our scope: agents staying within technical permissions while exceeding business authority, and instructions injected through documents or tool results. Start with the Authority and Human Control and Adversarial Resilience bundles. Findings come with the test details, observed behavior, affected tools and evidence, so engineers can locate the issue.
You're right that in-chat identity checks don't fit an out-of-band auth design; "right tool, wrong customer" is the real thing to probe. And since your agents will keep changing, you can run a fresh audit after each change to see what improved or regressed. Don't forget code PH1000IF for your first audit!