Agenomics is an open-source methodology and toolkit for evaluating AI agents through structured agent genomes, Trust Score, Compatibility Score, behavioral drift monitoring, and real-world evidence collection. Connect your agents, record observations and incidents, and test whether declared trust actually corresponds to behavior in production.
AI agents are becoming more autonomous — but our ability to measure whether we should trust them hasn't kept up.
We built Agenomics around a simple question:
**Can an AI agent's declared trust level be tested against how it actually behaves in the real world?**
Agenomics gives agents a structured genome, calculates Trust and Compatibility scores, monitors behavioral drift, and — starting with v0.7.0 — provides persistent evidence collection through EvidenceStore.
But there is an important part we cannot manufacture:
**real-world data.**
We are looking for developers running real AI agents who are willing to connect them to Agenomics, collect observations and incidents, and help test whether trust scores actually correspond to production behavior.
No user conversations need to be sent to Agenomics.
Your agent is the experiment.
Your data is the evidence.
Agenomics is the methodology.
I'd love to hear what you think — especially what signals you would use to measure whether an AI agent is actually trustworthy.