
Silene Bench
A primary care benchmark you can play yourself
4 followers
A primary care benchmark you can play yourself
4 followers
Can an AI agent run a primary care clinic over time? SileneBench is an experimental, playable benchmark exploring that question. Patients return, resources are limited, and rare diseases can hide among everyday cases. Run the clinic yourself or watch GPT-6 Astra make clinical and operational decisions. This hackathon demo is a first step toward a year-long simulation, with support for connecting your own agent harness planned.











Hey everyone! I’m Adrian, a researcher with a PhD background in machine learning for healthcare. My previous work includes using ML to support the diagnosis of rare diseases.
With SileneBench, I’m exploring a question: can an AI agent make good clinical and operational decisions while running a primary care practice over time?
Patients return, appointment slots are limited, and earlier decisions affect what happens next. I’m particularly interested in whether an agent can manage everyday care while picking up on rare diseases as evidence gradually emerges.
You can play the demo yourself or watch GPT-6 Astra run the clinic. I wanted humans to be able to play the same environment from the start, with the longer-term goal of comparing human and agent performance.
This is an early hackathon demo. The goal is a roughly year-long simulation, including chronic disease management and consequences that unfold over months. There’s still work ahead on calibration with doctors and evaluation. Support for connecting your own agent harness is planned, but isn’t available yet.
If you give it a try, I’d love to hear whether the simulation feels too easy and what would make it more challenging. Would you be interested in connecting your own agent harness to SileneBench? And does the idea of a visual benchmark, where you can watch an agent’s decisions play out or try it yourself, appeal to you?