Can an AI agent run a primary care clinic over time? SileneBench is an experimental, playable benchmark exploring that question. Patients return, resources are limited, and rare diseases can hide among everyday cases. Run the clinic yourself or watch GPT-6 Astra make clinical and operational decisions. This hackathon demo is a first step toward a year-long simulation, with support for connecting your own agent harness planned.
How did Astra change the scope or ambition of what you built?
Maker
Building a reliable benchmark is a substantial research challenge. Astra expanded my ambition by making it feasible to turn that benchmark into a visual, playable application within a challenge.
I brought my research background in AI-assisted rare disease diagnosis and the questions I wanted to explore. Astra helped me design the benchmark, built almost the entire application, and processed and integrated the visual assets into a clinic that people can interact with. Being able to move between research design, implementation, and visual work with the same model changed what I felt I could attempt.
The result is an early environment where you can run the practice yourself or watch an AI agent make decisions. Establishing the benchmark’s reliability still requires calibration and evaluation, but Astra helped me make the underlying idea something people can actually play and inspect.
Report
Maker
📌
Hey everyone! I’m Adrian, a researcher with a PhD background in machine learning for healthcare. My previous work includes using ML to support the diagnosis of rare diseases.
With SileneBench, I’m exploring a question: can an AI agent make good clinical and operational decisions while running a primary care practice over time?
Patients return, appointment slots are limited, and earlier decisions affect what happens next. I’m particularly interested in whether an agent can manage everyday care while picking up on rare diseases as evidence gradually emerges.
You can play the demo yourself or watch GPT-6 Astra run the clinic. I wanted humans to be able to play the same environment from the start, with the longer-term goal of comparing human and agent performance.
This is an early hackathon demo. The goal is a roughly year-long simulation, including chronic disease management and consequences that unfold over months. There’s still work ahead on calibration with doctors and evaluation. Support for connecting your own agent harness is planned, but isn’t available yet.
If you give it a try, I’d love to hear whether the simulation feels too easy and what would make it more challenging. Would you be interested in connecting your own agent harness to SileneBench? And does the idea of a visual benchmark, where you can watch an agent’s decisions play out or try it yourself, appeal to you?
Hey everyone! I’m Adrian, a researcher with a PhD background in machine learning for healthcare. My previous work includes using ML to support the diagnosis of rare diseases.
With SileneBench, I’m exploring a question: can an AI agent make good clinical and operational decisions while running a primary care practice over time?
Patients return, appointment slots are limited, and earlier decisions affect what happens next. I’m particularly interested in whether an agent can manage everyday care while picking up on rare diseases as evidence gradually emerges.
You can play the demo yourself or watch GPT-6 Astra run the clinic. I wanted humans to be able to play the same environment from the start, with the longer-term goal of comparing human and agent performance.
This is an early hackathon demo. The goal is a roughly year-long simulation, including chronic disease management and consequences that unfold over months. There’s still work ahead on calibration with doctors and evaluation. Support for connecting your own agent harness is planned, but isn’t available yet.
If you give it a try, I’d love to hear whether the simulation feels too easy and what would make it more challenging. Would you be interested in connecting your own agent harness to SileneBench? And does the idea of a visual benchmark, where you can watch an agent’s decisions play out or try it yourself, appeal to you?