Your model changed and nobody told you.
There's this weird thing with LLMs: they can quietly start behaving differently and you'd never know. A daily Claude Opus user I read about noticed one day that it suddenly agreed with everything he said — sycophancy that crept in without a version note, a changelog, anything. If you've got production workflows sitting on top of a model, that kind of silent drift is a real reliability hole.
The idea we've been kicking around is basically a health dashboard for your models. Run a fixed set of probe prompts every day, watch for tone drift, refusal-rate changes, sycophancy, output pattern shifts, and flag it when the behavior wanders off its baseline. Less "trust me it works" and more "here's exactly what moved and when."
It fits how we think about AI business opportunities at SoloVault — spotting the gaps that only show up once people actually depend on these tools.
If you run an LLM in production, how do you currently notice when it's changed?

Replies