We spent the last few days trying to break our own AI coach
We gave APEX conflicting follow-ups, new restrictions, old workouts, and the classic:
“Make it harder.”
The interesting part wasn’t making the workout harder.
It was making sure the AI didn’t forget what it had already been told.
A harder request shouldn’t erase a shoulder restriction.
A new limitation shouldn’t let an old workout sneak back in.
And sometimes the correct AI response is simply: don’t generate the workout.
That sounds obvious.
It wasn’t.
We found several edge cases where the AI technically gave a “good” answer - but the system logic underneath wasn’t good enough.
So we rebuilt those paths.
Still a lot to improve, but this is probably the part of building APEX I find most interesting:
making AI know when not to be helpful.
Curious how other builders test this kind of failure in their products.
Replies