Some lessons only land when you have to retire something you were proud of.
What we shipped: a fully autonomous agent that could approve low-risk customer service refunds without human review. Cost-justified, well-scoped, passed every test.
Genuine question because the answer keeps changing.
Six months ago GPT-4o was the default for most teams. Then Claude 3.5 Sonnet started winning reasoning-heavy use cases. Now Llama 3 is showing up in production for cost-sensitive or on-prem deployments.