When your agents get complex
Simulation-based AI agent testing and evaluation that turns unpredictable agents into reliable production systems.
This is the 4th launch from LangWatch. View more
Claude Code usage tracking by LangWatch
Launched this week
Track Claude Code usage: cost, cache, session replay. Run `npx langwatch claude` once. Every session gets cost with cache reads/writes as separate token classes, every bash and MCP call as a span, theoretical vs billed for your Max plan, and a full terminal replay in the UI. Works for Codex too.









Free
Launch Team




I might have missed this, but the cache-hit view is the part I’d check first. Claude Code costs feel random until you can separate cache misses from actual model usage.
GrowMeOrganic
Most observability tools focus on applications. This one focuses on the developers building them. Good to have LangWatch back on PH.
LangWatch
@iamanantgupta Thanks! It has been some time since our latest launch and indeed, this was just so much value to our own team that we had to launch it here again to other developers!
Its a much broader trend we see moving from building custom agents/apps towards anyone building with claude, co-pilot, codex solutions, so broader view / visibilty is required to keep control! More to come!
I'm bookmarking this. Comparing theoretical cost against what I'm actually billed on Max is something I've wanted for a while.
LangWatch
@ramish_saje go try it once you've found some time to track it, it will only take. you acouple of minutes!
LangWatch
@hamza_afzal_butt that's indeed the goal, so additionally to the current usage insights for the users, we do provide Admin dashboards as well to view the spend accross teams, projects and so on.
I can imagine engineering managers using this to better understand AI adoption across teams. Good launch!!
LangWatch
@zerotox yes so this part of the launch is pure visibility for dev's but we do give Admin dashboards and insights as well for engineering leaders managing budgets, for prediction and optimizing spend accross
Looking forward to seeing historical trends showing how coding workflows evolve over months. :D
LangWatch
@ragsyme exactly! We all have a feeling that the token expending per PR has been increasing, the cost per token too, sometimes "cheaper" models but actually spending double the tokens and so on.
It will be great to actually measure that effect for your team and having all the data behind it to read back and optimize
Lancepilot
LangWatch
@priyankamandal interesting use-case!