Inferock Bench - An independent receipt for every LLM API call

by•
Inferock-bench is a local proxy that sits between your app and OpenAI, Anthropic, Gemini, or OpenRouter shaped calls. It captures per-call token usage, failures, and retries, then generates an independent receipt showing what you were billed and how much you're actually overpaying for.

Add a comment

Replies

Best

Hey PH

:wave:

We built inferock-bench because we kept paying for AI answers that died mid sentence, and nobody could tell us where the money went.

Providers give you totals. They don't give you the per call receipt you'd need to prove which answer broke, which retry ran, or which token count changed. The company that charges you also decides what counts as a failure and keeps the only detailed records.

inferock-bench runs locally as a proxy in front of OpenAI, Anthropic, Gemini, OpenRouter shaped calls. Point your existing SDK at it (change two settings: apiKey and baseURL), and it captures every call as an independent, per call record. Your provider key never touches our servers, it's used locally only, attached to provider requests.

What it catches:
- Answers cut off mid stream that still got billed
- Empty replies with billed tokens attached
- Token counts that don't match visible output
- Retries that may have silently doubled a charge
- Cache discounts you may be missing on your invoice

Every run reports a receipt: spend observed, bill-bounded money loss, time loss, and a separate "invoice-check exposure" line that never gets summed into money loss, because we don't want a louder headline at the cost of a weaker claim.

Run it in about a minute: npx inferock-bench

Its open source (FSL-1.1-Apache-2.0, converts to Apache-2.0 in 2 years).

Question for this community: has anyone here actually disputed an AI provider bill and gotten a credit? What worked?

Getting billed for mid sentence drop offs and silent retries is such a massive headache for devs having local per call receipts proxying OpenAI/Anthropic/Gemini is brilliant. To answer your question: Disputing AI bills with providers usually gets ignored unless you have concrete per-request log receipts like this. This tool gives teams actual leverage!

   Appreciate you actually answering the question, that matches what we kept hearing from others too. Without per request records the conversation with a provider ends pretty fast. You walk in with evidence instead of a feeling.

Exactly data always wins over estimates when dealing with provider support! Love the angle you guys took here.

 Thank you!

:raised_hands:

Genuinely hoping it saves a few of those support conversations.

Congrats on the launch . Good find .

Question regarding accuracy, is it possible it might not be able to distinguish a genuinely billable provider failure from valid hidden token usage, such as reasoning, refusal, cache, or tool-call tokens?

 thank you! Anything that could be legit hidden usage like reasoning, cache, or tool tokens doesn't count as loss, it just gets flagged for you to check. Only what we can prove makes the loss number.

The failed-calls and retries breakdown is the part I'd use first. One case I keep hitting might not show up there: a tool call that returns 200 with a silently corrupted value. I measured this on Anthropic models, 40 calls, none flagged the value was wrong, so it bills as a clean success and the retry logic never fires. Can the receipt catch a call that looked fine but wasn't? Or is that out of scope by design?

that one's mostly out of scope. The receipt proves what you were billed and what actually came back, so truncations, empty replies, token mismatches, retries. A well-formed 200 with a wrong value inside would pass as a clean call, we simply can't tell from the outside that the value was wrong.

  Thanks, that's a clear answer and the right boundary to hold. The corrupted-200 case really does belong to a different layer. It needs a schema or an oracle on the payload and only the caller knows what the value should be. What you're proving from the outside is exactly the part I can't check myself today, so I'd keep it that sharp. Following the launch.

This solves a problem I didn't know I could solve, I always assumed API billing discrepancies were just the cost of doing business. Turns out I was wrong.

Thanks  , honestly, this is exactly what we told ourselves for months, until we actually looked at the per call receipts and couldn't unsee it. Really glad this clicked for you. Pls run it on your own traffic, come back and tell us what your receipts show, genuinely curious!

The independent receipt idea is genuinely useful. LLM costs are still surprisingly hard to audit at the individual request level.

 thank you! That's exactly the gap we kept running into, you can see the total, but never the audit at the individual request level. Glad it resonates.

This is relevant to a problem I've had for months. I run agents that call out to multiple models depending on the task complexity and every so often the bill jumps in a way I can't explain from usage alone. If this can pinpoint whether that's failed calls, redundant retries or just legitimate scaling, I'd finally have an answer instead of a guess.

 Multi model agent setups are where this gets most interesting. Every call gets its own record whichever provider it went to, so a jump breaks down into failed calls, retries, etc. Would love to hear what you find.

I like the positioning here, it's not trying to replace the provider's billing system, just verify it. That's a smarter pitch than "cost optimization," which every tool claims.

 Thank you, that's how we think about it as well. Your bill already exists, you should just be able to check it.

I've been burned by silent retries inflating my OpenAI bill before. Having an independent receipt for that would've saved me a painful invoice conversation.

 Ugh, the invoice conversation where you both know something's off. Sorry you've lived that one. Next time at least you'd be the one holding the itemized version.

The independent receipt idea is really practical. It's nice to have a way to verify token usage instead of relying only on provider dashboards.

 Thank you! It always felt a little odd that the only record of what we bought came from the company selling it, so now you can keep your own. Glad it's practical!

Love that this sits as a local proxy instead of asking me to route traffic through another cloud service. Keeps my API keys and data where they belong.

 everything stays on your machine, we never see it.

123
Next