Inferock Bench - An independent receipt for every LLM API call

by•
Inferock-bench is a local proxy that sits between your app and OpenAI, Anthropic, Gemini, or OpenRouter shaped calls. It captures per-call token usage, failures, and retries, then generates an independent receipt showing what you were billed and how much you're actually overpaying for.

Add a comment

Replies

Best

The retry tracking caught my attention. I’ve seen failed requests become surprisingly expensive, so being able to trace each call would be useful to me.

 Failed requests being the expensive ones is such a strange truth of this stuff, you pay and get nothing usable back. Hope it earns a spot in your setup!

changing the baseURL and apiKey instead of rewriting the application is a nice touch. makes this much easier to test on an existing project.

 that was a rule we set for ourselves, if trying it takes more than a minute nobody will ever find out what their calls really cost. Two settings in, and your app never knows the difference. Would be curious how it goes on your existing project.

Thanks for hunting us. Feel free to ask if anybody has any questions about Inferock Bench.

The retry tracking caught my eye. silent retries are probably one of the easiest ways for API costs to creep up without anyone noticing.

 Right, retries are the sneakiest of the bunch: the answer still arrives, everything looks fine, and the bill just grows. Making those visible was one of the first things we wanted for ourselves.

What I really want to know is how this handles historical data. Can I feed it a month of past logs and get a retroactive receipt or is it strictly forward-looking from install? I ask because the overspending I'm most curious about already happened and I'd love a way to audit it after the fact.

 really good question. Today it works by sitting in front of your live traffic, so the receipts start from the moment you point your SDK at it, it can't vouch for calls it never saw.

I like that Inferock Bench focuses on evidence rather than simply showing another dashboard. Having a separate record of every API call could make unexpected billing much easier to investigate.

 thank you! Dashboards summarize, and summaries are where the weird stuff goes to hide. We wanted something closer to a paper trail, boring on purpose, there for the day you need to investigate.

The two setting setup makes this feel unusally easy to try.

 thank you! Hopefully the receipts earn it the permanent spot.

The bill dispute question is interesting. have you personally managed to get a provider to credit a charge after showing them one of these per call receipts or is that still something you are testing?

 honest answer: no credit to brag about yet, that's exactly why we threw the question to the community. What we can stand behind today is the receipt itself, knowing which call failed and what it cost, instead of arguing from a monthly total.

Does anyone know of any tools that can work with web-based logins (claud.ai etc)

 we don't touch that layer. inferock-bench works where there's an API key and a baseURL to point somewhere, and web app subscriptions don't expose the per call detail we'd need. If someone has cracked that curious to see it too.

auditing your own inference bill is such an obvious gap 🔍 nice one. seeing much overbilling in the wild yet?

it really is one of those gaps hiding in plain sight. And yes, we see it, and it scales with your spend, every cut off answer and retry gets billed like a success, and at production volume that's money leaking with zero visibility. What bothers us most is that without your own records you can't even know how much it's costing you.