Tool · 01
Inference cost calculator
Estimate the monthly cost of an inference workload.
The workload
Override the rates
Using custom pricing? Enter your actual rates.
Self-host assumptions
Utilization is the number optimistic models fake. Measure it, or assume 30–50%.
Monthly
$5,000
output is 50% of it · $0.0200 per call
input 2,000 tokens = $2,500
output 400 tokens = $2,500
Cache break-even
2 requests
Reads cost 0.1×. The write costs 1.25×. You need this many uses of the same prefix before caching pays off.
Batch would save
$2,500 / mo
Flat 50% off input and output for work that tolerates a queue.
Same workload, other tiers
| gpt-5.6-luna | $220 | −96% |
| gpt-5.4-nano | $225 | −96% |
| deepseek-v4-flash | $352 | −93% |
| gemini-3.5-flash-lite | $400 | −92% |
| mistral-large-3 | $400 | −92% |
| claude-opus-5 · | $5,000 | — |
Cost is only half the decision. Run an eval before moving traffic.
Self-host break-even
Self-hosting wins above 175M tokens/month. Capacity ceiling is 828M/month.
$1,460/mo fixed, $1.76/MTok against a blended API rate of $8.33/MTok. Based on the GPU cost, utilization, and throughput assumptions on the left.
Have a real workload?
Send me the workload and the bill. I’ll dig into the engineering.
Look at my bill ↗