AI API cost calculator
Put in your traffic, get your monthly bill across every major model. Everything runs in your browser — no signup, no data leaves the page.
| Model | In / Out per 1M | Monthly | vs cheapest |
|---|---|---|---|
| GPT-4o miniOpenAI | $0.15 / $0.6 | $2.40 | — |
| DeepSeek ChatDeepSeek | $0.27 / $1.1 | $4.35 | 1.8× |
| Gemini FlashGoogle | $0.3 / $2.5 | $6.75 | 2.8× |
| Llama 70B (hosted)Together / Groq | $0.6 / $0.7 | $7.05 | 2.9× |
| Claude Haiku 4.5Anthropic | $1 / $5 | $17.50 | 7.3× |
| Gemini ProGoogle | $1.25 / $10 | $27.50 | 11.5× |
| Mistral LargeMistral | $2 / $6 | $29.00 | 12.1× |
| GPT-4oOpenAI | $2.5 / $10 | $40.00 | 16.7× |
| Claude Sonnet 5Anthropic | $3 / $15 | $52.50 | 21.9× |
| Claude Opus 5Anthropic | $15 / $75 | $263 | 109.4× |
Prices last checked 2026-08-13, in USD per million tokens, standard API rates. Batch processing and prompt caching can cut input costs substantially and aren't modelled here — neither is the free tier some providers offer. Confirm on the vendor's pricing page before you budget on this.
How to estimate your tokens
A token is roughly four characters of English, so about 750 words per 1,000 tokens. Code and non-English text run token-heavier — budget 30–50% more than the word count suggests.
Input is everything you send: system prompt, conversation history, retrieved documents, tool definitions. This is where costs hide. A chatbot that resends a 20-message history on every turn pays for that history every single turn.
Output is what the model writes back. It costs three to five times more per token than input on most providers, so capping response length is usually the fastest saving available.
Three ways to cut the bill
- Prompt caching. If your system prompt or documents repeat across requests, most providers will cache them at a large discount. On a RAG app this is often the single biggest saving.
- Route by difficulty. Send the easy 80% to a small model and escalate only what needs it. The table above shows the gap is frequently 50× or more, not 2×.
- Batch anything not user-facing. Classification, enrichment and summarisation jobs usually run at roughly half price if you can wait.
Comparing the providers themselves?
The directory covers the inference platforms behind these models — and what each is actually good at.