Token spend is the line item nobody forecasts until the invoice lands. Size a real prompt, compare models, project the month โ and check whether a flat plan beats paying per token.
This is a character-density approximation, not a tokenizer. Expect ยฑ10โ15%. For exact counts, use your provider's token-counting endpoint before you commit to a budget.
| Model | Context | $/1M in โ out | Per request | Per day | Per month |
|---|
Green marks the cheapest monthly total. Cached input is billed at 0.1ร the base rate; the write that populates a cache costs 1.25ร, so caching pays for itself from the second read onward โ which is why routing cheap work to a smaller model and caching a stable prefix are the two levers that actually move this number.
Heavy users routinely find an order-of-magnitude gap between metered API billing and a flat plan for identical work. The break-even figure is the request volume where the two meet โ below it you are overpaying for the plan, above it you are overpaying for tokens.
Seeded with published Anthropic list prices. Rates move โ edit any cell, or add the model you actually use, and the whole page recalculates. Nothing is sent anywhere.