o4-mini API pricing (August 2026)
Reasoning model — billed reasoning tokens count as output.
What it costs at real volumes
A chat-shaped workload: 800 input tokens and 300 output tokens per request, no caching, no batch.
Specifications
- Provider
- OpenAI
- Context window
- 200K tokens
- Max output
- 100K tokens
- Tier
- mid
- Status
- GA
- Blended $/1M (3:1)
- $1.93
- Cache break-even hit rate
- 0% — caching always saves
- Prices verified
- 4 Aug 2026
Source: OpenAI pricing. Finitizer is not affiliated with OpenAI.
Similar models
Other gpt models
- GPT-5+$1.51
- GPT-5 mini-$1.24
- GPT-5 nano-$1.79
- GPT-4.1+$1.57
- GPT-4.1 mini-$1.23
- GPT-4.1 nano-$1.75
- GPT-4o+$2.45
- GPT-4o mini-$1.66
- o3+$1.57
Closest prices elsewhere
- Claude Haiku 4.5+$0.075
- Claude Haiku 3.5$-0.325
- Grok 4.3$-0.363
- Amazon Nova Pro$-0.525
Frequently asked questions
How much does o4-mini cost per 1M tokens?
o4-mini costs $1.10 per million input tokens and $4.40 per million output tokens at OpenAI list prices. Output is 4.0x the input rate, so the input:output ratio of your workload drives the real cost more than the headline price.
Does o4-mini support prompt caching?
Yes. Cached input is billed at $0.275 per million tokens, 4x cheaper than uncached input. There is no cache-write premium, so caching is never a net loss.
What is the batch discount for o4-mini?
Requests submitted through the asynchronous batch endpoint are billed at 50% off both input and output rates, bringing input to $0.550 and output to $2.20 per million tokens.
What is o4-mini's context window?
200K tokens of context with up to 100K output tokens per request. A larger context window raises the ceiling on what you can send, not the price per token — but filling it on every request is one of the most common causes of an unexpected bill.
Running o4-mini in production?
Finitizer TokenOps attributes real token spend to teams, features, and prompts — so you find out which workload moved the bill before the invoice does.