o3 API pricing (August 2026)
Reasoning model — billed reasoning tokens count as output and can dominate the bill.
What it costs at real volumes
A chat-shaped workload: 800 input tokens and 300 output tokens per request, no caching, no batch.
Specifications
- Provider
- OpenAI
- Context window
- 200K tokens
- Max output
- 100K tokens
- Tier
- frontier
- Status
- GA
- Blended $/1M (3:1)
- $3.50
- Cache break-even hit rate
- 0% — caching always saves
- Prices verified
- 4 Aug 2026
Source: OpenAI pricing. Finitizer is not affiliated with OpenAI.
Similar models
Other gpt models
- GPT-5$-0.063
- GPT-5 mini-$2.81
- GPT-5 nano-$3.36
- GPT-4.1+$0
- GPT-4.1 mini-$2.80
- GPT-4.1 nano-$3.33
- GPT-4o+$0.875
- GPT-4o mini-$3.24
- o4-mini-$1.57
Closest prices elsewhere
- Gemini 2.5 Pro$-0.063
- Claude Sonnet 5+$0.500
- Grok 4.5$-0.500
- Claude Haiku 4.5-$1.50
Frequently asked questions
How much does o3 cost per 1M tokens?
o3 costs $2.00 per million input tokens and $8.00 per million output tokens at OpenAI list prices. Output is 4.0x the input rate, so the input:output ratio of your workload drives the real cost more than the headline price.
Does o3 support prompt caching?
Yes. Cached input is billed at $0.500 per million tokens, 4x cheaper than uncached input. There is no cache-write premium, so caching is never a net loss.
What is the batch discount for o3?
Requests submitted through the asynchronous batch endpoint are billed at 50% off both input and output rates, bringing input to $1.00 and output to $4.00 per million tokens.
What is o3's context window?
200K tokens of context with up to 100K output tokens per request. A larger context window raises the ceiling on what you can send, not the price per token — but filling it on every request is one of the most common causes of an unexpected bill.
Running o3 in production?
Finitizer TokenOps attributes real token spend to teams, features, and prompts — so you find out which workload moved the bill before the invoice does.