Mistral Large 3 API pricing (August 2026)
Open-weights flagship. Mistral Medium 3.5 ($1.50 in / $7.50 out) is the premium hosted tier.
What it costs at real volumes
A chat-shaped workload: 800 input tokens and 300 output tokens per request, no caching, no batch.
Specifications
- Provider
- Mistral AI
- Context window
- 131K tokens
- Max output
- 8.2K tokens
- Tier
- frontier
- Status
- GA
- Blended $/1M (3:1)
- $0.750
- Cache break-even hit rate
- 0% — caching always saves
- Prices verified
- 4 Aug 2026
Source: Mistral AI pricing. Finitizer is not affiliated with Mistral AI.
Similar models
Other mistral models
- Mistral Small 4$-0.488
Closest prices elsewhere
- GPT-4.1 mini$-0.050
- GPT-5 mini$-0.063
- Gemini 2.5 Flash+$0.100
- Llama 3.3 70B (Groq)$-0.110
Frequently asked questions
How much does Mistral Large 3 cost per 1M tokens?
Mistral Large 3 costs $0.500 per million input tokens and $1.50 per million output tokens at Mistral AI list prices. Output is 3.0x the input rate, so the input:output ratio of your workload drives the real cost more than the headline price.
Does Mistral Large 3 support prompt caching?
Yes. Cached input is billed at $0.050 per million tokens, 10x cheaper than uncached input. There is no cache-write premium, so caching is never a net loss.
What is the batch discount for Mistral Large 3?
Requests submitted through the asynchronous batch endpoint are billed at 50% off both input and output rates, bringing input to $0.250 and output to $0.750 per million tokens.
What is Mistral Large 3's context window?
131K tokens of context with up to 8.2K output tokens per request. A larger context window raises the ceiling on what you can send, not the price per token — but filling it on every request is one of the most common causes of an unexpected bill.
Running Mistral Large 3 in production?
Finitizer TokenOps attributes real token spend to teams, features, and prompts — so you find out which workload moved the bill before the invoice does.