Price your own workload on this model

A rate per million tokens is not a budget. The calculator takes your request volume, prompt size and cache hit rate and gives you a monthly figure.

Mistral Large 3 API pricing (September 2026)

Prices verified 11 Sep 2026

Open-weights flagship. Mistral Medium 3.5 ($1.50 in / $7.50 out) is the premium hosted tier.

Input
$0.500
per 1M tokens
Output
$1.50
per 1M tokens
Cached input
$0.050
per 1M tokens
Batch discount
50% off
async batch endpoint

What it costs at real volumes

A chat-shaped workload: 800 input tokens and 300 output tokens per request, no caching, no batch.

VolumePer requestPer monthPer year
10K requests/day$0.00085$258.74$3,104.88Adjust →
100K requests/day$0.00085$2,587.40$31.0KAdjust →
1M requests/day$0.00085$25.9K$310.5KAdjust →

Specifications

Provider
Mistral AI
Context window
131K tokens
Max output
8.2K tokens
Tier
frontier
Status
GA
Blended $/1M (3:1)
$0.750
Cache break-even hit rate
0%, caching always saves
Prices verified
11 Sep 2026

Source: Mistral AI pricing. Finitizer is not affiliated with Mistral AI.

Similar models

Other mistral models

Closest prices elsewhere

Frequently asked questions

How much does Mistral Large 3 cost per 1M tokens?

Mistral Large 3 costs $0.500 per million input tokens and $1.50 per million output tokens at Mistral AI list prices. Output is 3.0x the input rate, so the input:output ratio of your workload drives the real cost more than the headline price.

Does Mistral Large 3 support prompt caching?

Yes. Cached input is billed at $0.050 per million tokens, 10x cheaper than uncached input. There is no cache-write premium, so caching is never a net loss.

What is the batch discount for Mistral Large 3?

Requests submitted through the asynchronous batch endpoint are billed at 50% off both input and output rates, bringing input to $0.250 and output to $0.750 per million tokens.

What is Mistral Large 3's context window?

131K tokens of context with up to 8.2K output tokens per request. A larger context window raises the ceiling on what you can send, not the price per token, but filling it on every request is one of the most common causes of an unexpected bill.

Running Mistral Large 3 in production?

Finitizer TokenOps attributes real token spend to teams, features, and prompts, so you find out which workload moved the bill before the invoice does.