Claude Sonnet 5 API pricing (August 2026)

Prices verified 4 Aug 2026

Introductory rate through 31 Aug 2026 — rises to $3 in / $15 out on 1 Sep 2026.

Input
$2.00
per 1M tokens
Output
$10.00
per 1M tokens
Cached input
$0.200
per 1M tokens
Batch discount
50% off
async batch endpoint

What it costs at real volumes

A chat-shaped workload: 800 input tokens and 300 output tokens per request, no caching, no batch.

VolumePer requestPer monthPer year
10K requests/day$0.00460$1,400.24$16.8KAdjust →
100K requests/day$0.00460$14.0K$168.0KAdjust →
1M requests/day$0.00460$140.0K$1.68MAdjust →

Specifications

Provider
Anthropic Claude
Context window
1.0M tokens
Max output
128K tokens
Tier
mid
Status
GA
Blended $/1M (3:1)
$4.00
Cache break-even hit rate
22%
Prices verified
4 Aug 2026

Source: Anthropic Claude pricing. Finitizer is not affiliated with Anthropic Claude.

Similar models

Closest prices elsewhere

Frequently asked questions

How much does Claude Sonnet 5 cost per 1M tokens?

Claude Sonnet 5 costs $2.00 per million input tokens and $10.00 per million output tokens at Anthropic Claude list prices. Output is 5.0x the input rate, so the input:output ratio of your workload drives the real cost more than the headline price.

Does Claude Sonnet 5 support prompt caching?

Yes. Cached input is billed at $0.200 per million tokens, 10x cheaper than uncached input. Cache writes carry a premium, so caching only pays above roughly a 22% hit rate.

What is the batch discount for Claude Sonnet 5?

Requests submitted through the asynchronous batch endpoint are billed at 50% off both input and output rates, bringing input to $1.00 and output to $5.00 per million tokens.

What is Claude Sonnet 5's context window?

1.0M tokens of context with up to 128K output tokens per request. A larger context window raises the ceiling on what you can send, not the price per token — but filling it on every request is one of the most common causes of an unexpected bill.

Running Claude Sonnet 5 in production?

Finitizer TokenOps attributes real token spend to teams, features, and prompts — so you find out which workload moved the bill before the invoice does.