Claude Haiku 3.5 API pricing (August 2026)

Prices verified 4 Aug 2026

Retired 19 Feb 2026 — shown at its final price for historical comparison only.

Input
$0.800
per 1M tokens
Output
$4.00
per 1M tokens
Cached input
$0.080
per 1M tokens
Batch discount
50% off
async batch endpoint

What it costs at real volumes

A chat-shaped workload: 800 input tokens and 300 output tokens per request, no caching, no batch.

VolumePer requestPer monthPer year
10K requests/day$0.00184$560.10$6,721.15Adjust →
100K requests/day$0.00184$5,600.96$67.2KAdjust →
1M requests/day$0.00184$56.0K$672.1KAdjust →

Specifications

Provider
Anthropic Claude
Context window
200K tokens
Max output
8.2K tokens
Tier
small
Status
DEPRECATED
Blended $/1M (3:1)
$1.60
Cache break-even hit rate
22%
Prices verified
4 Aug 2026

Source: Anthropic Claude pricing. Finitizer is not affiliated with Anthropic Claude.

Similar models

Other claude models

Closest prices elsewhere

Frequently asked questions

How much does Claude Haiku 3.5 cost per 1M tokens?

Claude Haiku 3.5 costs $0.800 per million input tokens and $4.00 per million output tokens at Anthropic Claude list prices. Output is 5.0x the input rate, so the input:output ratio of your workload drives the real cost more than the headline price.

Does Claude Haiku 3.5 support prompt caching?

Yes. Cached input is billed at $0.080 per million tokens, 10x cheaper than uncached input. Cache writes carry a premium, so caching only pays above roughly a 22% hit rate.

What is the batch discount for Claude Haiku 3.5?

Requests submitted through the asynchronous batch endpoint are billed at 50% off both input and output rates, bringing input to $0.400 and output to $2.00 per million tokens.

What is Claude Haiku 3.5's context window?

200K tokens of context with up to 8.2K output tokens per request. A larger context window raises the ceiling on what you can send, not the price per token — but filling it on every request is one of the most common causes of an unexpected bill.

Running Claude Haiku 3.5 in production?

Finitizer TokenOps attributes real token spend to teams, features, and prompts — so you find out which workload moved the bill before the invoice does.