Claude Sonnet 5 vs GPT-5

The default choice on each side for general production work.

Prices verified 4 Aug 2026
Input / 1M$2.00
Output / 1M$10.00
Cached input / 1M$0.200
Batch50% off
Context1.0M tokens
Blended (3:1)$4.00
Input / 1M$1.25
Output / 1M$10.00
Cached input / 1M$0.125
Batch50% off
Context400K tokens
Blended (3:1)$3.44
Break-even. These two swap places at an output share of 100% of total tokens. Below that, GPT-5 is cheaper; above it, the other one is. Output-heavy work — code generation, long-form writing, reasoning models that bill their thinking — sits on the far side of that line more often than teams expect.

Monthly cost by workload shape

The same volume costs very different amounts depending on the shape of the traffic. No caching or batch applied.

WorkloadClaude Sonnet 5GPT-5Difference
Customer support chatbot
50K/day · 800 in / 300 out
$7,001.20$6,088.00GPT-5saves $913.20Model it →
RAG / search answers
20K/day · 3000 in / 400 out
$6,088.00$4,718.20GPT-5saves $1,369.80Model it →
Agent / tool-use loop
8.0K/day · 2500 in / 600 out
$2,678.72$2,222.12GPT-5saves $456.60Model it →
Batch summarization
100K/day · 5000 in / 500 out
$45.7K$34.2KGPT-5saves $11.4KModel it →
Code assistant
15K/day · 2000 in / 800 out
$5,479.20$4,794.30GPT-5saves $684.90Model it →

Frequently asked questions

Which is cheaper, Claude Sonnet 5 or GPT-5?

At a typical 3:1 input-to-output ratio, GPT-5 is cheaper: $3.44 blended per 1M tokens against $4.00. That flips once output exceeds 100% of your total tokens, because the two price input and output differently.

What is the price difference between Claude Sonnet 5 and GPT-5?

Claude Sonnet 5: $2.00 input, $10.00 output per 1M tokens. GPT-5: $1.25 input, $10.00 output. That is 1.6x on input and 1.0x on output.

Does prompt caching change the answer?

Claude Sonnet 5 caches input at $0.200/1M and GPT-5 caches at $0.125/1M. For input-heavy workloads with repeated context, caching can matter more than the base rate difference — model both in the calculator rather than comparing rate cards.

This page compares published prices only. It makes no claim about which model performs better on any task — that depends entirely on your evaluations.

Rate cards decide nothing on their own.

Most teams run several models at once. Finitizer TokenOps shows what each one actually costs you in production, by team and feature, so a routing decision is made on data rather than on a price page.