Gemini 2.5 Flash-Lite API pricing (August 2026)

Prices verified 4 Aug 2026
Input
$0.100
per 1M tokens
Output
$0.400
per 1M tokens
Cached input
$0.010
per 1M tokens
Batch discount
50% off
async batch endpoint

What it costs at real volumes

A chat-shaped workload: 800 input tokens and 300 output tokens per request, no caching, no batch.

VolumePer requestPer monthPer year
10K requests/day$0.00020$60.88$730.56Adjust →
100K requests/day$0.00020$608.80$7,305.60Adjust →
1M requests/day$0.00020$6,088.00$73.1KAdjust →

Specifications

Provider
Google Gemini
Context window
1.0M tokens
Max output
66K tokens
Tier
small
Status
GA
Blended $/1M (3:1)
$0.175
Cache break-even hit rate
0% — caching always saves
Prices verified
4 Aug 2026

Source: Google Gemini pricing. Finitizer is not affiliated with Google Gemini.

Similar models

Other gemini models

Closest prices elsewhere

Frequently asked questions

How much does Gemini 2.5 Flash-Lite cost per 1M tokens?

Gemini 2.5 Flash-Lite costs $0.100 per million input tokens and $0.400 per million output tokens at Google Gemini list prices. Output is 4.0x the input rate, so the input:output ratio of your workload drives the real cost more than the headline price.

Does Gemini 2.5 Flash-Lite support prompt caching?

Yes. Cached input is billed at $0.010 per million tokens, 10x cheaper than uncached input. There is no cache-write premium, so caching is never a net loss.

What is the batch discount for Gemini 2.5 Flash-Lite?

Requests submitted through the asynchronous batch endpoint are billed at 50% off both input and output rates, bringing input to $0.050 and output to $0.200 per million tokens.

What is Gemini 2.5 Flash-Lite's context window?

1.0M tokens of context with up to 66K output tokens per request. A larger context window raises the ceiling on what you can send, not the price per token — but filling it on every request is one of the most common causes of an unexpected bill.

Running Gemini 2.5 Flash-Lite in production?

Finitizer TokenOps attributes real token spend to teams, features, and prompts — so you find out which workload moved the bill before the invoice does.