Llama 3.3 70B (Groq) API pricing (August 2026)
What it costs at real volumes
A chat-shaped workload: 800 input tokens and 300 output tokens per request, no caching, no batch.
Specifications
- Provider
- Groq
- Context window
- 128K tokens
- Max output
- 33K tokens
- Tier
- mid
- Status
- GA
- Blended $/1M (3:1)
- $0.640
- Cache break-even hit rate
- No caching
- Prices verified
- 4 Aug 2026
Source: Groq pricing. Finitizer is not affiliated with Groq.
Similar models
Other llama models
- Llama 3.1 8B (Groq)$-0.583
Closest prices elsewhere
- GPT-5 mini+$0.047
- GPT-4.1 mini+$0.060
- DeepSeek V4 Pro$-0.096
- Mistral Large 3+$0.110
Frequently asked questions
How much does Llama 3.3 70B (Groq) cost per 1M tokens?
Llama 3.3 70B (Groq) costs $0.590 per million input tokens and $0.790 per million output tokens at Groq list prices. Output is 1.3x the input rate, so the input:output ratio of your workload drives the real cost more than the headline price.
Does Llama 3.3 70B (Groq) support prompt caching?
Llama 3.3 70B (Groq) has no published prompt-caching rate, so repeated context is billed at the full input rate on every request.
What is the batch discount for Llama 3.3 70B (Groq)?
Requests submitted through the asynchronous batch endpoint are billed at 50% off both input and output rates, bringing input to $0.295 and output to $0.395 per million tokens.
What is Llama 3.3 70B (Groq)'s context window?
128K tokens of context with up to 33K output tokens per request. A larger context window raises the ceiling on what you can send, not the price per token — but filling it on every request is one of the most common causes of an unexpected bill.
Running Llama 3.3 70B (Groq) in production?
Finitizer TokenOps attributes real token spend to teams, features, and prompts — so you find out which workload moved the bill before the invoice does.