Input and output tokens for hosted models.
Token charges for hosted generative models on Vertex AI, billed separately for input and output. Output tokens cost substantially more than input tokens, and agentic workloads invert the usual intuition by generating enormous input volume as they re-read their own context on every step.
Billed: Per 1,000 or 1,000,000 tokens, priced separately for input and output. Output is typically several times the input rate.