Google Cloud TPU pricing is no longer a single “TPU price.” Cost depends on TPU generation, region, purchasing model, number of chips, and the accelerator time required to complete a workload at the required quality and service level.
For enterprise AI teams, the right question is not simply “Is a TPU cheaper than a GPU?” A better question is:
What is the cost of completing the same AI workload at the same quality, latency, and reliability target on each accelerator?
That distinction matters because TPU pricing is commonly expressed per chip-hour, while GPU services may be packaged as complete virtual machines that include GPUs, CPUs, memory, storage, and networking.
Representative TPU pricing
Published U.S. on-demand rates in 2026 illustrate the range, alongside the discounted purchasing modes available for the same hardware. All figures are per chip-hour in USD.
| TPU generation | On demand | Flex-start | 1-year commitment | 3-year commitment |
|---|---|---|---|---|
| TPU v5e | $1.20 | $0.60 | $0.84 | $0.54 |
| Trillium (v6e) | $2.70 | $1.35 | $1.89 | $1.22 |
| TPU v5p | $4.20 | $2.10 | $2.94 | $1.89 |
| Ironwood (v7) | $12.00 | $6.00 | $8.40 | $5.40 |
Baseline U.S. regions: v5e and Ironwood in us-central1 (Iowa), Trillium in us-east1 and us-east5, v5p in us-east5 and us-east1. Flex-start refers to Dynamic Workload Scheduler flex-start capacity. Rates outside these regions run higher : Trillium is $2.97 in Amsterdam and $3.24 in Tokyo, and Ironwood is $13.20 in London. Spot pricing is dynamic and quoted separately. Always confirm against the live price list before modelling.
The spread inside each row is the point. The same chip can differ by more than 2x in unit price depending on how it is purchased, which frequently outweighs the difference between adjacent hardware generations.
The TPU generations in practice
TPU v5e
TPU v5e is a cost-oriented option for workloads where efficiency matters more than maximum accelerator performance. It can be attractive for batch training, fine-tuning, embeddings, and inference workloads that map efficiently to TPU infrastructure. At $0.54 per chip-hour on a three-year commitment, it is the cheapest published TPU capacity Google sells.
Trillium (v6e)
Trillium illustrates why procurement strategy matters almost as much as chip selection. The same hardware ranges from $2.70 per chip-hour on demand down to $1.22 on a three-year commitment : the same silicon, less than half the unit price.
TPU v5p
TPU v5p targets larger training workloads. A higher hourly rate does not automatically mean higher workload cost. If it reaches the same training objective in fewer chip-hours, the total run can still be cheaper than a nominally cheaper accelerator that takes longer.
Ironwood (v7)
Ironwood is the highest-performance TPU option and the most expensive per hour. Its pricing makes the utilization question even more important: underused high-end accelerators can have poor economics even when they are technically faster. An idle Ironwood chip burns roughly ten times the hourly cost of an idle v5e.
Calculate TPU workload cost
A simple capacity formula is:
For example, 64 chips for 10 hours of Trillium at the on-demand rate:
If flex-start scheduling reduces the rate to $1.35, the same capacity time costs:
But capacity price is not the same as workload economics. If another accelerator takes longer to reach the same target, hourly price alone is misleading.
TPU price is not cost per token
AI teams often compare accelerators with peak FLOPS, memory, bandwidth, tokens per second, or hourly rate. These are engineering metrics, not business outcomes. Substitute the metric that matches the workload:
Inference
Cost per one million successful tokens at the required latency and quality.
Training
Cost to reach the target model quality, not cost per hour of capacity.
Embeddings
Cost per million documents embedded, end to end.
Agentic applications
Cost per successful end-to-end task, including retries and failures.
This connects directly to how LLM token cost is modelled for production systems running large language models.
TPU vs GPU cost
A GPU VM and a TPU chip-hour are not directly comparable units. GPU configurations may include multiple accelerators, CPU, RAM, storage, and different networking characteristics. Benchmark the same workload on each platform and measure:
- Model quality.
- Time to completion.
- Accelerator utilization.
- Throughput.
- Latency.
- Failure and retry cost.
- Engineering and migration cost.
- Networking and storage overhead.
When TPUs can be attractive
- Workloads map well to supported frameworks such as JAX and PyTorch/XLA.
- Utilization is high and capacity is not left idle.
- Matrix-heavy computation dominates the workload profile.
- The workload benefits from Google’s TPU pod scaling architecture.
When GPUs can make more sense
- The organization already has a mature CUDA stack.
- The workload depends on GPU-specific libraries or kernels.
- Portability across cloud and on-prem environments is required.
- Migration cost would be substantial relative to the saving.
Evaluate all purchasing models
Do not model only on-demand capacity. Compare:
- On demand : hourly, based on actual usage; best for short experiments and benchmarks.
- Flexible scheduling : Dynamic Workload Scheduler flex-start, roughly half the on-demand rate across every current generation.
- Committed use : one- or three-year reservations billed monthly against reserved quota.
- Interruptible or Spot-style capacity : dynamic pricing for batch and fault-tolerant work.
Practical decision framework
Define the workload
Specify training or inference, model size, precision, batch size, context length, latency target, availability target, and expected volume.
Identify feasible accelerators
Do not benchmark hardware that cannot efficiently run the workload.
Benchmark performance
Measure the real model and application rather than relying only on vendor specifications.
Apply realistic pricing
Use the procurement mode the organization can actually buy, in the region it will actually run in.
Calculate outcome economics
Compare cost per successful inference, per million tokens, per training run, or to reach target quality.
Bottom line
The lowest hourly accelerator rate does not necessarily produce the lowest AI workload cost.
That turns accelerator selection from a hardware-specification debate into an engineering and economic optimization problem.