AI Infrastructure · TPU / GPU Economics

Google TPU Pricing in 2026: Cost per Hour, Discounts, and TPU vs GPU Economics

Google Cloud TPU pricing is no longer a single “TPU price.” Cost depends on TPU generation, region, purchasing model, chip count, and the accelerator time required to complete a workload at the required quality and service level. This page sets out the current published rates and a framework for comparing TPU and GPU economics on the basis that actually matters : cost per completed outcome.
TPU v5e
$1.20
per chip-hour, on demand
Trillium (v6e)
$2.70
per chip-hour, on demand
TPU v5p
$4.20
per chip-hour, on demand
Ironwood (v7)
$12.00
per chip-hour, on demand
Published list rates
Commitment & flex-start discounts
Cost per outcome, not per hour

Google Cloud TPU pricing is no longer a single “TPU price.” Cost depends on TPU generation, region, purchasing model, number of chips, and the accelerator time required to complete a workload at the required quality and service level.

For enterprise AI teams, the right question is not simply “Is a TPU cheaper than a GPU?” A better question is:

What is the cost of completing the same AI workload at the same quality, latency, and reliability target on each accelerator?

That distinction matters because TPU pricing is commonly expressed per chip-hour, while GPU services may be packaged as complete virtual machines that include GPUs, CPUs, memory, storage, and networking.

Representative TPU pricing

Published U.S. on-demand rates in 2026 illustrate the range, alongside the discounted purchasing modes available for the same hardware. All figures are per chip-hour in USD.

TPU generation On demand Flex-start 1-year commitment 3-year commitment
TPU v5e $1.20 $0.60 $0.84 $0.54
Trillium (v6e) $2.70 $1.35 $1.89 $1.22
TPU v5p $4.20 $2.10 $2.94 $1.89
Ironwood (v7) $12.00 $6.00 $8.40 $5.40

Baseline U.S. regions: v5e and Ironwood in us-central1 (Iowa), Trillium in us-east1 and us-east5, v5p in us-east5 and us-east1. Flex-start refers to Dynamic Workload Scheduler flex-start capacity. Rates outside these regions run higher : Trillium is $2.97 in Amsterdam and $3.24 in Tokyo, and Ironwood is $13.20 in London. Spot pricing is dynamic and quoted separately. Always confirm against the live price list before modelling.

The spread inside each row is the point. The same chip can differ by more than 2x in unit price depending on how it is purchased, which frequently outweighs the difference between adjacent hardware generations.

The TPU generations in practice

TPU v5e

TPU v5e is a cost-oriented option for workloads where efficiency matters more than maximum accelerator performance. It can be attractive for batch training, fine-tuning, embeddings, and inference workloads that map efficiently to TPU infrastructure. At $0.54 per chip-hour on a three-year commitment, it is the cheapest published TPU capacity Google sells.

Trillium (v6e)

Trillium illustrates why procurement strategy matters almost as much as chip selection. The same hardware ranges from $2.70 per chip-hour on demand down to $1.22 on a three-year commitment : the same silicon, less than half the unit price.

TPU v5p

TPU v5p targets larger training workloads. A higher hourly rate does not automatically mean higher workload cost. If it reaches the same training objective in fewer chip-hours, the total run can still be cheaper than a nominally cheaper accelerator that takes longer.

Ironwood (v7)

Ironwood is the highest-performance TPU option and the most expensive per hour. Its pricing makes the utilization question even more important: underused high-end accelerators can have poor economics even when they are technically faster. An idle Ironwood chip burns roughly ten times the hourly cost of an idle v5e.

Calculate TPU workload cost

A simple capacity formula is:

Capacity cost
TPU cost = number of chips × chip-hours × price per chip-hour

For example, 64 chips for 10 hours of Trillium at the on-demand rate:

On demand
64 × 10 × $2.70 = $1,728

If flex-start scheduling reduces the rate to $1.35, the same capacity time costs:

Flex-start
64 × 10 × $1.35 = $864

But capacity price is not the same as workload economics. If another accelerator takes longer to reach the same target, hourly price alone is misleading.

TPU price is not cost per token

AI teams often compare accelerators with peak FLOPS, memory, bandwidth, tokens per second, or hourly rate. These are engineering metrics, not business outcomes. Substitute the metric that matches the workload:

Inference

Cost per one million successful tokens at the required latency and quality.

Training

Cost to reach the target model quality, not cost per hour of capacity.

Embeddings

Cost per million documents embedded, end to end.

Agentic applications

Cost per successful end-to-end task, including retries and failures.

This connects directly to how LLM token cost is modelled for production systems running large language models.

TPU vs GPU cost

A GPU VM and a TPU chip-hour are not directly comparable units. GPU configurations may include multiple accelerators, CPU, RAM, storage, and different networking characteristics. Benchmark the same workload on each platform and measure:

  1. Model quality.
  2. Time to completion.
  3. Accelerator utilization.
  4. Throughput.
  5. Latency.
  6. Failure and retry cost.
  7. Engineering and migration cost.
  8. Networking and storage overhead.

When TPUs can be attractive

  • Workloads map well to supported frameworks such as JAX and PyTorch/XLA.
  • Utilization is high and capacity is not left idle.
  • Matrix-heavy computation dominates the workload profile.
  • The workload benefits from Google’s TPU pod scaling architecture.

When GPUs can make more sense

  • The organization already has a mature CUDA stack.
  • The workload depends on GPU-specific libraries or kernels.
  • Portability across cloud and on-prem environments is required.
  • Migration cost would be substantial relative to the saving.

Evaluate all purchasing models

Do not model only on-demand capacity. Compare:

  • On demand : hourly, based on actual usage; best for short experiments and benchmarks.
  • Flexible scheduling : Dynamic Workload Scheduler flex-start, roughly half the on-demand rate across every current generation.
  • Committed use : one- or three-year reservations billed monthly against reserved quota.
  • Interruptible or Spot-style capacity : dynamic pricing for batch and fault-tolerant work.
The purchasing model can change economics as much as the hardware generation. A three-year Ironwood commitment at $5.40 per chip-hour costs less than on-demand v5p at $4.20 once you account for the work completed per hour.

Practical decision framework

Define the workload

Specify training or inference, model size, precision, batch size, context length, latency target, availability target, and expected volume.

Identify feasible accelerators

Do not benchmark hardware that cannot efficiently run the workload.

Benchmark performance

Measure the real model and application rather than relying only on vendor specifications.

Apply realistic pricing

Use the procurement mode the organization can actually buy, in the region it will actually run in.

Calculate outcome economics

Compare cost per successful inference, per million tokens, per training run, or to reach target quality.

Bottom line

The lowest hourly accelerator rate does not necessarily produce the lowest AI workload cost.

Benchmark the workload, normalize quality and service levels, measure time-to-completion, and calculate cost per outcome.

That turns accelerator selection from a hardware-specification debate into an engineering and economic optimization problem.

Pricing source and currency. On-demand, flex-start and commitment rates on this page are taken from Google Cloud’s published Cloud TPU price list and were verified on 21 September 2026. Cloud accelerator pricing changes without notice and varies by region; treat these figures as a planning baseline and confirm current rates in the Google Cloud pricing documentation or with your account team before committing budget.

Measure which settings actually move cost, quality, and latency

Use KNOBS to measure which model, infrastructure, retrieval, prompt, and runtime settings materially affect cost, quality, latency, and business outcomes. Accelerator choice is one knob among many : and rarely the one with the largest effect until the rest are tuned.

Explore KNOBS →
FAQ

Google TPU pricing questions

Short answers to the questions teams ask most often when they start modelling TPU and GPU spend.

How much does a Google TPU cost per hour?

Pricing varies by TPU generation, region, and purchasing model. As of September 2026, published U.S. on-demand list rates are about $1.20 per chip-hour for TPU v5e, $2.70 for Trillium, $4.20 for TPU v5p, and $12.00 for Ironwood. Flexible scheduling and one- or three-year commitments reduce those rates substantially : often by half or more.

Is TPU cheaper than GPU?

Sometimes, but hourly price alone is not enough. TPU rates are quoted per chip-hour while GPU offerings are often packaged as complete virtual machines. Compare the cost of completing the same workload at the same quality and latency target.

What is the best metric for comparing TPU and GPU cost?

Use workload metrics such as cost per million successful tokens, cost per training run, or cost to reach target model quality : not peak FLOPS or hourly rate.

How much does the Ironwood TPU cost?

Google Cloud lists Ironwood at $12.00 per chip-hour on demand in us-central1 (Iowa) and $13.20 in europe-west2 (London). Flex-start capacity is $6.00 per hour, a one-year commitment is $8.40 in Iowa, and a three-year commitment is $5.40.

How do you calculate TPU workload cost?

Capacity cost is chips × chip-hours × price per chip-hour. Sixty-four chips for ten hours at $2.70 costs $1,728. That is capacity cost, not workload economics : a slower accelerator at a lower hourly rate can still cost more per completed job.