Updated September 4, 2026 · AI infrastructure decision guide

Google TPU vs NVIDIA vs AMD: choose by workload, not peak FLOPS.

The 2026 accelerator market changed quickly: Ironwood is priced and generally available, Blackwell Ultra is shipping, Rubin is ramping, AMD has moved to MI350/MI400, and Google has split TPU 8 into training- and inference-optimized systems. The right choice now depends on memory fit, software migration cost, scale topology, availability, and performance on your actual model.

Ironwood pricing is public$12/chip-hour on demand in us-central1; DWS and commitment rates are lower.
NVIDIA Rubin is in productionPartner systems are rolling out in the second half of 2026; GB300 Blackwell Ultra is already available.
AMD moved to MI400MI455X: 432 GB HBM4 and 23.3 TB/s bandwidth; Helios deployments ramp in 2H 2026.
TPU 8 remains “coming soon”8t targets pre-training; 8i targets post-training, inference and reasoning.
2026 decision stack

Hardware is only one Knob.

A fair accelerator decision holds the model, precision, batch/sequence shape, quality target, serving stack and SLO constant—then compares end-to-end goodput and cost.

1. Workload
Pre-training · fine-tuning · RL · reasoning · high-throughput inference · embeddings
2. System fit
Memory · interconnect · framework · kernels · orchestration · availability
3. Evidence
Quality · goodput · TTFT/TPOT · p99 · utilization · cost per successful task
Current landscape

What is actually available as of September 2026?

Do not compare a shipping platform with an announced architecture as if procurement risk were equal. The status column matters as much as the spec column.

Vendor / platformStatusMemoryBandwidthPublished low-precision peakSystem emphasis
Google · Trillium v6eGA32 GB HBM1.638 TB/s918 TFLOPS BF16Cost-efficient training/inference; 256-chip slices
Google · Ironwood TPU7xGA192 GiB HBM3E7.38 TB/s2.307 PFLOPS BF16 · 4.614 PFLOPS FP8Large-scale training, reasoning and inference; 9,216-chip pod
Google · TPU 8tComing soon216 GB HBM + 128 MB SRAM6.528 TB/s12.6 PFLOPS FP4Pre-training and embedding-heavy workloads; 9,600-chip superpod
Google · TPU 8iComing soon288 GB HBM + 384 MB SRAM8.601 TB/s10.1 PFLOPS FP4Post-training, serving, reasoning; Boardfly with up to 1,152 connected chips
NVIDIA · GB300 Blackwell UltraAvailable now288 GB HBM3e / GPURack: up to 576 TB/s across 72 GPUsRack: 1,080 PFLOPS dense FP4 (1,440 with sparsity)Rack-scale reasoning and inference; mature CUDA ecosystem
NVIDIA · Vera RubinProduction ramp288 GB HBM4 / GPU22 TB/s / GPU50 PFLOPS NVFP4 inference · 35 PFLOPS NVFP4 training / GPUNext-gen agentic inference + training; partner rollout in 2H 2026
AMD · MI355XAvailable288 GB HBM3E8 TB/s10.1 PFLOPS MXFP4Large-memory training/inference on CDNA4; ROCm ecosystem
AMD · MI455X2026 rollout432 GB HBM423.3 TB/s40.3 PFLOPS MXFP4Helios rack-scale training/inference; volume deployments expected in 2H 2026

Peak figures are vendor specifications and are not directly comparable across different numeric formats, sparsity assumptions, tensor/matrix definitions, chips, racks or system configurations. Use them for capacity screening—not to declare a workload winner.

Decision framework

Three strong starting points—not three universal winners.

Treat these as hypotheses to benchmark, not purchasing conclusions.

Start with Google TPU when…

Your workload is already on Google Cloud, JAX/XLA or TPU-native tooling; you need pod-scale training; embeddings/SparseCore matter; or Google’s integrated AI Hypercomputer is an architectural advantage.

Start with NVIDIA when…

You want the lowest ecosystem and migration friction, broad PyTorch/kernel support, mature profiling/serving tools, or need the widest choice of cloud, OEM and software integrations.

Start with AMD when…

Memory capacity and bandwidth are first-order constraints, you value an increasingly open ROCm stack, or you want to qualify an alternative to CUDA for training and high-throughput inference.

Important: “GPU vs TPU” is increasingly the wrong level of abstraction. Modern deployments are rack- and cluster-scale systems. Compare the complete system: host CPUs, memory, scale-up fabric, scale-out network, compiler/runtime, serving stack, observability and commercial availability.
Software ecosystem

The gap is narrowing—but the stacks are still different.

NVIDIA

CUDA remains the compatibility default

CUDA, cuDNN, NCCL, TensorRT, Triton, Nsight and the third-party kernel ecosystem still minimize migration risk for many PyTorch teams. That developer and operations familiarity is a real TCO factor.

Google TPU

JAX is established; native PyTorch is advancing

Google now presents TPUs as supporting high-performance PyTorch and JAX, plus vLLM. TorchTPU’s native eager-first path was announced in April 2026; Google’s TPU 8 deep dive still labels native PyTorch support as preview, so qualification remains workload-specific.

AMD

ROCm is now a serious production stack

MI350/MI400 support spans PyTorch, TensorFlow, JAX, ONNX Runtime, vLLM and Triton. ROCm 10, released in August 2026, emphasizes repeatable production deployment, profiling, distributed communication and operational tooling.

Current Google Cloud price anchors

Ironwood pricing changed the picture.

The previous version said Ironwood had no public price. Google Cloud now publishes regional Ironwood rates. Compare the same consumption mode whenever possible.

Cloud TPU · per chip-hour · us regions

TPU v5eOn demand / DWS flex / DWS calendar
$1.20 / $0.60 / $0.84
Trillium v6eOn demand / DWS flex / DWS calendar
$2.70 / $1.35 / $1.89
TPU v5pOn demand / DWS flex / DWS calendar
$4.20 / $2.10 / $2.94
Ironwoodus-central1 · on demand / DWS flex / DWS calendar
$12.00 / $6.00 / $8.40

Google Cloud GPU VMs · 8 GPUs / VM

A3 High · 8× H100On demand / DWS flex / DWS calendar
$88.49 / $38.32 / $41.60
A3 Ultra · 8× H200On demand / DWS flex / DWS calendar
$84.81 / $42.40 / $59.36
A4 High · 8× B200DWS flex / DWS calendar; on-demand not listed
$64.44 / $90.22
Do not divide price by “accelerator” and stop there. TPU prices are per chip-hour; GPU VM prices include the complete VM configuration. Different accelerators deliver different throughput, memory, networking and utilization. The meaningful unit is cost per completed training target or cost per successful inference task at a fixed quality/SLO.

Same-mode illustration

DWS calendar capacity: 8× Trillium ≈ $15.12/hr; 8× Ironwood ≈ $67.20/hr; 8×H100 A3 High = $41.60/hr; 8×H200 A3 Ultra = $59.36/hr; 8×B200 A4 High = $90.22/hr.

What this does tell you

It gives a current Google Cloud procurement anchor for capacity blocks under one consumption mode.

What it does not tell you

It does not normalize model quality, tokens/sec, training goodput, interconnect scale, CPU/host resources or engineering cost.

Benchmark protocol

How to make the decision defensible.

Public benchmarks are useful for shortlisting. Procurement should be based on your model, your stack and your SLO.

Freeze the quality target

Use the same model checkpoint, evaluation set and acceptable quality regression. Lower precision counts only if quality remains acceptable.

Measure end-to-end goodput

Training: examples/tokens per wall-clock hour at target convergence. Inference: completed requests/tokens at the required latency and concurrency.

Separate TTFT and generation latency

For reasoning/agents, track time-to-first-token, time-per-output-token, p95/p99 and long-context behavior—not just aggregate tokens/sec.

Stress the memory boundary

Measure the largest practical batch, context/KV cache, optimizer state and model shard that fits without damaging utilization.

Scale beyond one node

Test collective efficiency, fabric bottlenecks, failure recovery and checkpoint behavior at the scale you plan to operate.

Price the operating model

Include reservations/DWS/spot assumptions, host CPU, storage, network, idle time, engineer effort, migration work and support burden.

DataKnobs view

Accelerator selection is an experimentable Knob.

Hardware should not be a one-time architecture opinion. Treat the accelerator, precision, parallelism, batch shape, serving engine and cost controls as measurable choices.

KREATE

Build a repeatable workload harness: model, data, deployment recipe, test traffic, quality evaluation and cost capture.

KONTROLS

Govern spend, allowed regions, model/data lineage, reliability thresholds, security boundaries and evidence for production approval.

KNOBS

Compare accelerator family, precision, batch size, context length, tensor/pipeline parallelism, routing, quantization and serving configuration.

FAQ

Current questions, current answers.

Which AI accelerator should an enterprise choose in 2026?

There is no universal winner. NVIDIA is the lowest-friction default for broad ecosystem compatibility; Google TPU is compelling for Google Cloud/JAX/XLA and very large TPU-native workloads; AMD is increasingly compelling for memory-heavy systems and teams qualifying an open alternative. Benchmark the real workload.

Does Google Cloud now publish Ironwood pricing?

Yes. Google Cloud lists Ironwood at $12.00 per chip-hour on demand in us-central1, $6.00 DWS flex-start and $8.40 DWS calendar mode. Other regions and commitments differ.

Are TPU 8t and TPU 8i available?

Google Cloud currently lists both as “coming soon.” 8t is optimized for pre-training and embedding-heavy workloads; 8i is optimized for post-training, serving, reasoning and low-latency MoE workloads.

Is Rubin the latest NVIDIA platform?

Rubin is NVIDIA’s next-generation platform and is in full production, with partner availability during the second half of 2026. For broadly available current systems, Blackwell/Blackwell Ultra—such as GB300 NVL72—remains the immediate procurement reference.

Which platform has the most memory per accelerator?

Among the platforms covered here, AMD MI455X has the largest published per-accelerator memory at 432 GB HBM4. TPU 8i, Rubin and MI355X are each specified at 288 GB, though TPU 8i remains coming soon.