Start with Google TPU when…
Your workload is already on Google Cloud, JAX/XLA or TPU-native tooling; you need pod-scale training; embeddings/SparseCore matter; or Google’s integrated AI Hypercomputer is an architectural advantage.
The 2026 accelerator market changed quickly: Ironwood is priced and generally available, Blackwell Ultra is shipping, Rubin is ramping, AMD has moved to MI350/MI400, and Google has split TPU 8 into training- and inference-optimized systems. The right choice now depends on memory fit, software migration cost, scale topology, availability, and performance on your actual model.
A fair accelerator decision holds the model, precision, batch/sequence shape, quality target, serving stack and SLO constant—then compares end-to-end goodput and cost.
Do not compare a shipping platform with an announced architecture as if procurement risk were equal. The status column matters as much as the spec column.
| Vendor / platform | Status | Memory | Bandwidth | Published low-precision peak | System emphasis |
|---|---|---|---|---|---|
| Google · Trillium v6e | GA | 32 GB HBM | 1.638 TB/s | 918 TFLOPS BF16 | Cost-efficient training/inference; 256-chip slices |
| Google · Ironwood TPU7x | GA | 192 GiB HBM3E | 7.38 TB/s | 2.307 PFLOPS BF16 · 4.614 PFLOPS FP8 | Large-scale training, reasoning and inference; 9,216-chip pod |
| Google · TPU 8t | Coming soon | 216 GB HBM + 128 MB SRAM | 6.528 TB/s | 12.6 PFLOPS FP4 | Pre-training and embedding-heavy workloads; 9,600-chip superpod |
| Google · TPU 8i | Coming soon | 288 GB HBM + 384 MB SRAM | 8.601 TB/s | 10.1 PFLOPS FP4 | Post-training, serving, reasoning; Boardfly with up to 1,152 connected chips |
| NVIDIA · GB300 Blackwell Ultra | Available now | 288 GB HBM3e / GPU | Rack: up to 576 TB/s across 72 GPUs | Rack: 1,080 PFLOPS dense FP4 (1,440 with sparsity) | Rack-scale reasoning and inference; mature CUDA ecosystem |
| NVIDIA · Vera Rubin | Production ramp | 288 GB HBM4 / GPU | 22 TB/s / GPU | 50 PFLOPS NVFP4 inference · 35 PFLOPS NVFP4 training / GPU | Next-gen agentic inference + training; partner rollout in 2H 2026 |
| AMD · MI355X | Available | 288 GB HBM3E | 8 TB/s | 10.1 PFLOPS MXFP4 | Large-memory training/inference on CDNA4; ROCm ecosystem |
| AMD · MI455X | 2026 rollout | 432 GB HBM4 | 23.3 TB/s | 40.3 PFLOPS MXFP4 | Helios rack-scale training/inference; volume deployments expected in 2H 2026 |
Peak figures are vendor specifications and are not directly comparable across different numeric formats, sparsity assumptions, tensor/matrix definitions, chips, racks or system configurations. Use them for capacity screening—not to declare a workload winner.
Treat these as hypotheses to benchmark, not purchasing conclusions.
Your workload is already on Google Cloud, JAX/XLA or TPU-native tooling; you need pod-scale training; embeddings/SparseCore matter; or Google’s integrated AI Hypercomputer is an architectural advantage.
You want the lowest ecosystem and migration friction, broad PyTorch/kernel support, mature profiling/serving tools, or need the widest choice of cloud, OEM and software integrations.
Memory capacity and bandwidth are first-order constraints, you value an increasingly open ROCm stack, or you want to qualify an alternative to CUDA for training and high-throughput inference.
CUDA, cuDNN, NCCL, TensorRT, Triton, Nsight and the third-party kernel ecosystem still minimize migration risk for many PyTorch teams. That developer and operations familiarity is a real TCO factor.
Google now presents TPUs as supporting high-performance PyTorch and JAX, plus vLLM. TorchTPU’s native eager-first path was announced in April 2026; Google’s TPU 8 deep dive still labels native PyTorch support as preview, so qualification remains workload-specific.
MI350/MI400 support spans PyTorch, TensorFlow, JAX, ONNX Runtime, vLLM and Triton. ROCm 10, released in August 2026, emphasizes repeatable production deployment, profiling, distributed communication and operational tooling.
The previous version said Ironwood had no public price. Google Cloud now publishes regional Ironwood rates. Compare the same consumption mode whenever possible.
DWS calendar capacity: 8× Trillium ≈ $15.12/hr; 8× Ironwood ≈ $67.20/hr; 8×H100 A3 High = $41.60/hr; 8×H200 A3 Ultra = $59.36/hr; 8×B200 A4 High = $90.22/hr.
It gives a current Google Cloud procurement anchor for capacity blocks under one consumption mode.
It does not normalize model quality, tokens/sec, training goodput, interconnect scale, CPU/host resources or engineering cost.
Public benchmarks are useful for shortlisting. Procurement should be based on your model, your stack and your SLO.
Use the same model checkpoint, evaluation set and acceptable quality regression. Lower precision counts only if quality remains acceptable.
Training: examples/tokens per wall-clock hour at target convergence. Inference: completed requests/tokens at the required latency and concurrency.
For reasoning/agents, track time-to-first-token, time-per-output-token, p95/p99 and long-context behavior—not just aggregate tokens/sec.
Measure the largest practical batch, context/KV cache, optimizer state and model shard that fits without damaging utilization.
Test collective efficiency, fabric bottlenecks, failure recovery and checkpoint behavior at the scale you plan to operate.
Include reservations/DWS/spot assumptions, host CPU, storage, network, idle time, engineer effort, migration work and support burden.
Hardware should not be a one-time architecture opinion. Treat the accelerator, precision, parallelism, batch shape, serving engine and cost controls as measurable choices.
Build a repeatable workload harness: model, data, deployment recipe, test traffic, quality evaluation and cost capture.
Govern spend, allowed regions, model/data lineage, reliability thresholds, security boundaries and evidence for production approval.
Compare accelerator family, precision, batch size, context length, tensor/pipeline parallelism, routing, quantization and serving configuration.
There is no universal winner. NVIDIA is the lowest-friction default for broad ecosystem compatibility; Google TPU is compelling for Google Cloud/JAX/XLA and very large TPU-native workloads; AMD is increasingly compelling for memory-heavy systems and teams qualifying an open alternative. Benchmark the real workload.
Yes. Google Cloud lists Ironwood at $12.00 per chip-hour on demand in us-central1, $6.00 DWS flex-start and $8.40 DWS calendar mode. Other regions and commitments differ.
Google Cloud currently lists both as “coming soon.” 8t is optimized for pre-training and embedding-heavy workloads; 8i is optimized for post-training, serving, reasoning and low-latency MoE workloads.
Rubin is NVIDIA’s next-generation platform and is in full production, with partner availability during the second half of 2026. For broadly available current systems, Blackwell/Blackwell Ultra—such as GB300 NVL72—remains the immediate procurement reference.
Among the platforms covered here, AMD MI455X has the largest published per-accelerator memory at 432 GB HBM4. TPU 8i, Rubin and MI355X are each specified at 288 GB, though TPU 8i remains coming soon.
Vendor claims are labeled as such; cross-vendor “winner” claims are intentionally avoided.