Precision is a workload contract
FP8, BF16, FP16, TF32, and FP4 trade range or precision for throughput. Compare the same datatype and accuracy policy.
NVIDIA, AMD, Google TPU, and AWS Trainium compared on the constraints that shape real clusters: precision-normalized peak compute, HBM capacity and bandwidth, scale-up fabric, power boundary, form factor, and access.
re-verify on every major silicon launch · global availability; export-control variants exist
| Source | |||||||||
|---|---|---|---|---|---|---|---|---|---|
| NVIDIA B300 SXMBlackwell Ultra · shipping | 4.5 PF dense 9 PF sparse |
2.25 PF dense 4.5 PF sparse |
288 GB HBM3e | 8.0 TB/s | 1.8 TB/s NVLink 5 | 8 GPUs HGX | ≤1,000 W exposed capa | SXM; systems + cloud | N |
| NVIDIA B200 SXMBlackwell · shipping | 4.5 PF dense 9 PF sparse |
2.25 PF dense 4.5 PF sparse |
180 GB HBM3e | up to 8.0 TB/s | 1.8 TB/s NVLink 5 | 8 GPUs HGX | ≤1,000 W cap | SXM; systems + cloud | N |
| NVIDIA H100 SXMHopper · prior generation | 1.979 PF dense 3.958 PF sparse |
0.989 PF dense 1.979 PF sparse |
80 GB HBM3 | 3.35 TB/s | 900 GB/s NVLink 4 | 8 GPUs HGX/DGX | up to 700 W | SXM; systems + cloud | N |
| AMD Instinct MI355XCDNA 4 · launched 2025-06-12 | 5.0 PF dense 10.1 PF sparse |
2.5 PF dense 5.0 PF sparse |
288 GB HBM3e | 8.0 TB/s | 7 × 153 GB/s IF | 8 GPUs UBB | 1,400 W TBP | OAM; systems + cloud | A |
| AMD Instinct MI300XCDNA 3 · launched 2023-12-06 | 2.61 PF dense 5.22 PF sparse |
1.30 PF dense 2.61 PF sparse |
192 GB HBM3 | 5.3 TB/s | 7 × 128 GB/s IF | 8 GPUs UBB | 750 W peak TBP | OAM; systems + cloud | A |
| Google TPU7xIronwood · latest Cloud TPU | 4.614 PF vendor peak no multiplier stated |
2.307 PF vendor peak | 192 GiB HBM | 7.38 TB/s | 1.2 TB/s ICI | 9,216 chips pod | not disclosed | 4-chip VM; Google Cloud | G |
| Google TPU v5pprior training generation | 459 TF vendor peak no multiplier stated |
459 TF vendor peak | 95 GiB HBM | 2.765 TB/s | 1.2 TB/s ICI | 8,960 chips pod | not disclosed | 4-chip VM; Google Cloud | G |
| AWS Trainium33 nm · Trn3 UltraServer | 2.52 PF vendor peak no multiplier stated |
peak not disclosed | 144 GB HBM3e | 4.9 TB/s | 2.0 TB/s NeuronLink 4 | 144 chips UltraServer | not disclosed | UltraServer; EC2 only | W |
| AWS Trainium2Trn2 · generally available | 1.3 PF dense 5.2 PF sparse |
peak not disclosed | 96 GiB HBM3 | 2.9 TB/s | 1.0 TB/s NeuronLink | 64 chips UltraServer | not disclosed | 1/16-chip EC2; 64-chip US | W |
a B300/B200 power is a configured platform power-limit boundary, not a promise of constant draw. NVIDIA’s B300 guide exposes a 1,000 W allowable maximum and a lower default setpoint; system policy and workload determine observed power.
Dense/sparse labels are explicit only where the vendor publishes both. Google/AWS “vendor peak” values are shown without inventing a density conversion. GB and GiB follow each vendor’s unit. Fabric numbers are bidirectional vendor figures and are not latency-equivalent.
FP8, BF16, FP16, TF32, and FP4 trade range or precision for throughput. Compare the same datatype and accuracy policy.
Structured sparsity figures assume a supported pattern and kernels able to exploit it. Dense models cannot claim the multiplier.
Capacity controls model shards, KV-cache headroom, optimizer state, and batch size before compute enters the discussion.
Decode-heavy inference often streams weights and KV state, so delivered tokens can be limited by HBM traffic rather than tensor throughput.
NVLink, Infinity Fabric, ICI, and NeuronLink create different fast communication domains; crossing to the data-center network changes latency and contention.
TBP/TDP/power limits size electrical and cooling envelopes. Workload, clocks, caps, and chassis overhead determine observed facility draw.
CUDA, NCCL, TensorRT-LLM, and broad framework/kernel coverage reduce porting friction. The trade is a tightly coupled proprietary platform and often constrained supply.
MI355X pairs 288 GB HBM3e with ROCm and an eight-GPU UBB domain. Validate every critical operator and collective on the exact ROCm/container release.
TPU7x offers a very large ICI pod and JAX/PyTorch support through Google Cloud. TensorFlow is not supported on TPU7x per the current product page.
Trainium integrates Neuron, EC2 UltraServers, EFA, SageMaker/EKS pathways, and PyTorch/JAX. It is cloud-only and requires Neuron compilation/profiling work.
bandwidth per watt = GB/s ÷ W · memory per kW = GB ÷ (W/1000) · domain device kW = chips × W ÷ 1000. Accelerator-only figures omit host, networking, cooling, and power-conversion overhead.
| Accelerator | HBM BW / device W | HBM / device kW | Documented scale-up unit | Accelerator-only power at cap | Interpretation / gotcha |
|---|---|---|---|---|---|
| B300 SXM | 8,000 ÷ 1,000 = 8.00 GB/s/W | 288 ÷ 1 = 288 GB/kW | 8-GPU HGX | 8 × 1,000 = 8.0 kW | Uses exposed max cap, not default or measured draw. |
| B200 SXM | 8,000 ÷ 1,000 = 8.00 GB/s/W | 180 ÷ 1 = 180 GB/kW | 8-GPU HGX | 8 × 1,000 = 8.0 kW | “Up to” bandwidth; platform may publish 7.7 TB/s in earlier material. |
| H100 SXM | 3,350 ÷ 700 = 4.79 GB/s/W | 80 ÷ .7 = 114 GB/kW | 8-GPU HGX/DGX | 8 × 700 = 5.6 kW | TDP is configurable; H100 NVL is a different product. |
| MI355X | 8,000 ÷ 1,400 = 5.71 GB/s/W | 288 ÷ 1.4 = 206 GB/kW | 8-GPU UBB | 8 × 1,400 = 11.2 kW | AMD also documents up to 64 air-cooled or 128 liquid-cooled GPUs/rack; facility design is separate. |
| MI300X | 5,300 ÷ 750 = 7.07 GB/s/W | 192 ÷ .75 = 256 GB/kW | 8-GPU UBB | 8 × 750 = 6.0 kW | Peak TBP and peak bandwidth are simultaneous arithmetic inputs, not a benchmark. |
| TPU7x | N/D—power undisclosed | N/D | up to 9,216-chip pod | N/D | Do not reverse-engineer chip watts from facility or marketing efficiency claims. |
| TPU v5p | N/D—power undisclosed | N/D | 8,960-chip pod | N/D | Pod size is not a rack count and is not directly comparable to HGX. |
| Trainium3 | N/D—power undisclosed | N/D | 144-chip UltraServer | N/D | AWS publishes performance/W ratios, not comparable per-chip watts. |
| Trainium2 | N/D—power undisclosed | N/D | 64-chip UltraServer | N/D | Do not substitute instance power estimates for vendor chip power. |
| Platform | Primary stack | Framework path | Access model | Lock-in boundary | Proof before commitment |
|---|---|---|---|---|---|
| NVIDIA | CUDA, NCCL, cuDNN, TensorRT-LLM, Triton Inference Server | Broad PyTorch/JAX/TensorFlow ecosystem | Buy systems or rent from many clouds | CUDA kernels, NCCL topology, TensorRT engine artifacts | Profile kernels, collectives, memory, and delivered tokens on the target GPU. |
| AMD | ROCm, HIP, RCCL, composable kernels | PyTorch/JAX paths; validate exact support matrix | Buy systems or selected clouds | HIP/ROCm versions, operator gaps, tuning recipes | Run production containers and the complete operator graph at target scale. |
| Google TPU | XLA, JAX, PJRT, GKE/Compute Engine tooling | JAX + PyTorch on TPU7x; no TensorFlow on TPU7x | Google Cloud only | XLA compilation, sharding specs, Cloud TPU topology | Measure compile time, recompiles, collectives, and host↔device stalls. |
| AWS Trainium | AWS Neuron SDK, NKI, Neuron Explorer | PyTorch/JAX plus Optimum Neuron and AWS services | AWS EC2 only | Neuron compiler/runtime, NKI kernels, EC2 topology | Compile the full graph; measure fallback ops, accepted-token cost, and EFA scaling. |
usable HBM-hours ÷ total bill from a dated quote instead of inventing a chip MSRP. Include checkpoint/restart loss and reservation utilization.Quoting 10.1 sparse PF for MI355X against 4.5 dense PF for B200 silently changes the workload. Compare dense-to-dense or clearly qualify the sparsity pattern.
FP4 inference, FP8 training, and BF16 training have different accuracy and software constraints. Match datatype, accumulation, quantization recipe, and quality target.
A faster chip that requires another shard can lose to a higher-memory device through added communication. Budget weights, KV cache, activations, optimizer, fragmentation, and runtime workspaces.
Device limits omit CPU, NIC, memory, storage, fans/pumps, and conversion losses. Use measured server/rack power at the intended cap and workload for facility planning.
Kernel efficiency, bubbles, collectives, memory stalls, and host work consume time. Publish end-to-end tokens/s or time-to-train at a stated quality target, then explain utilization.
Eight HGX GPUs, 144 Trainium chips, and 9,216 TPU chips differ in topology, scheduling, latency, and failure scope. A domain count alone is not a communication benchmark.
Use MLPerf only when system category, benchmark version, scenario, accuracy division, availability status, and result boundary match your question. Vendor peak specs are not MLPerf results.
Every hardware figure above is volatile and was checked against these vendor pages on 2026-08-09. Re-open the linked page before procurement or topology freeze.