Architecture desk // dense numbers first

AI Accelerator Comparison Table

NVIDIA, AMD, Google TPU, and AWS Trainium compared on the constraints that shape real clusters: precision-normalized peak compute, HBM capacity and bandwidth, scale-up fabric, power boundary, form factor, and access.

re-verify on every major silicon launch · global availability; export-control variants exist

NVIDIAAMDGoogleAWS

Verified comparison bus

Sort any column. “Winner” highlights the strongest numeric value per comparable column—not the best product for every workload.
Source
NVIDIA B300 SXMBlackwell Ultra · shipping 4.5 PF dense
9 PF sparse
2.25 PF dense
4.5 PF sparse
288 GB HBM3e 8.0 TB/s 1.8 TB/s NVLink 5 8 GPUs HGX ≤1,000 W exposed capa SXM; systems + cloud N
NVIDIA B200 SXMBlackwell · shipping 4.5 PF dense
9 PF sparse
2.25 PF dense
4.5 PF sparse
180 GB HBM3e up to 8.0 TB/s 1.8 TB/s NVLink 5 8 GPUs HGX ≤1,000 W cap SXM; systems + cloud N
NVIDIA H100 SXMHopper · prior generation 1.979 PF dense
3.958 PF sparse
0.989 PF dense
1.979 PF sparse
80 GB HBM3 3.35 TB/s 900 GB/s NVLink 4 8 GPUs HGX/DGX up to 700 W SXM; systems + cloud N
AMD Instinct MI355XCDNA 4 · launched 2025-06-12 5.0 PF dense
10.1 PF sparse
2.5 PF dense
5.0 PF sparse
288 GB HBM3e 8.0 TB/s 7 × 153 GB/s IF 8 GPUs UBB 1,400 W TBP OAM; systems + cloud A
AMD Instinct MI300XCDNA 3 · launched 2023-12-06 2.61 PF dense
5.22 PF sparse
1.30 PF dense
2.61 PF sparse
192 GB HBM3 5.3 TB/s 7 × 128 GB/s IF 8 GPUs UBB 750 W peak TBP OAM; systems + cloud A
Google TPU7xIronwood · latest Cloud TPU 4.614 PF vendor peak
no multiplier stated
2.307 PF vendor peak 192 GiB HBM 7.38 TB/s 1.2 TB/s ICI 9,216 chips pod not disclosed 4-chip VM; Google Cloud G
Google TPU v5pprior training generation 459 TF vendor peak
no multiplier stated
459 TF vendor peak 95 GiB HBM 2.765 TB/s 1.2 TB/s ICI 8,960 chips pod not disclosed 4-chip VM; Google Cloud G
AWS Trainium33 nm · Trn3 UltraServer 2.52 PF vendor peak
no multiplier stated
peak not disclosed 144 GB HBM3e 4.9 TB/s 2.0 TB/s NeuronLink 4 144 chips UltraServer not disclosed UltraServer; EC2 only W
AWS Trainium2Trn2 · generally available 1.3 PF dense
5.2 PF sparse
peak not disclosed 96 GiB HBM3 2.9 TB/s 1.0 TB/s NeuronLink 64 chips UltraServer not disclosed 1/16-chip EC2; 64-chip US W

a B300/B200 power is a configured platform power-limit boundary, not a promise of constant draw. NVIDIA’s B300 guide exposes a 1,000 W allowable maximum and a lower default setpoint; system policy and workload determine observed power.

Dense/sparse labels are explicit only where the vendor publishes both. Google/AWS “vendor peak” values are shown without inventing a density conversion. GB and GiB follow each vendor’s unit. Fabric numbers are bidirectional vendor figures and are not latency-equivalent.

Two-chip column picker

Read the table without lying to yourself

Precision is a workload contract

FP8, BF16, FP16, TF32, and FP4 trade range or precision for throughput. Compare the same datatype and accuracy policy.

H100 FP8 dense 1.979 PF ≠ H100 BF16 dense 0.989 PF; the hardware did not become “2× faster” for an unchanged numerical contract.

Sparsity is not free compute

Structured sparsity figures assume a supported pattern and kernels able to exploit it. Dense models cannot claim the multiplier.

MI355X OCP-FP8: 5.0 PF dense / 10.1 PF sparse. A dense-vs-sparse chart must show both labels.

HBM capacity sets fit

Capacity controls model shards, KV-cache headroom, optimizer state, and batch size before compute enters the discussion.

A 180 GB B200 has 100 GB more HBM than an 80 GB H100; that can remove a shard even when peak FLOPS are irrelevant.

Bandwidth sets the token roof

Decode-heavy inference often streams weights and KV state, so delivered tokens can be limited by HBM traffic rather than tensor throughput.

Arithmetic intensity = operations ÷ bytes moved. If a kernel needs 2 TB/s but achieves 1 TB/s, more nominal FLOPS cannot close that memory deficit.

Scale-up is a locality boundary

NVLink, Infinity Fabric, ICI, and NeuronLink create different fast communication domains; crossing to the data-center network changes latency and contention.

Eight B200s share an HGX NVLink domain. TPU7x exposes a 9,216-chip pod topology—but those are not equivalent fabrics or scheduling units.

Power is a boundary, not a meter

TBP/TDP/power limits size electrical and cooling envelopes. Workload, clocks, caps, and chassis overhead determine observed facility draw.

8 × 1,400 W MI355X = 11.2 kW accelerator-only. It is not an 11.2 kW server: CPUs, NICs, fans, pumps, and conversion losses remain.

Architecture decision shortcuts

NVIDIA

Use when ecosystem risk dominates

CUDA, NCCL, TensorRT-LLM, and broad framework/kernel coverage reduce porting friction. The trade is a tightly coupled proprietary platform and often constrained supply.

Choose B200/B300 when an existing CUDA fleet, custom CUDA kernels, or vendor-certified appliances matter more than raw HBM per device.
AMD

Use when memory and open tooling dominate

MI355X pairs 288 GB HBM3e with ROCm and an eight-GPU UBB domain. Validate every critical operator and collective on the exact ROCm/container release.

Choose MI355X when a 200+ GB shard avoids tensor parallelism—after proving your model on production ROCm, not only a framework compatibility list.
Google Cloud TPU

Use when XLA-scale co-design fits

TPU7x offers a very large ICI pod and JAX/PyTorch support through Google Cloud. TensorFlow is not supported on TPU7x per the current product page.

Choose TPU7x for a JAX-native training program prepared to shard over a 3D torus; do not choose it when on-prem ownership is required.
AWS Trainium

Use when AWS-native economics dominate

Trainium integrates Neuron, EC2 UltraServers, EFA, SageMaker/EKS pathways, and PyTorch/JAX. It is cloud-only and requires Neuron compilation/profiling work.

Choose Trn3 when the workload already lives in AWS and a Neuron proof shows better cost per accepted token—not from peak PFLOPS alone.
Selection gate: benchmark the exact model, sequence-length mix, batch policy, accuracy target, compiler/runtime version, and scale. Peak specifications shortlist hardware; they do not award a purchase.

Derived efficiency, with the algebra exposed

bandwidth per watt = GB/s ÷ W · memory per kW = GB ÷ (W/1000) · domain device kW = chips × W ÷ 1000. Accelerator-only figures omit host, networking, cooling, and power-conversion overhead.

AcceleratorHBM BW / device WHBM / device kWDocumented scale-up unitAccelerator-only power at capInterpretation / gotcha
B300 SXM8,000 ÷ 1,000 = 8.00 GB/s/W288 ÷ 1 = 288 GB/kW8-GPU HGX8 × 1,000 = 8.0 kWUses exposed max cap, not default or measured draw.
B200 SXM8,000 ÷ 1,000 = 8.00 GB/s/W180 ÷ 1 = 180 GB/kW8-GPU HGX8 × 1,000 = 8.0 kW“Up to” bandwidth; platform may publish 7.7 TB/s in earlier material.
H100 SXM3,350 ÷ 700 = 4.79 GB/s/W80 ÷ .7 = 114 GB/kW8-GPU HGX/DGX8 × 700 = 5.6 kWTDP is configurable; H100 NVL is a different product.
MI355X8,000 ÷ 1,400 = 5.71 GB/s/W288 ÷ 1.4 = 206 GB/kW8-GPU UBB8 × 1,400 = 11.2 kWAMD also documents up to 64 air-cooled or 128 liquid-cooled GPUs/rack; facility design is separate.
MI300X5,300 ÷ 750 = 7.07 GB/s/W192 ÷ .75 = 256 GB/kW8-GPU UBB8 × 750 = 6.0 kWPeak TBP and peak bandwidth are simultaneous arithmetic inputs, not a benchmark.
TPU7xN/D—power undisclosedN/Dup to 9,216-chip podN/DDo not reverse-engineer chip watts from facility or marketing efficiency claims.
TPU v5pN/D—power undisclosedN/D8,960-chip podN/DPod size is not a rack count and is not directly comparable to HGX.
Trainium3N/D—power undisclosedN/D144-chip UltraServerN/DAWS publishes performance/W ratios, not comparable per-chip watts.
Trainium2N/D—power undisclosedN/D64-chip UltraServerN/DDo not substitute instance power estimates for vendor chip power.

Software and lock-in profile

PlatformPrimary stackFramework pathAccess modelLock-in boundaryProof before commitment
NVIDIACUDA, NCCL, cuDNN, TensorRT-LLM, Triton Inference ServerBroad PyTorch/JAX/TensorFlow ecosystemBuy systems or rent from many cloudsCUDA kernels, NCCL topology, TensorRT engine artifactsProfile kernels, collectives, memory, and delivered tokens on the target GPU.
AMDROCm, HIP, RCCL, composable kernelsPyTorch/JAX paths; validate exact support matrixBuy systems or selected cloudsHIP/ROCm versions, operator gaps, tuning recipesRun production containers and the complete operator graph at target scale.
Google TPUXLA, JAX, PJRT, GKE/Compute Engine toolingJAX + PyTorch on TPU7x; no TensorFlow on TPU7xGoogle Cloud onlyXLA compilation, sharding specs, Cloud TPU topologyMeasure compile time, recompiles, collectives, and host↔device stalls.
AWS TrainiumAWS Neuron SDK, NKI, Neuron ExplorerPyTorch/JAX plus Optimum Neuron and AWS servicesAWS EC2 onlyNeuron compiler/runtime, NKI kernels, EC2 topologyCompile the full graph; measure fallback ops, accepted-token cost, and EFA scaling.

Advanced boundaries

Scale-up first; scale-out second
Keep the highest-traffic tensor/pipeline parallel groups inside the fast local domain where possible. Crossing from NVLink/Infinity Fabric/ICI/NeuronLink to InfiniBand, Ethernet, or EFA changes bandwidth, latency, oversubscription, failure domains, and collective tuning. Example: an eight-way tensor-parallel group fits one HGX B200; a ninth shard crosses the node fabric.
Model FLOP utilization is an outcome, not a chip spec
MFU divides useful model math by theoretical peak over elapsed time. It changes with sequence length, batch size, recomputation, communication, compiler maturity, and numerical format. Do not apply a generic “30–60%” band as if it were verified for every model; calculate it from the exact run and disclose the FLOP accounting convention.
Bandwidth roofs need achieved bytes, not vendor peak
A roofline check pairs operational intensity with achieved memory bandwidth. Profile device counters: if an attention decode step moves 400 GB and takes 100 ms, its achieved bandwidth is 4 TB/s. Compare that with the device’s peak only after confirming the counter boundary and whether compression/cache hits reduce HBM traffic.
Capacity-per-dollar is usually not a chip metric
Public cloud pricing bundles accelerators, host CPUs, RAM, networking, commitments, region, and service terms; system purchase prices are negotiated. Compute usable HBM-hours ÷ total bill from a dated quote instead of inventing a chip MSRP. Include checkpoint/restart loss and reservation utilization.
Availability is part of architecture
A cloud-only ASIC cannot satisfy an on-prem requirement; a purchasable OAM still needs qualified servers, firmware, cooling, and delivery. Verify region, quota, reservation type, lead time, support matrix, and export controls before freezing topology. China-market variants exist but are intentionally outside this global comparison.

Common mistakes / anti-patterns

Sparse vs. dense headline

Quoting 10.1 sparse PF for MI355X against 4.5 dense PF for B200 silently changes the workload. Compare dense-to-dense or clearly qualify the sparsity pattern.

Cross-precision leaderboard

FP4 inference, FP8 training, and BF16 training have different accuracy and software constraints. Match datatype, accumulation, quantization recipe, and quality target.

Ignoring model fit

A faster chip that requires another shard can lose to a higher-memory device through added communication. Budget weights, KV cache, activations, optimizer, fragmentation, and runtime workspaces.

Treating TBP/TDP as the electric bill

Device limits omit CPU, NIC, memory, storage, fans/pumps, and conversion losses. Use measured server/rack power at the intended cap and workload for facility planning.

Peak FLOPS = model FLOPS

Kernel efficiency, bubbles, collectives, memory stalls, and host work consume time. Publish end-to-end tokens/s or time-to-train at a stated quality target, then explain utilization.

Scale-up domain = identical network

Eight HGX GPUs, 144 Trainium chips, and 9,216 TPU chips differ in topology, scheduling, latency, and failure scope. A domain count alone is not a communication benchmark.

Spec-sheet benchmark shopping

Use MLPerf only when system category, benchmark version, scenario, accuracy division, availability status, and result boundary match your question. Vendor peak specs are not MLPerf results.

Primary sources and freshness register

Every hardware figure above is volatile and was checked against these vendor pages on 2026-08-09. Re-open the linked page before procurement or topology freeze.

David Veksler is a Principal AI Engineer in Denver. He leads agentic AI engineering at Antech, a Mars company, and builds AI platforms for regulated financial firms. This page was produced by a governed, multi-agent Claude Code pipeline with a git audit trail. How it's built Case studies