AI 加速卡规格库
本页收录 26 款 AI 加速卡的公开规格,按厂商与发布顺序排列。FP16/FP8/FP4 均为厂商公布的密集(dense)算力,未采用稀疏(sparse)宣传值;标价为空表示厂商未公开官方定价,需向渠道询价。如需按算力挑选,可关注 Cerebras WSE-3 Turbo(FP16 250,000 TFLOPS)。
| # | 型号 | 厂商 | 家族 | 制程 | 发布 | 显存 | 带宽 | FP16 (TFLOPS) | FP8 (TFLOPS) | FP4 (TFLOPS) | 功耗 | 互联 | 标价 | 官网 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | NVIDIA Rubin (Vera Rubin) | NVIDIA | Rubin | Not disclosed by NVIDIA | 2026-08 | 288 GB | 19.2 TB/s | 4,000 | 17,500 | 35,000 | — | NVLink 6 (3 TB/s per GPU) + Vera CPU NVLink-C2C (1.8 TB/s) | 未公开 | 规格页 |
| 2 | NVIDIA B300 (Blackwell Ultra) | NVIDIA | Blackwell Ultra | TSMC 4NP | 2025 | 288 GB | 8 TB/s | 2,500 | 5,000 | 15,000 | 1400 W | NVLink 5 (1.8 TB/s) | 未公开 | 规格页 |
| 3 | NVIDIA GB200 (Blackwell) | NVIDIA | Blackwell | TSMC 4NP | 2025 | 372 GB | 16 TB/s | 5,000 | 10,000 | 20,000 | 2700 W | NVLink 5 (1.8 TB/s) + Grace CPU C2C (900 GB/s) | 未公开 | 规格页 |
| 4 | NVIDIA B200 | NVIDIA | Blackwell | TSMC 4NP | 2025 | 180 GB | 8 TB/s | 2,500 | 5,000 | 10,000 | 1000 W | NVLink 5 (1.8 TB/s) | 未公开 | 规格页 |
| 5 | NVIDIA H200 | NVIDIA | Hopper | TSMC 4N | 2024 | 141 GB | 4.8 TB/s | 989 | 1,979 | — | 700 W | NVLink 4 (900 GB/s) | 未公开 | 规格页 |
| 6 | NVIDIA H100 | NVIDIA | Hopper | TSMC 4N | 2022 | 80 GB | 3.35 TB/s | 989 | 1,979 | — | 700 W | NVLink 4 (900 GB/s) | $30000 | 规格页 |
| 7 | NVIDIA A100 80GB | NVIDIA | Ampere | TSMC 7nm | 2020 | 80 GB | 2 TB/s | 312 | — | — | 400 W | NVLink 3 (600 GB/s) | $18000 | 规格页 |
| 8 | NVIDIA RTX PRO 6000 Blackwell | NVIDIA | Blackwell | TSMC 4NP | 2025 | 96 GB | 1.79 TB/s | 360 | 720 | 1,440 | 600 W | PCIe 5.0 | $8500 | 规格页 |
| 9 | NVIDIA RTX 4090 | NVIDIA | Ada Lovelace | TSMC 4N | 2022 | 24 GB | 1.01 TB/s | 165 | 660 | — | 450 W | PCIe 4.0 | $1600 | 规格页 |
| 10 | NVIDIA RTX 5090 | NVIDIA | Blackwell | TSMC 4NP | 2025 | 32 GB | 1.79 TB/s | 209 | 838 | 3,352 | 575 W | PCIe 5.0 | $2000 | 规格页 |
| 11 | AMD MI455X | AMD | Instinct | Not disclosed by AMD | 2026-07 | 432 GB | 23.3 TB/s | 5,030 | 20,130 | 40,260 | — | UALink over Ethernet (3.6 TB/s bidirectional scale-up per GPU) plus 2,400 Gb/s scale-out | 未公开 | 规格页 |
| 12 | AMD MI355X | AMD | Instinct | TSMC N3 / N6 | 2025-06 | 288 GB | 8 TB/s | 2,500 | 5,000 | 10,100 | 1400 W | Infinity Fabric (7x 153 GB/s scale-up + 1x 128 GB/s scale-out) | 未公开 | 规格页 |
| 13 | AMD Instinct MI350P | AMD | Instinct | TSMC N3 / N6 | 2026-05 | 144 GB | 4 TB/s | 1,150 | 2,300 | 4,600 | 600 W | PCIe 5.0 x16 (no GPU-to-GPU Infinity Fabric links) | 未公开 | 规格页 |
| 14 | AMD MI325X | AMD | Instinct | TSMC N5 / N6 | 2024-Q4 | 256 GB | 6 TB/s | 1,300 | 2,600 | — | 1000 W | Infinity Fabric (8x 153 GB/s) | 未公开 | 规格页 |
| 15 | AMD MI300X | AMD | Instinct | TSMC N5 / N6 | 2023-Q4 | 192 GB | 5.3 TB/s | 1,300 | 2,600 | — | 750 W | Infinity Fabric | $15000 | 规格页 |
| 16 | Google TPU7x (Ironwood) | TPU | Not disclosed by Google | 2026-03 | 192 GB | 7.38 TB/s | 2,307 | 4,614 | — | — | ICI (1.2 TB/s bidirectional per chip), 9,216-chip pods | 未公开 | 规格页 | |
| 17 | Google TPU v6e (Trillium) | TPU | Not disclosed by Google | 2024-12 | 32 GB | 1.64 TB/s | 918 | 918 | — | — | ICI (800 GB/s bidirectional per chip), 256-chip pods | 未公开 | 规格页 | |
| 18 | Google TPU v5p | TPU | TSMC 3nm | 2023-12 | 95 GB | 2.77 TB/s | 459 | 459 | — | 700 W | ICI (Inter-Chip Interconnect, 1.2 TB/s bidirectional per chip) | 未公开 | 规格页 | |
| 19 | Google TPU v5e | TPU | TSMC 5nm | 2023-08 | 16 GB | 0.8 TB/s | 197 | 393 | — | 170 W | ICI (400 GB/s bidirectional per chip) | 未公开 | 规格页 | |
| 20 | AWS Trainium 2 | AWS | Trainium | TSMC 5nm | 2024-12 | 96 GB | 2.9 TB/s | 667 | 1,299 | — | 500 W | NeuronLink-v3 (1.28 TB/s per chip) | 未公开 | 规格页 |
| 21 | AWS Trainium3 | AWS | Trainium | TSMC 3nm | 2025-12 | 144 GB | 4.9 TB/s | 671 | 2,517 | 2,517 | — | NeuronLink-v4 (2.56 TB/s per chip) | 未公开 | 规格页 |
| 22 | AWS Inferentia 2 | AWS | Inferentia | TSMC 5nm | 2023-04 | 32 GB | 0.82 TB/s | 190 | — | — | — | NeuronLink | 未公开 | 规格页 |
| 23 | Apple M4 Max | Apple | Apple Silicon | TSMC 3nm | 2024-10 | 128 GB | 0.55 TB/s | 38 | — | — | 60 W | On-die (unified memory) | $4699 | 规格页 |
| 24 | Cerebras WSE-3 Turbo | Cerebras | WSE | Not disclosed by Cerebras | 2026-08 | 44 GB | 43200 TB/s | 250,000 | — | — | — | On-wafer fabric (53.5 PB/s), 2.4 Tb/s off-wafer I/O | 未公开 | 规格页 |
| 25 | Cerebras WSE-3 | Cerebras | WSE | TSMC 5nm | 2024 | 44 GB | 21000 TB/s | 125,000 | — | — | 23000 W | On-wafer (no interconnect needed) | 未公开 | 规格页 |
| 26 | Groq LPU | Groq | LPU | GlobalFoundries 14nm (gen 1) | 2023 | 0 GB | 80 TB/s | 188 | — | — | 350 W | GroqLink | 未公开 | 规格页 |
NVIDIA Rubin (Vera Rubin):Blackwell successor. 288GB HBM4 per GPU. NVIDIA began production shipments in August 2026. FLOPS are NVIDIA dense figures (FP8 is FP8/FP6 training, FP4 is NVFP4 training); the 50 PFLOPS NVFP4 inference headline is a sparse figure and is not used here.
NVIDIA B300 (Blackwell Ultra):Mid-cycle Blackwell refresh: 208 billion transistors across two dies, with NVIDIA quoting 1.5x the HBM3e capacity and 1.5x the dense NVFP4 rate of Blackwell. Dense figures shown; NVIDIA rates the NVL72 rack at 1,440 PFLOPS NVFP4 with sparsity and 1.1 exaFLOPS dense.
NVIDIA GB200 (Blackwell):Blackwell flagship: dual B200 + Grace CPU. FP4 native. The chip every major lab is buying for 2025-2026 frontier training.
NVIDIA B200:Single-die Blackwell. 180GB HBM3e per GPU (1,440GB across an 8-GPU DGX B200). Inference workhorse for serving 100B+ MoE models on a single GPU.
NVIDIA H200:Hopper refresh with HBM3e. The sweet spot for production inference in 2025. Extra memory means 70B-class models fit single-GPU at FP8.
NVIDIA H100:The chip that built the 2023-2024 LLM boom. Still the most-deployed AI GPU. FP8 native. 80GB HBM3.
NVIDIA A100 80GB:Workhorse of the late-2010s ML wave. Still widely used for training smaller models and inference where H100/H200 supply is tight.
NVIDIA RTX PRO 6000 Blackwell:Workstation Blackwell. 96GB GDDR7. Good fit for on-prem agent dev environments where multi-H100 rentals are overkill.
NVIDIA RTX 4090:Best price/perf for local agent dev. 24GB VRAM fits Q4-quantized 70B models. The default consumer-tier choice for self-hosted Ollama/llama.cpp.
NVIDIA RTX 5090:Consumer Blackwell flagship. 32GB GDDR7. FP4 native. Strong pick for local agent dev with current-frontier features.
AMD MI455X:AMD's MI400-series flagship, launched at Advancing AI on July 23, 2026 with Helios racks already in production. 432GB HBM4. Per-GPU figures are MXFP4, MXFP8 and dense BF16 peaks; AMD publishes the rack numbers of 2.9 exaflops FP4, 1.4 exaflops FP8, 31TB HBM4 and 1.7 PB/s across 72 GPUs. OpenAI expects to bring Helios online starting in Q4 2026.
AMD MI355X:AMD's MI350-series flagship. 288GB HBM3e with native MXFP4 and MXFP6 support. 1400W typical board power. The step between MI325X and the Helios-rack MI400 series.
AMD Instinct MI350P:CDNA 4 in a standard PCIe slot: 128 compute units, roughly half an MI350X, with 144GB HBM3e and a configurable 450W power mode. Peak dense figures shown; AMD doubles its 8-bit and 16-bit ratings with structured sparsity. Built to drop inference capacity into existing air-cooled enterprise servers instead of liquid-cooled racks.
AMD MI325X:AMD's answer to H200. 256GB HBM3e (highest in the catalog at single-die). Strong inference value when supply is constrained on NVIDIA.
AMD MI300X:Wider availability than NVIDIA flagship; cheaper to rent. ROCm software stack matures; vLLM and PyTorch first-class on MI300X in 2026.
Google TPU7x (Ironwood):Seventh-generation TPU and the first in the Ironwood family. 192 GiB HBM per chip, six times Trillium. Google positions it for large-scale training and inference of LLMs, MoE, and diffusion models.
Google TPU v6e (Trillium):Sixth-generation TPU. Twice the per-chip BF16 compute of v5p with a third of the memory; Google used it to train Gemini 2.0.
Google TPU v5p:Google's previous-generation training TPU, now behind Trillium and Ironwood. Used for Gemini training. Only available on GCP; pod-scale (8960 chips) is a different beast than per-chip rentals.
Google TPU v5e:TPU inference tier. The 8-bit figure is Google's INT8 rating (393 TOPS). Cheaper than v5p; 256-chip pods. Best fit for Google Cloud customers serving Gemini-style inference at scale.
AWS Trainium 2:AWS custom training/inference silicon. Dense figures shown; AWS rates it at 2,563 TFLOPS with sparsity. Anthropic uses these for Claude training. Cheaper TCO on AWS than NVIDIA flagship for committed workloads.
AWS Trainium3:Third-generation AWS training and inference silicon. 144 GB HBM3e, 4.9 TB/s bandwidth. 8-bit and 4-bit figures are MXFP8 and MXFP4 per AWS; BF16 is roughly flat with Trainium 2. Roughly 30-50% lower cost per hour than H100/H200 on committed AWS workloads.
AWS Inferentia 2:AWS inference silicon. Lower compute than Trainium 2 but cheaper per-token for high-volume serving workloads on AWS.
Apple M4 Max:128GB unified memory means a 2024-2025 Mac Studio or MacBook Pro runs 70B-class models at Q4 quantization with usable speed. The dark-horse local-inference platform. Apple publishes memory and bandwidth for Apple Silicon but no TFLOPS rating for the M5 generation, so no M5 row is listed here.
Cerebras WSE-3 Turbo:Announced August 18, 2026. Same 46,225 mm^2 wafer, 4 trillion transistors and 900,000 cores as WSE-3, but Cerebras doubles AI compute to 250 PFLOPS and memory bandwidth to 43.2 PB/s (shown here as 43,200 TB/s). A CS-4 rack holds three wafers for 750 PFLOPS. Compute figure is on the same basis as the 125 PFLOPS FP16 rating of WSE-3.
Cerebras WSE-3:Single 46,225 mm^2 wafer. Holds entire activations on-chip; eliminates many distributed-training pains. Cerebras rates aggregate memory bandwidth at 21 PB/s (shown here as 21,000 TB/s).
Groq LPU:Deterministic single-thread inference silicon. SRAM-only memory model means small per-chip capacity but massive bandwidth. Behind Groq's 700+ tokens/sec Llama 4 Scout serving.
同步频率:每 3 小时同步一次,规格变更后自动更新,无需人工维护。
口径说明:显存带宽单位 TB/s,算力单位 TFLOPS;TPU/LPU 等非 GPU 架构的算力口径与 NVIDIA 不完全可比,跨厂商对比时请以显存容量与带宽为主。
免责:本页为公开资料汇总,仅供选型参考,最终规格以厂商官方文档为准。