AI 加速卡规格库

显存 / 算力 / 带宽 / 功耗 / 互联 全参数对比 · 共 26 款
数据同步于 2026-10-06 14:13(北京时间) · 每 3 小时更新
算力总览 加速卡规格库 GPU 租用比价

本页收录 26 款 AI 加速卡的公开规格,按厂商与发布顺序排列。FP16/FP8/FP4 均为厂商公布的密集(dense)算力,未采用稀疏(sparse)宣传值;标价为空表示厂商未公开官方定价,需向渠道询价。如需按算力挑选,可关注 Cerebras WSE-3 Turbo(FP16 250,000 TFLOPS)。

厂商
# 型号 厂商 家族 制程 发布 显存 带宽 FP16 (TFLOPS) FP8 (TFLOPS) FP4 (TFLOPS) 功耗 互联 标价 官网
1 NVIDIA Rubin (Vera Rubin) NVIDIA Rubin Not disclosed by NVIDIA 2026-08 288 GB 19.2 TB/s 4,000 17,500 35,000 — NVLink 6 (3 TB/s per GPU) + Vera CPU NVLink-C2C (1.8 TB/s) 未公开 规格页
2 NVIDIA B300 (Blackwell Ultra) NVIDIA Blackwell Ultra TSMC 4NP 2025 288 GB 8 TB/s 2,500 5,000 15,000 1400 W NVLink 5 (1.8 TB/s) 未公开 规格页
3 NVIDIA GB200 (Blackwell) NVIDIA Blackwell TSMC 4NP 2025 372 GB 16 TB/s 5,000 10,000 20,000 2700 W NVLink 5 (1.8 TB/s) + Grace CPU C2C (900 GB/s) 未公开 规格页
4 NVIDIA B200 NVIDIA Blackwell TSMC 4NP 2025 180 GB 8 TB/s 2,500 5,000 10,000 1000 W NVLink 5 (1.8 TB/s) 未公开 规格页
5 NVIDIA H200 NVIDIA Hopper TSMC 4N 2024 141 GB 4.8 TB/s 989 1,979 — 700 W NVLink 4 (900 GB/s) 未公开 规格页
6 NVIDIA H100 NVIDIA Hopper TSMC 4N 2022 80 GB 3.35 TB/s 989 1,979 — 700 W NVLink 4 (900 GB/s) $30000 规格页
7 NVIDIA A100 80GB NVIDIA Ampere TSMC 7nm 2020 80 GB 2 TB/s 312 — — 400 W NVLink 3 (600 GB/s) $18000 规格页
8 NVIDIA RTX PRO 6000 Blackwell NVIDIA Blackwell TSMC 4NP 2025 96 GB 1.79 TB/s 360 720 1,440 600 W PCIe 5.0 $8500 规格页
9 NVIDIA RTX 4090 NVIDIA Ada Lovelace TSMC 4N 2022 24 GB 1.01 TB/s 165 660 — 450 W PCIe 4.0 $1600 规格页
10 NVIDIA RTX 5090 NVIDIA Blackwell TSMC 4NP 2025 32 GB 1.79 TB/s 209 838 3,352 575 W PCIe 5.0 $2000 规格页
11 AMD MI455X AMD Instinct Not disclosed by AMD 2026-07 432 GB 23.3 TB/s 5,030 20,130 40,260 — UALink over Ethernet (3.6 TB/s bidirectional scale-up per GPU) plus 2,400 Gb/s scale-out 未公开 规格页
12 AMD MI355X AMD Instinct TSMC N3 / N6 2025-06 288 GB 8 TB/s 2,500 5,000 10,100 1400 W Infinity Fabric (7x 153 GB/s scale-up + 1x 128 GB/s scale-out) 未公开 规格页
13 AMD Instinct MI350P AMD Instinct TSMC N3 / N6 2026-05 144 GB 4 TB/s 1,150 2,300 4,600 600 W PCIe 5.0 x16 (no GPU-to-GPU Infinity Fabric links) 未公开 规格页
14 AMD MI325X AMD Instinct TSMC N5 / N6 2024-Q4 256 GB 6 TB/s 1,300 2,600 — 1000 W Infinity Fabric (8x 153 GB/s) 未公开 规格页
15 AMD MI300X AMD Instinct TSMC N5 / N6 2023-Q4 192 GB 5.3 TB/s 1,300 2,600 — 750 W Infinity Fabric $15000 规格页
16 Google TPU7x (Ironwood) Google TPU Not disclosed by Google 2026-03 192 GB 7.38 TB/s 2,307 4,614 — — ICI (1.2 TB/s bidirectional per chip), 9,216-chip pods 未公开 规格页
17 Google TPU v6e (Trillium) Google TPU Not disclosed by Google 2024-12 32 GB 1.64 TB/s 918 918 — — ICI (800 GB/s bidirectional per chip), 256-chip pods 未公开 规格页
18 Google TPU v5p Google TPU TSMC 3nm 2023-12 95 GB 2.77 TB/s 459 459 — 700 W ICI (Inter-Chip Interconnect, 1.2 TB/s bidirectional per chip) 未公开 规格页
19 Google TPU v5e Google TPU TSMC 5nm 2023-08 16 GB 0.8 TB/s 197 393 — 170 W ICI (400 GB/s bidirectional per chip) 未公开 规格页
20 AWS Trainium 2 AWS Trainium TSMC 5nm 2024-12 96 GB 2.9 TB/s 667 1,299 — 500 W NeuronLink-v3 (1.28 TB/s per chip) 未公开 规格页
21 AWS Trainium3 AWS Trainium TSMC 3nm 2025-12 144 GB 4.9 TB/s 671 2,517 2,517 — NeuronLink-v4 (2.56 TB/s per chip) 未公开 规格页
22 AWS Inferentia 2 AWS Inferentia TSMC 5nm 2023-04 32 GB 0.82 TB/s 190 — — — NeuronLink 未公开 规格页
23 Apple M4 Max Apple Apple Silicon TSMC 3nm 2024-10 128 GB 0.55 TB/s 38 — — 60 W On-die (unified memory) $4699 规格页
24 Cerebras WSE-3 Turbo Cerebras WSE Not disclosed by Cerebras 2026-08 44 GB 43200 TB/s 250,000 — — — On-wafer fabric (53.5 PB/s), 2.4 Tb/s off-wafer I/O 未公开 规格页
25 Cerebras WSE-3 Cerebras WSE TSMC 5nm 2024 44 GB 21000 TB/s 125,000 — — 23000 W On-wafer (no interconnect needed) 未公开 规格页
26 Groq LPU Groq LPU GlobalFoundries 14nm (gen 1) 2023 0 GB 80 TB/s 188 — — 350 W GroqLink 未公开 规格页
共 26 款。FP16/FP8/FP4 为厂商公布的密集(dense)算力;「未公开」表示厂商未公布官方定价。按厂商筛选仅作用于当前页面显示,不影响搜索引擎收录的完整列表。
规格与选型说明

NVIDIA Rubin (Vera Rubin):Blackwell successor. 288GB HBM4 per GPU. NVIDIA began production shipments in August 2026. FLOPS are NVIDIA dense figures (FP8 is FP8/FP6 training, FP4 is NVFP4 training); the 50 PFLOPS NVFP4 inference headline is a sparse figure and is not used here.

NVIDIA B300 (Blackwell Ultra):Mid-cycle Blackwell refresh: 208 billion transistors across two dies, with NVIDIA quoting 1.5x the HBM3e capacity and 1.5x the dense NVFP4 rate of Blackwell. Dense figures shown; NVIDIA rates the NVL72 rack at 1,440 PFLOPS NVFP4 with sparsity and 1.1 exaFLOPS dense.

NVIDIA GB200 (Blackwell):Blackwell flagship: dual B200 + Grace CPU. FP4 native. The chip every major lab is buying for 2025-2026 frontier training.

NVIDIA B200:Single-die Blackwell. 180GB HBM3e per GPU (1,440GB across an 8-GPU DGX B200). Inference workhorse for serving 100B+ MoE models on a single GPU.

NVIDIA H200:Hopper refresh with HBM3e. The sweet spot for production inference in 2025. Extra memory means 70B-class models fit single-GPU at FP8.

NVIDIA H100:The chip that built the 2023-2024 LLM boom. Still the most-deployed AI GPU. FP8 native. 80GB HBM3.

NVIDIA A100 80GB:Workhorse of the late-2010s ML wave. Still widely used for training smaller models and inference where H100/H200 supply is tight.

NVIDIA RTX PRO 6000 Blackwell:Workstation Blackwell. 96GB GDDR7. Good fit for on-prem agent dev environments where multi-H100 rentals are overkill.

NVIDIA RTX 4090:Best price/perf for local agent dev. 24GB VRAM fits Q4-quantized 70B models. The default consumer-tier choice for self-hosted Ollama/llama.cpp.

NVIDIA RTX 5090:Consumer Blackwell flagship. 32GB GDDR7. FP4 native. Strong pick for local agent dev with current-frontier features.

AMD MI455X:AMD's MI400-series flagship, launched at Advancing AI on July 23, 2026 with Helios racks already in production. 432GB HBM4. Per-GPU figures are MXFP4, MXFP8 and dense BF16 peaks; AMD publishes the rack numbers of 2.9 exaflops FP4, 1.4 exaflops FP8, 31TB HBM4 and 1.7 PB/s across 72 GPUs. OpenAI expects to bring Helios online starting in Q4 2026.

AMD MI355X:AMD's MI350-series flagship. 288GB HBM3e with native MXFP4 and MXFP6 support. 1400W typical board power. The step between MI325X and the Helios-rack MI400 series.

AMD Instinct MI350P:CDNA 4 in a standard PCIe slot: 128 compute units, roughly half an MI350X, with 144GB HBM3e and a configurable 450W power mode. Peak dense figures shown; AMD doubles its 8-bit and 16-bit ratings with structured sparsity. Built to drop inference capacity into existing air-cooled enterprise servers instead of liquid-cooled racks.

AMD MI325X:AMD's answer to H200. 256GB HBM3e (highest in the catalog at single-die). Strong inference value when supply is constrained on NVIDIA.

AMD MI300X:Wider availability than NVIDIA flagship; cheaper to rent. ROCm software stack matures; vLLM and PyTorch first-class on MI300X in 2026.

Google TPU7x (Ironwood):Seventh-generation TPU and the first in the Ironwood family. 192 GiB HBM per chip, six times Trillium. Google positions it for large-scale training and inference of LLMs, MoE, and diffusion models.

Google TPU v6e (Trillium):Sixth-generation TPU. Twice the per-chip BF16 compute of v5p with a third of the memory; Google used it to train Gemini 2.0.

Google TPU v5p:Google's previous-generation training TPU, now behind Trillium and Ironwood. Used for Gemini training. Only available on GCP; pod-scale (8960 chips) is a different beast than per-chip rentals.

Google TPU v5e:TPU inference tier. The 8-bit figure is Google's INT8 rating (393 TOPS). Cheaper than v5p; 256-chip pods. Best fit for Google Cloud customers serving Gemini-style inference at scale.

AWS Trainium 2:AWS custom training/inference silicon. Dense figures shown; AWS rates it at 2,563 TFLOPS with sparsity. Anthropic uses these for Claude training. Cheaper TCO on AWS than NVIDIA flagship for committed workloads.

AWS Trainium3:Third-generation AWS training and inference silicon. 144 GB HBM3e, 4.9 TB/s bandwidth. 8-bit and 4-bit figures are MXFP8 and MXFP4 per AWS; BF16 is roughly flat with Trainium 2. Roughly 30-50% lower cost per hour than H100/H200 on committed AWS workloads.

AWS Inferentia 2:AWS inference silicon. Lower compute than Trainium 2 but cheaper per-token for high-volume serving workloads on AWS.

Apple M4 Max:128GB unified memory means a 2024-2025 Mac Studio or MacBook Pro runs 70B-class models at Q4 quantization with usable speed. The dark-horse local-inference platform. Apple publishes memory and bandwidth for Apple Silicon but no TFLOPS rating for the M5 generation, so no M5 row is listed here.

Cerebras WSE-3 Turbo:Announced August 18, 2026. Same 46,225 mm^2 wafer, 4 trillion transistors and 900,000 cores as WSE-3, but Cerebras doubles AI compute to 250 PFLOPS and memory bandwidth to 43.2 PB/s (shown here as 43,200 TB/s). A CS-4 rack holds three wafers for 750 PFLOPS. Compute figure is on the same basis as the 125 PFLOPS FP16 rating of WSE-3.

Cerebras WSE-3:Single 46,225 mm^2 wafer. Holds entire activations on-chip; eliminates many distributed-training pains. Cerebras rates aggregate memory bandwidth at 21 PB/s (shown here as 21,000 TB/s).

Groq LPU:Deterministic single-thread inference silicon. SRAM-only memory model means small per-chip capacity but massive bandwidth. Behind Groq's 700+ tokens/sec Llama 4 Scout serving.

相关频道
数据说明
数据来源:第三方公开聚合服务 TensorFeed,汇总各芯片厂商公开的产品规格页。
同步频率:每 3 小时同步一次,规格变更后自动更新,无需人工维护。
口径说明:显存带宽单位 TB/s,算力单位 TFLOPS;TPU/LPU 等非 GPU 架构的算力口径与 NVIDIA 不完全可比,跨厂商对比时请以显存容量与带宽为主。
免责:本页为公开资料汇总,仅供选型参考,最终规格以厂商官方文档为准。
常见问题
加速卡规格库收录了哪些芯片?
目前收录 26 款,覆盖 NVIDIA(H100/H200/B200/GB200/B300/Rubin、RTX 系列)、AMD(MI300X/MI325X/MI350/MI355X/MI455X)、Google TPU(v5e/v5p/v6e/Ironwood)、AWS Trainium/Inferentia、Apple M4 Max、Cerebras WSE-3 与 Groq LPU 等。
显存带宽为什么比算力更值得关注?
大模型推理与训练常受显存容量与带宽限制(memory-bound),尤其是长上下文与 MoE 场景。带宽决定数据搬运速度,容量决定单卡能装下的模型规模;算力再高,显存装不下也无法运行。
「标价」为空代表什么?
代表厂商未公开该型号的官方建议零售价(如数据中心 GPU 通常只对 ODM/云厂商报价)。这类芯片的实际成本需参考云厂商的租用价格,可跳转到「GPU 租用价格对比」查看。
制程「Not disclosed by NVIDIA」是什么意思?
表示厂商未公布该型号的具体制程工艺节点。本页如实展示数据源内容,不做推测性补全。
讨论