MINERSDEALSMinersDeals

NVIDIA

NVIDIA H200 141GB HBM3e AI GPU

141 GB of HBM3e at 4.8 TB/s — the memory-bandwidth workhorse for LLM training and high-throughput inference.

Memory
141 GB HBM3e
Bandwidth
4.8 TB/s
TDP
700 W SXM / 600 W NVL
Architecture
Hopper
FP8 (sparse)
Up to 1,979 TFLOPS
Interconnect
NVLink 900 GB/s
Talk to Sales

Need 10, 50 or 100+ units? Request volume pricing.

NVIDIA H200 data center AI GPU accelerator

What Is the NVIDIA H200?

The NVIDIA H200 is a Hopper-architecture data center GPU with 141 GB of HBM3e memory and 4.8 TB/s of memory bandwidth. It ships as an SXM module (up to 700 W, for HGX baseboards) and as the PCIe H200 NVL (configurable up to 600 W), and is used primarily for large language model training, fine-tuning and high-throughput inference.

The H200's defining feature is memory. At 141 GB per GPU with 4.8 TB/s of bandwidth, it holds substantially larger model shards and KV caches than the H100 it succeeds, which translates directly into higher inference throughput per GPU and fewer GPUs needed to serve a given model.

Two variants matter when specifying a build. The SXM module goes onto an HGX baseboard with NVSwitch, giving full 900 GB/s NVLink fabric between eight GPUs — the right choice for training. The H200 NVL is a dual-slot PCIe card with NVLink bridges, air-cooled and configurable from 400 to 600 W, designed to drop into mainstream servers for inference and mixed workloads.

Published tensor figures — 989 TFLOPS FP16 and 1,979 TFLOPS FP8 on SXM — are with sparsity; dense values are roughly half. Hopper has no FP4 tensor path, so workloads targeting FP4 need Blackwell-generation hardware instead.

NVIDIA H200 Specifications

NVIDIA H200 specifications
ManufacturerNVIDIA
ArchitectureHopper
GPU memory141 GB HBM3e
Memory bandwidth4.8 TB/s
FP16 tensor (sparse)989 TFLOPS (SXM) / 835 TFLOPS (NVL)
FP8 tensor (sparse)1,979 TFLOPS (SXM) / 1,671 TFLOPS (NVL)
FP4 tensorNot supported on Hopper
TDP700 W (SXM) / up to 600 W configurable (NVL)
Form factorSXM5 module or dual-slot PCIe (NVL)
InterconnectNVLink 900 GB/s; PCIe Gen5 host interface on NVL
CoolingAir or liquid depending on chassis
AvailabilityAnnounced Nov 2023; broadly available from 2024

H200 Power Consumption & Electricity Cost

Rated wall power for the NVIDIA H200 is not published in the sources we track. Contact sales for the current manufacturer specification sheet.

A verified rated wattage for the NVIDIA H200 is not published by the manufacturer in the sources we track, so we do not publish an electricity-cost estimate for it. Contact sales for the current specification sheet.

H200 Power, Cooling & Rack Planning

An eight-GPU HGX H200 baseboard alone draws around 5.6 kW before CPUs, memory, storage and fans, so a populated node commonly lands between 8 and 10 kW. That is well beyond the 3–5 kW per rack many older facilities were designed for, and it is the single most common blocker in H200 deployments.

Air cooling works at these densities only with high-static-pressure chassis fans, strict hot-aisle containment and cold-aisle inlet temperatures held low. Above roughly 30 kW per rack, direct-to-chip liquid cooling is usually the practical route, and rear-door heat exchangers are a common middle path for retrofits.

The PCIe H200 NVL is explicitly designed for air-cooled mainstream servers with a configurable 400–600 W TDP, which lets integrators trade peak performance for thermal headroom in racks that cannot support SXM densities.

Why Consider the NVIDIA H200?

141 GB per GPU

Fits larger models and longer context windows on fewer GPUs, cutting both node count and inter-GPU communication overhead.

4.8 TB/s bandwidth

Memory bandwidth is the binding constraint for token generation; the H200 raises it substantially over prior Hopper parts.

Two deployment paths

SXM for maximum training performance on HGX, or PCIe NVL for air-cooled inference nodes in standard racks.

Who Is the H200 Best For?

  • LLM training and fine-tuning clusters
  • High-throughput inference and serving
  • Retrieval-augmented generation with large caches
  • HPC and scientific workloads needing HBM bandwidth

NVIDIA H200 Price

Pricing for the NVIDIA H200 can vary depending on quantity, availability, market conditions and fulfillment requirements. MinersDeals provides individual and volume quotes for qualified buyers, so we publish current pricing on request rather than as a fixed list price.

Buying multiple units? Ask about volume pricing.

Buy NVIDIA H200

Looking to purchase the NVIDIA H200? MinersDeals helps individual buyers, mining operators and enterprise customers source specialized hardware. Submit your requirements to receive current availability and pricing.

Wholesale & Bulk H200 Orders

Planning a larger deployment? Request pricing for 10, 50, 100 or more units. MinersDeals supports hardware procurement for mining farms, hosting providers and infrastructure operators, including staged delivery across multiple sites.

Frequently Asked Questions

How much memory does the NVIDIA H200 have?

141 GB of HBM3e with 4.8 TB/s of memory bandwidth.

What is the difference between H200 SXM and H200 NVL?

SXM is a 700 W module for HGX baseboards with full 900 GB/s NVLink fabric. H200 NVL is a dual-slot PCIe card, air cooled, with a configurable TDP up to 600 W and NVLink bridges for 2- or 4-way pairing.

Does the H200 support FP4?

No. Hopper has no FP4 tensor path. FP4 workloads require Blackwell-generation hardware.

How much power does an eight-GPU H200 node need?

The GPUs alone account for about 5.6 kW; a complete node typically lands between 8 and 10 kW.

How much does an NVIDIA H200 cost?

Enterprise GPU pricing depends on configuration, quantity and availability. MinersDeals quotes current pricing for single units, HGX nodes and full clusters on request.

Explore Related Categories

Get H200 Pricing from MinersDeals

Tell us the quantity, destination and deployment timeline, and our team will come back with current availability and a quotation.