New Why ionstream exists — the infrastructure teams can depend on Read more

GPU Compute · Blackwell

NVIDIA B200 — redefining AI and HPC with one of the most advanced GPUs yet

192GB of HBM3e and 8 TB/s of memory bandwidth deliver up to 15X faster real-time inference than NVIDIA Hopper — reserve dedicated B200 power, in limited supply.

Pricing starts at $2.69 per GPU/hr

  • Generative AI
  • Training
  • Real-time inference
  • HPC
NVIDIA B200 GPU

Specifications

Built for the largest models

Massive memory and bandwidth, with a dedicated decompression engine for data-heavy pipelines.

GPU memory 192GB HBM3e
Memory interface 2x 4096-bit
Memory bandwidth 8 TB/s
Decompression LZ4 · Snappy · Deflate

Server configuration

NVIDIA HGX B200 GPU Server (Intel)

A fully integrated 8-GPU node — CPU, memory, storage and networking provisioned as one dedicated server.

GPUs
8x NVIDIA B200 SXM (HGX B200), 1,440 GB total HBM3e, NVLink-connected
CPU
2x Intel Xeon 6972P — 192 cores total
Memory
1.5 TB DDR5
Local storage
8x 3.84 TB NVMe SSD (~30 TB)
GPU fabric
8x NVIDIA ConnectX-7 InfiniBand, up to 400 Gb/s each (3.2 Tb/s aggregate) for multi-node scaling
Data / storage network
2x 100 GbE
Internet uplink
Dual 100 Gbps — two independent 100GbE uplinks at the DC edge deliver redundant, high-throughput public connectivity to our upstream providers, with no single point of failure.
Management
Dedicated 10 GbE + IPMI/BMC

Performance

A generational leap over Hopper

Compared against NVIDIA Hopper and H100, the B200 accelerates both training and inference.

15X
Real-time inference vs Hopper
3X
Training speed vs Hopper
6X
Data analytics vs H100

Also up to 18X faster query performance for data analytics versus traditional CPUs.

Use cases

What the B200 powers

From frontier-model training to real-time production inference and HPC.

01

Generative AI

Train and serve large language and multi-modal models with headroom for the biggest context windows.

02

Training & inference

End-to-end machine learning — from large-scale training runs to low-latency production inference.

03

Real-time inference

Computer vision, NLP and recommendation systems served in real time at scale.

04

Data analytics

Up to 18X faster query performance versus traditional CPUs for analytics and predictive modeling.

05

High-performance computing

Simulation, modeling and scientific workloads that demand maximum throughput.

06

Large-scale data processing

A dedicated decompression engine accelerates data pipelines feeding your GPUs.

Reserve your NVIDIA B200 access

Supply is limited. Talk to our team to secure dedicated B200 capacity for your training and inference workloads.