B200.
NVIDIA B200-SXM-180GB · 180 GB · SXM
Logical ID accelerator:nvidia-b200-sxm-180-gb
Content version c98239344525428cadbbc1836c8e7e2eabb1ca7869bdd6f23993172939dc8d11
Submitted systems
Dell · XE9780_x86_B200-SXM-180GBx8_TRT
1 nodes × 8 accelerators = 8 total
Topology: 18x 5th Gen NVLink, 14.4 TB/s aggregated bandwidth
Software: TensorRT 10.14, CUDA 13.1
deepseek-r1 · Offline: 58,235.4 tokens/s · source
deepseek-r1 · Server: 40,949.5 tokens/s · source
Google · B200-SXM-180GBx8_TRT
1 nodes × 8 accelerators = 8 total
Topology: 18x 5th Gen NVLink, 14.4 TB/s aggregated bandwidth
Software: TensorRT 10.14, CUDA 13.1
deepseek-r1 · Offline: 57,116.9 tokens/s · source
deepseek-r1 · Server: 41,027.9 tokens/s · source
HPE · HPE_ProLiant_XD685_B200_SXM_192GBx8_TRT
1 nodes × 8 accelerators = 8 total
Topology: 18x 5th Gen NVLink, 14.4 TB/s aggregated bandwidth
Software: TensorRT 10.11, CUDA 13.0
llama3.1-8b · Interactive: 126,143 tokens/s · source
llama3.1-8b · Offline: 152,625 tokens/s · source
llama3.1-8b · Server: 131,270 tokens/s · source
Lambda_SIT · B200-SXM-180GBx8_TRT
1 nodes × 8 accelerators = 8 total
Topology: 18x 5th Gen NVLink, 14.4 TB/s aggregated bandwidth
Software: TensorRT 10.14, CUDA 13.1
llama3.1-8b · Interactive: 128,750 tokens/s · source
llama3.1-8b · Offline: 160,403 tokens/s · source
llama3.1-8b · Server: 130,008 tokens/s · source
Nebius · nebius_b200_n1
1 nodes × 8 accelerators = 8 total
Topology: 18x 5th Gen NVLink, 14.4 TB/s aggregated bandwidth
Software: TensorRT 10.14, CUDA 13.1, cuDNN 9.17, TensorRT-LLM feat/1.2-mlpinf, NVIDIA Dynamo mlperf-v6.0-dynamo-v0.8.0, vLLM CentML:mlperf-inf-mm-q3vl-v6.0
deepseek-r1 · Offline: 58,581.6 tokens/s · source
deepseek-r1 · Server: 51,692.9 tokens/s · source
gpt-oss-120b · Interactive: 13,155.3 tokens/s · source
gpt-oss-120b · Offline: 85,921.2 tokens/s · source
gpt-oss-120b · Server: 87,444.2 tokens/s · source
RedHat · 8xB200-LLM-D-Openshift
1 nodes × 8 accelerators = 8 total
Topology: 18x 5th Gen NVLink, 14.4 TB/s aggregated bandwidth
Software: LLM-D v0.5.0 ,vLLM 0.14.1, RHEL 9.6
gpt-oss-120b · Offline: 93,070.7 tokens/s · source
gpt-oss-120b · Server: 71,588.1 tokens/s · source