Opening the index…
Resolving the selected comparison slice.
Resolving the selected comparison slice.
NVIDIA RTX PRO 6000 Blackwell Server Edition · 96 GB · Blackwell Server Edition
Logical ID accelerator:nvidia-rtx-pro-6000-blackwell-server-edition-96-gb
Content version 34f5470880ba60d5dbae3f82d7cbd345778be7ef79047ebf24231631c9e6f069
1 nodes × 8 accelerators = 8 total
Topology: PCIe Gen5 x16
Software: TensorRT 10.14, CUDA 13.1
llama3.1-8b · Offline: 48,613.8 tokens/s · source
llama3.1-8b · Server: 47,581 tokens/s · source
1 nodes × 8 accelerators = 8 total
Topology: PCIe Gen5 x16
Software: TensorRT 10.14, CUDA 13.1
llama3.1-8b · Interactive: 34,241 tokens/s · source
llama3.1-8b · Offline: 48,791 tokens/s · source
llama3.1-8b · Server: 48,531.3 tokens/s · source
1 nodes × 4 accelerators = 4 total
Topology: N/A
Software: TensorRT 10.11, CUDA 12.9
llama3.1-8b · Offline: 24,013.5 tokens/s · source
llama3.1-8b · Server: 23,423.9 tokens/s · source
1 nodes × 8 accelerators = 8 total
Topology: 18x 5th Gen NVLink, 14.4 TB/s aggregated bandwidth
Software: TensorRT 10.14, CUDA 13.0
llama3.1-8b · Offline: 49,036.1 tokens/s · source
llama3.1-8b · Server: 45,008.6 tokens/s · source
1 nodes × 10 accelerators = 10 total
Topology: N/A
Software: TensorRT 10.14.1.48, CUDA 13.0
gpt-oss-120b · Offline: 19,040.8 tokens/s · source
gpt-oss-120b · Server: 17,737.6 tokens/s · source
llama3.1-8b · Interactive: 53,054.7 tokens/s · source
llama3.1-8b · Offline: 60,679.9 tokens/s · source
llama3.1-8b · Server: 59,632.7 tokens/s · source
1 nodes × 8 accelerators = 8 total
Topology: N/A
Software: TensorRT 10.14.1.48, CUDA 13.0
gpt-oss-120b · Offline: 15,189.9 tokens/s · source
gpt-oss-120b · Server: 14,258.9 tokens/s · source
llama3.1-8b · Interactive: 44,087.3 tokens/s · source
llama3.1-8b · Offline: 49,580.2 tokens/s · source
llama3.1-8b · Server: 48,794.3 tokens/s · source
1 nodes × 2 accelerators = 2 total
Topology: N/A
Software: TensorRT 10.11.0.33, CUDA 13.0
llama3.1-8b · Interactive: 10,459.6 tokens/s · source
llama3.1-8b · Offline: 12,433.3 tokens/s · source
llama3.1-8b · Server: 11,898.8 tokens/s · source
1 nodes × 2 accelerators = 2 total
Topology: N/A
Software: TensorRT 10.14.1.48, CUDA 13.0
llama3.1-8b · Interactive: 10,589.3 tokens/s · source
llama3.1-8b · Offline: 12,495.2 tokens/s · source
llama3.1-8b · Server: 11,899.7 tokens/s · source
1 nodes × 4 accelerators = 4 total
Topology: N/A
Software: TensorRT 10.11.0.33, CUDA 13.0
llama3.1-8b · Interactive: 19,991.1 tokens/s · source
llama3.1-8b · Offline: 23,265.7 tokens/s · source
llama3.1-8b · Server: 22,916.3 tokens/s · source
1 nodes × 4 accelerators = 4 total
Topology: N/A
Software: TensorRT 10.14.1.48, CUDA 13.0
gpt-oss-120b · Offline: 7,297.05 tokens/s · source
gpt-oss-120b · Server: 6,687.25 tokens/s · source
llama3.1-8b · Interactive: 20,118.9 tokens/s · source
llama3.1-8b · Offline: 23,556.7 tokens/s · source
llama3.1-8b · Server: 23,227.2 tokens/s · source