Methodology.
One question per slice
Every comparison fixes release, division, workload, scenario, accuracy target, metric and unit. Offline measures batch throughput. Server and Interactive have different latency constraints and are never combined. Hardware without evidence in your slice is missing, not zero.
What “verified” means here
We check pinned Git blob identity, SHA-256 provenance, exact accelerator model and memory mapping, explicit node and accelerator counts, topology presence, LoadGen VALID status, duration/query/early stopping flags, performance and accuracy log errors, official accuracy thresholds and checked-in compliance pass artifacts. We do not rerun the benchmarks or claim MLCommons certification of this app. Full upstream submission checks also cover training/calibration, measurements and packaging; this index consumes the published closed-division release.
Identity and system boundaries
Accelerator variants are distinct from systems and results. SXM and NVL variants stay separate. Counts come from number_of_nodes × accelerators_per_node, never from x8 or NVL72 in names. Unlisted models, ambiguous dual boards and partitioned devices go to review-required quarantine.
Ranking and derived values
The default ranks the official submitted system throughput, best result per accelerator variant. All-systems mode preserves every accepted submission. Equal scores share competition rank (1, 1, 3); stable result IDs break display-order ties. Derived mode divides by documented count and is explicitly labeled. It is not a single-card measurement or a scaling claim.
Provenance and updates
Logical IDs identify records across rebuilds. SHA-256 content versions identify immutable snapshots. Each record links to HTTPS source files at a pinned commit with file hashes. Updates are staged, schema validated, reviewed as a diff and promoted atomically. Removed records become tombstones; rollback restores a saved version.
Pinned results repository · Official v6.0 metric and accuracy rules