If you were building an inference server today, would you buy one H200 or spend the same budget on multiple RTX PRO 6000s? -> The PRO 6000 has 96GB of GDDR7 at 1.792 TB/s (1.6 on the Server Edition). Native FP4, no NVLink. -> The H200 has 141GB of HBM3e at 4.8 TB/s, with NVLink and FP8 as its lowest precision. After speccing both, I think it comes down to fit and interconnect, rather than picking by brand or spec sheet alone. Where the PRO 6000 wins: Single-card and small multi-card inference on models up to roughly 70B at sensible quantisation. Cost per card is a fraction of an H200. Power draw is manageable in a normal rack. Availability is far better. Native FP4 helps on 4-bit models. Where the H200 wins: When inference is memory-bandwidth-bound. Long-context workloads, big models where you don’t want to shard across PCIe, and tensor parallelism, where NVLink between cards actually earns its keep. The extra memory and bandwidth can also help with long-context serving and fine-tuning, depending on the model and workload. Just don't compare raw FLOPS. Decode is usually memory bandwidth bound, not compute bound, so the TFLOPS line on the datasheet tells you very little about tokens/sec.   submitted by   /u/recentheartbroken [link]   [comments]