Disclosure: I run Kitani.AI, an open model inference provider. Hey guys so we've been experimenting a lot with AMD GPUs lately, and I'm starting to think the "gap" everyone says between AMD and NVIDIA for inference is a lot smaller than people assume once the software stack is actually optimized. We've been working on optimized kernels and serving configs for some of the newer MoE/open models. The economics have gotten good enough that we're currently serving MiMo V2.6 below Xiaomi's own standard API pricing: MiMo V2.6 Pro: $0.425/M input $0.825/M output MiMo V2.6 Flash: $0.125/M input $0.25/M output We've also been experimenting with GLM 5.3 Flash/Uncensored and other sparse models, where the relatively small number of active parameters makes the hardware economics especially interesting. After the testings etc, it wasn't just the tok/s that was suprising. Memory capacity/bandwidth + newer ROCm kernels can make AMD extremely competitive on cost per generated token especially when you're batching instead of optimizing purely for a single user's TPS. Obviously NVIDIA still has the much more mature ecosystem and there are workloads where CUDA is just easier. But for people actually running production inference: has anyone else seriously tested MI300X/MI355X against H100/H200/B200 lately? It's definitely a lot cheaper and can make a inference platform a lot more profitable easily. Whats your guys real cost/token and throughput looked like after optimization, rather than just comparing GPU hourly rental prices.   submitted by   /u/bakatristan [link]   [comments]