So here is my server with four MI100, each with 32 GB of VRAM, with their infinity fabric bridge directly connecting all of them. MI100 is not well supported out of box in most software, so I made a fork of vLLM to get the performance I was hoping for. VLLM was running at 15 tok/s stock on this machine and now it runs at over 1000 tok/s decode and over 5000 tok/s prefill at 8 concurrency. This level of performance took tons of new kernels and optimization techniques. The list of features I added from scratch for this arch is huge. About half of the features are useful for other int8-oriented architectures, so this fork is an excellent starting point or reference for other architectures. I've had a report that someone with a 2x MI210 server also was getting over 1k tok/s on my fork with minor changes. I can say I'm now satisfied with where its at, as it outperforms the public APIs by a wide margin and is a pleasure to vibe code with. It replaces about $80 per day of API costs, so I feel like if you know what you're doing, MI100s are a great deal for mid range setups. The infinity fabric bridge (aka XGMI) is a great and unique feature at this price range, giving much lower latency and much higher bandwidth during accumulation steps than PCIe. Here's my fork, along with custom version of AITER and custom 27B checkpoints (int8 group size 128 with rounding fine tuning for the numerics of the platform -- PTQR). https://github.com/curvedinf/int8-vllm Qwen 27B has been flakey on vllm since it came out on many architectures, with weird garble issues especially at long context, so most recently I spent a long time tracking down the bugs in upstream vLLM, AITER, and rocm-systems. I found and fixed something like 30 of these numerics and memory tracking issues. After testing for months now at high context, unsynced concurrency, and KLD at every block in the whole model, the quality is now great and reliability is excellent. I run it on a low power profile at C6 with full 256k contexts for all streams, and this is used by my hermes agent, business process workers, and as backups for my coding agents. It uses 540 watts at the wall and is a bit louder than a desktop. I also use the server for ML research and model training. Couldn't be more pleased as I originally bought this server for a bit over $6k, and looking at prices now it seems like this was a great decision. ☠️ Anyway, this was a lot of work but it was fun to learn about LLMs at this depth and I'd do it again if given the chance. I hope the fork will come in handy for some people with similar builds.   submitted by   /u/1ncehost [link]   [comments]