We built an engine, Kyojin, on top of ExLlamaV3 for Strix Halo (gfx1151, ROCm), and packed two 300B-class MoE models so each fits one 128 GB machine. First release, all measured on Ryzen AI Max+ 395. Model GLM-5.3-Flash MiMo-V2.6-Flash-MOPD Size 99.7 GB 105 GB Prefill 580 tok/s at 3.5K, 546 at 64K about 650 tok/s at 4K Decode 26 to 30 tok/s (MTP) 32 prose / 35 chat / 44 code (speculative), 29 plain KLD vs official FP8 0.151 0.0713 Top-1 agreement with FP8 89.3 % 92.0 % Where the weights come from. MiMo is our own quantisation. The GLM pack mixes turboderp's public 2.05 and 3.05 bpw EXL3 tensors, with our layer mix and a small tuning stage. On the same 129 rows, his 2.05 bpw pack (85 GB) gets KLD 0.275; our mix (100 GB) gets 0.190. His is smaller and decodes about 10 % faster. Uncensored variants. Separate -Uncensored repos: same weights plus one small file the engine applies at load, one switch turns it off. Not measured yet. Task-suite scores for MiMo, GLM at 128K context, any GPU other than gfx1151. The conversion pipeline stays private. Quickstart. Clone, ./build.sh, hf download yamz-labs/GLM-5.3-Flash-EXL3-Yamz, python tools/glm/serve.py --model ./glm-pack -c 131072 --num-draft 2. You get an OpenAI-style API. Models: https://huggingface.co/yamz-labs Engine: https://github.com/Yamz-Labs/kyojin Built on turboderp's ExLlamaV3, with ROCm work from sdougbrown and vcruz305. If you own a Strix Halo machine, we'd love to see your tok/s. Issues, benchmarks and PRs are all welcome. Which model should we do next?   submitted by   /u/Yaniss916 [link]   [comments]