Hi guys I keep seeing people talk about ninfer, so I wanted to know if switching from llama.cpp is actually worth it. This was the prompt that I was using (physics, spin, full rules, the works) Setup: RTX 5090, Qwen3.8 27B, thinking on (xhigh), 120k context, default sampling settings, one run oneshot Speed Setup Output tokens Time Decode tokens/s ninfer, [precision of the non-NVFP4 build], MTP 67,539 7m 40s ~147 llama.cpp Q4_K_M + MTP (draft-n-max 3) 69,749 8m 13s ~141 ninfer, NVFP4, MTP 93,663 10m 3s ~155 llama.cpp Q4_K_M, no MTP 78,976 19m 31s ~68 A few things stood out. Stock llama.cpp without MTP is less than half as fast as ninfer. But once you turn on MTP in llama.cpp it jumps from 68 to 141 t/s and lands very close to ninfer, so a big part of the "ninfer is fast" story is really "MTP is fast". NVFP4 had the highest t/s, but it also wrote the most tokens (mostly thinking), so it only finished third on wall-clock time. For reasoning models I'd look at time-to-result, not just t/s. Quality I checked all four games with a script that fires about 2,400 random shots (random angle, power and spin) at each one, plus a few scripted rule scenarios. Good news: none of them crashed, produced NaNs or got stuck, so all four run. The differences are in the rules: ninfer NVFP4 ninfer [non NVFP4] llama Q4_K_M llama Q4_K_M + MTP Can you legally win by potting the 8? yes no stripes only no 8-ball on the break respotted re-rack counts as a loss counts as a loss Starting rack OK? yes yes balls overlap yes Sound no yes no yes Lines of code 933 1185 1127 965 All four run fine, but only the NVFP4 game can actually be won. The other three have small logic bugs in the win condition (and one has a broken starting rack), so none of them is quite finished. You can try them yourself: Qwen 3.8 27B ninfer NVFP4: https://claude.ai/artifact/1oj8LJkBRrLmHhQLSe9kCu Qwen 3.8 27B ninfer [non NVFP4]: https://claude.ai/artifact/WbW2XSmKDfxgAArKGEeihC Qwen 3.8 27B llama.cpp Q4_K_M: https://claude.ai/artifact/DqYMjSR7unZ8kicr3JKnSg Qwen 3.8 27B llama.cpp Q4_K_M + MTP: https://claude.ai/artifact/2sunpBgBJBRvbaJAJnkM3t Keep this in mind before you trust my numbers: One run per setup, so some of the bugs could just be bad luck. Everything ran on the default reasoning effort (xhigh), which inflates the token counts. NVFP4 and Q4_K_M are different quant schemes, so don't treat them as equivalent. My take: I'm sticking with the non-NVFP4 ninfer build for my next round of prompts. Of the four games, that one was my favorite to actually play. It had the most polish: sound, realistic ball size, the break rules, a proper kitchen for ball-in-hand. The only thing that bugged me is that you can't win a game legally, because potting the 8 after clearing your group counts as a foul. Funny enough, it turned out to be a one-line bug (an inverted check), so it was really close to being the best of the bunch.   submitted by   /u/theexile1337 [link]   [comments]