200 Task Custom Dataset Performance Result: 10 Popular Models, from 2B MoE to 27B Dense

Wait 5 sec.

https://preview.redd.it/8eny8jk3b4th1.png?width=1550&format=png&auto=webp&s=923f42ba5db0fd05e756bebc1799c07174bf58e2 Results are within the screenshot, but here's a TL;DR tierlist: S tier - Gemma-4-26B, even quantized down to iq3s it tops my charts. A tier - Qwen3.8-27B-q3/q5, this one really surprised me, as a dense model it crawls, but didn't snag the S tier slot. Somehow, qwen3.5-9b is also in this slot. B tier - Qwen3.5-4b, also incredibly Gemma-4-e2b, which punches far above its weight. C tier - Gemma-4-12B, Nanbeige, these both are too heavy for their performance, pass. F tier - Ling-3.0-tiny, minicpm, The test questions consisted on tasks that I do every day with my assistants, written by Fable 4.1. "Hive" is the llm cluster I'm working on, involving custom tools and executables called by the models for different functions. Calling (or miscalling) these is important, and running a heavier model than necessary hurt, so here we are, trying to figure out the best of both worlds. Anyway, I thought this was interesting. Hopefully you do too! YMMV.   submitted by   /u/Potential-Net-9375 [link]   [comments]