TL;DR -- You should probably just use base Qwen3.6-35B, as only Occamy-1.0 is competitive with it. Tiel is a major let-down, worse than Ornith. KAT surprises (good), Nex surprises (bad). This post is long. Sorry, lots to cover. I think we all want to see a next-generation small MoE from the Qwen team to replace 3.6-35B in our workflows. This model is a perfect fit for smaller gmaing laptops and mid-tier rigs. It sucks that Qwen seems to have abandoned this model, but at least there are fine-tunes that improve upon it... right? Well... maybe not. I ran benchmarks on the 3.6-35B-A3B base model, as well as five finetuunes: Occamy-1.0, Ornith-1.5, KAT-Coder-V2.5-Dev, Tiel-Coder, and Nex-N2.5-mini, and the results are quite surprising. I chose Aider Polyglot for a few reasons: it's an agentic coding benchmark I can run in ~10 hourso on my machine, it's not actively post-trained on by any of these models, and it provides a lot of useful information along with the raw accuracy scores. This includes: first-try and retry pass rates, token counts, solve times, and how well-formed the output diffs are. Here's the table: model First-try pass Retry pass tokens sec/case tok/solve well-formed diff Qwen3.6-35B-A3B BASE (STOCK template) 37.4% 71.0% 8650 285 14.1K 96.3% Occamy-1.0-35B-A3B (STOCK template) 29.0% 70.1% 6801 285 17.2K 86.9% Occamy-1.0-35B-A3B (froggeric medium) 30.8% 69.2% 6009 233 16.8K 94.4% Occamy-1.0-35B-A3B (froggeric, xhigh) 27.1% 67.3% 8631 310 20.0K 91.6% Ornith-1.5-35B-A3B 23.4% 63.6% 4813 226 16.4K 87.9% KAT-Coder-V2.5-Dev 20.6% 58.9% 2190 84 9.3K 86.9% Tiel-Coder-35B-A3B 18.7% 53.3% 4851 171 18.2K 89.7% Nex-N2.5-mini 10.3% 30.8% 5037 188 33.3K 95.3% As you can see, the only finetune that even competes with the base model is Occamy-1.0. The rest are utterly dominated by the base model, a grim disappointment for finetune enthusiasts. I was particualarly surprised by the performance of Tiel, which seems to get a lot of love in this subreddit. Speaking of Tiel, I want to clarify that Tiel is just Ornith-1.5 with a different chat template, Sharp, which is based on froggeric with an added "terse mode" instruction that's supposed to reduce excessive verbosity. I wanted to standardize for templates, so ALL models are using the base froggeric v22.5 template set to medium (which is equivalent to standard thinking on, no additional message sent). I used this because I wanted to test Tiel vs. Ornith-1.5, and Tiel is the chat template. Also, practially, I use froggeric in my real workflows. However, to ensure coverage, I also tested the STOCK template on the 2 highest-performing models, to make sure it wasn't affecting the scores. As you can see, the template doesn't make a significant difference in the scores, and the scores for base 35B with different templates are so close to identical that I excluded the froggeric one from the table. I also tested froggeric/Sharp's reasoning-effort toggle, and found xhigh -> medium significantly reduces token counts and solve times (by ~1/3), without affecting accuracy significantly. That stands in stark contrast to Tiel's 'terse mode' toggle, the core feature of Tiel over Ornith, which dramatically reduces accuracy along with the reduction in token counts. My results strongly suggest that if you want a less verbose model, you're better off lowering the reasoning effort than using Tiel with terseness on. Speaking of token use, that's probably the big differentiator here. A couple models stand out: Ornith and KAT-Coder-V2.5-Dev are the most efficient models, with KAT in particular having a brevity unmatched by anything else. KAT is fucking fast, and I think despite its lower accuracy than Occamy, it has a place in my lineup as a subagent because it just gets. shit. done. Occamy is also interesting, as it is the only model that perfoms on a similar level to the base, but it uses 20-30% fewer median tokens. However, Occamy also had a number of runaway generations where the token count blew up, so it's total tokens/solve is actually higher than base. In an effort to further distinguish Occamy from base, since Aider struggled to do that, I ran tau2-bench, an agentic tool-calling benchmark consisting of multi-turn interactions with a simulated counterparty. I figured this was a good bench to use as Occamy is post-trained specifically for 'co-work' scenarios, but not trained on this particular set. I used Qwen3.8-27B with reasoning effort set to low as the simulated customer in these conversations. The base model was able to pull away from Occamy in the harder retail domain of this benchmark, but Occamy resolved the issues in the airline domain at an equal rate while requiring fewer turns. Here's the results. model (Q8_0) domain pass1 tokens sec/task turns/task Occamy-1.0-35B-A3B (STOCK template, temp=1.0, presence=1.5) airline 80.0% 5390 290 11 Qwen3.6-35B-A3B BASE (STOCK template, temp=1.0, presence=1.5) airline 78.0% 6770 390 13 Qwen3.6-35B-A3B BASE (STOCK template, temp=1.0, presence=1.5) retail 86.0% 4098 333 14 Occamy-1.0-35B-A3B (STOCK template, temp=1.0, presence=1.5) retail 79.8% 3284 284 13 Overall, I think the results are clear, if unexpected: Occamy-1.0 is the only fine-tune that even competes with the base model on Aider Polyglot, but even it is not a clear winner. Tiel is noticebly worse than plain Ornith without the terseness toggle, and the terse mode doesn't even save any tokens. xhigh in froggeric/Sharp degrades accuracy slightly and bloats token use, which makes sense given the models were not RL'd for the extra thinking effort prompt. KAT-Coder-V2.5-Dev is the most efficient model, with accuracy nearly as good as Ornith and better than Tiel. Finally, Nex-N2.5-mini is a disaster.   submitted by   /u/returnity [link]   [comments]