Two local Qwen ( 3.8 27b unsloth Q6 and Qwen flash next strata coder ) models vs Claude Opus 4.6 on the same 3 coding tasks. One of them tied it. Not here to start a fight, just sharing numbers

Wait 5 sec.

First, two things so nobody reads this the wrong way. I'm not trying to argue that one model or one company is better. I have no stake in any of them. I just wanted to check for myself how they behave on the same work. And the reason I care: I use local models to write code, and I'd like to know how much I can rely on them instead of paying for frontier subscriptions. That's the whole motivation. What I did Three Python tasks, easy to hard: a log file analyzer, a parallel job runner, and a small interpreter for a toy language. Same instructions for every model, one attempt, no fixes afterwards. Then hidden tests the models never saw (162 in total), plus a code review with a fixed checklist: did it follow the instructions, is the code readable, does it break on weird input, are its notes honest. Results (out of 100, the hard task counts 3x) Claude Opus 4.6: 92.7 Qwen3.8-Flash-Next "Coder" (local): 92.7 Qwen 3.8 27B Q6 (local): 87.0 https://preview.redd.it/sz517mg73ath1.png?width=1920&format=png&auto=webp&s=d8651d288fc7e712fa3c37b66cbf3dc0b22368be https://preview.redd.it/kdtp1ed93ath1.png?width=2000&format=png&auto=webp&s=df7d60e9a0845e02cd75029abce498c3fdabb570 What I take from it The hidden tests are almost a tie: Opus 4.6 passed 162 out of 162, the Coder model 161, the 27B 160. The Coder model ended up level with Opus 4.6 overall, and they got there in different ways. On the easy task both local models beat Opus (97 and 92 against 88). On the medium task Opus wins (96 against 94 and 91). On the hard one, the interpreter, Opus and the Coder both scored 92, while the 27B dropped to 81. https://preview.redd.it/zhkga08c3ath1.png?width=2120&format=png&auto=webp&s=c60f0554864c105e3a72acd93e71a5efa354b695 Where Opus 4.6 is still better: its code is cleaner and easier to maintain. Where the local Coder model did better: it crashed less often on strange input. With one run each I would not say "a local model equals Opus 4.6". I'd say: on tasks of this size, I could not tell them apart by the results. For my wallet that's already interesting. Keep in mind Opus 4.6 is not the newest Claude; the current ones scored higher in my full test. The local setup My PC: Intel Core i5-14400, 48 GB of DDR4 RAM, two RTX 5060 Ti with 16 GB each (32 GB of VRAM in total), Windows 11. Qwen 3.8 27B, Unsloth Q6 quant: a steady 50 tokens/s. Qwen3.8-Flash-Next "Coder": between 50 and 90 tokens/s, around 60-65 on average. It's the coding variant from the Strata project, a pruned version that keeps half of the experts in each layer, as an IQ1_M GGUF. Details: https://github.com/Niko1221/Strata/blob/main/docs/MODELS.md#coder Limits, so you can judge the numbers yourself One run per model. Differences of 2 or 3 points mean nothing. I ran the Coder model twice: the first time my PC ran out of RAM while it was working, so I threw that run away and redid it from scratch. The numbers here are from the clean run. The review part was done by an AI (Claude Fable 5.1). All the grading sheets are public, so you can disagree with any call. The Coder model was checked with more weird-input cases than the other two, because I added checks over time. So it was graded a bit harder, not softer. Python only, and the tasks are small. This says nothing about working inside a big real project.   submitted by   /u/Short_Regular_7191 [link]   [comments]