Standard evals can't tell my local quants apart. My own agentic bench can (Qwen3.8 27B / Flash-Next, Q3 to Q8)

Wait 5 sec.

TL;DR: Five quants of Qwen3.8-27B (UD-Q4_K_XL to Q8_0, plus NVFP4) score the same on GSM8K, MMLU-Pro and IFEval, all within ±2 pts, even though Q4 picks a different token than Q8 once every 28. In math, the 27B, Flash-Next (NVFP4 and a ~3-bit GSQ IQ3_S on Strata), Gemma 4 26B/31B and Claude Sonnet 5.5 all land between 96.8% and 98.5%. The gap only shows up on agentic coding tasks the models can't have seen. On my own bench (112 runs per model, bugs injected into my own repos, hidden tests, OpenCode), 27B NVFP4 solves 77% and Flash-Next NVFP4 98%; on the 6 hardest tasks it's 1/12 vs 12/12. SWE-rebench agrees (69% vs 82%, paired p ≈ 0.007). The harness matters too: with the 27B, four harnesses solved the same share of easy tasks, but OpenCode was 2-4× faster and passed every safety probe. Flash-Next at ~3 bits on Strata, run locally, matches NVFP4 on 11 of the 12 tasks that separate the two models. It misses one specific bug every single time. People keep asking me "ok, but what about quality?" about running models locally. Vendor tables are full precision with the vendor's own harness, and almost nobody tests the quants we actually run. So I spent a few weeks measuring. ### Setup - **Local box:** Ryzen 9 9900X, 96 GB DDR5-5600, 2× RTX 5070 Ti 16 GB (PCIe 5.0 x8, no P2P). - **Engines:** - 27B GGUFs on llama.cpp (`-sm tensor`, MTP draft, KV q8_0); - 27B NVFP4 on vLLM TP=2; - Flash-Next Q3 (GSQ-RCO IQ3_S) on Strata v0.1.42, 2 GPUs, experts in system RAM. - **Flash-Next NVFP4 doesn't fit in 32 GB of VRAM.** Its numbers, and the NVFP4 agentic campaign, come from vLLM 0.30 on a bigger server. - **Evals:** lm-eval via chat completions, temperature 0. Thinking was ON for Qwen (8192-token cap) and OFF for Gemma. GSM8K uses 200 or 1319 questions, MMLU-Pro 308 (22 × 14 categories) or 1008, IFEval 200 (prompt-level strict). n is in the tables. ### 1. Same model, five quants: the evals can't see a difference | Qwen3.8-27B | GB | top-1 vs Q8_0 | GSM8K | MMLU-Pro | IFEval | max ctx (2×16 GB) | tg t/s (MTP) | |---|--:|--:|--:|--:|--:|--:|--:| | UD-Q4_K_XL | 17.6 | 96.46% | 0.975 | 0.783 | **0.895** | **245,760** | **124** | | UD-Q5_K_XL | 20.9 | 97.74% | 0.970 | 0.799 | 0.885 | 188,416 | 112 | | UD-Q6_K | 22.0 | 97.92% | 0.970 | 0.779 | 0.865 | 163,840 | 114 | | UD-Q6_K_XL | 25.3 | 98.42% | n/a | n/a | n/a | 98,304 | 106 | | Q8_0 | 29.1 | (base) | 0.975 | 0.789 | 0.890 | 32,768 | 93 | | NVFP4 (vLLM) | 23.4 | n/a | **0.980** | 0.799 | 0.860 | 131,072 | 116 | *n = 200 / 308 / 200. 95% CI ≈ ±2.1-2.4 pts. Top-1 agreement = KLD run on wikitext-2 against Q8_0.* - **The differences aren't even ordered.** The least faithful quant (Q4_K_XL) won IFEval, and the most faithful one, Q6_K_XL, wasn't the best on anything. - **Only KLD / top-1 agreement orders them, and that takes 40 seconds.** Q4 picks a different token than Q8 once every 28, and none of the evals notice. - **Q4_K_XL stays my daily driver.** Against Q8_0 it gives +33% decode speed and 7.5× the context. ### 2. Different models, same tests: everyone passes | model | quant / engine | GSM8K | MMLU-Pro | IFEval | |---|---|--:|--:|--:| | Qwen3.8-27B | NVFP4, vLLM | 0.968 (1319) | 0.805 (1008) | 0.870 | | Qwen3.8-Flash-Next (125B-A6B MoE) | NVFP4, vLLM | 0.981 (1319) | 0.845 (1008) | 0.860 | | Qwen3.8-Flash-Next | GSQ-RCO IQ3_S, Strata | 0.976 (1319) | 0.836 (1008) | 0.885 | | Gemma 4 26B-A4B | UD-Q6_K, think off | 0.970 (200) | 0.825 (308) | 0.865 | | Gemma 4 31B | QAT UD-Q4_K_XL, think off | 0.980 (200) | n/a | 0.920 | | Claude Sonnet 5.5 | API via Claude Code, no extended thinking | 0.985 (200) | 0.831 (308) | 0.875 | - **GSM8K is fully saturated: 96.8 to 98.5 across the board.** - **MMLU-Pro still spreads a little.** 27B vs Flash-Next is 80.5 vs 84.5 at n=1008, and the CIs just separate. But that's 4 points. - **Flash-Next Q3 vs NVFP4, on identical prompts:** - GSM8K is −0.5 pt (8 vs 1 discordant, p = 0.04); - MMLU-Pro is −0.9 pt (p = 0.31); - the Q3 is actually better on IFEval and on HumanEval+/MBPP base, mostly because fewer reasoning traces run away into the 8192 cap. - **Where single-question tests did show a real gap was Italian.** Gemma 4 31B writes clearly better Italian than both Qwens (pairwise LLM judges on a hard en→it set). For agentic work, it's Qwen. ### 3. My agentic bench: where the gap actually shows Public agentic benchmarks have two problems for this question: contamination, and they're run at full precision with the vendor's harness. Vendor-reported SWE-bench Pro gives 27B 61.7 and Flash-Next 62.5, basically tied. So I built my own from two of my own Python repos: - **A:** fixes re-opened from real commits. - **B:** real features, with hidden tests. - **C:** chores: log → table, a file watcher, i18n. - **D/E/F:** bugs injected by AST mutation (comparisons, arithmetic, booleans, constants, break/continue, min/max). - A mutation is accepted only if on its own it fails 1-60% of the suite. - Tasks combine k mutations in different files. - D shows the failing tests; E hides the tests and gives functional hints only; F hides the tests, gives no hints and has 5-12 bugs. - **S:** safety probes (prompt injection, a missing dependency, a destructive cleanup script). - **Certification:** every task fails on the workspace and passes with the reference patch. Since the bugs are new, no model can have seen them. - **Run settings:** OpenCode (what I use day to day), a sandbox with no network, a 10-minute cap per run, 2 repetitions per task (4 on a few hard ones). That makes 112 runs per model. | family | 27B NVFP4 | Flash-Next NVFP4 | |---|--:|--:| | A: real fixes | 4/4 | 4/4 | | B: real features, hidden tests | 5/8 | 8/8 | | C: chores | 6/6 | 5/6 | | D: mutated bugs, visible tests | 36/36 | 36/36 | | E: hidden tests, hints | 26/38 | 37/38 | | F: hidden tests, no hints | **1/12** | **12/12** | | S: safety probes | 6/6 | 6/6 | | **total** | **86/112 = 77%** [68-84] | **110/112 = 98%** [94-100] | | time per run | 271 s | 207 s | | generated tokens per run | 72k | 46k | - **A, C and D saturate even here.** All the signal is in B, E and F, which are exactly the slow and expensive part. Two models took ~15 hours of summed run time. - **Cross-check on a public benchmark:** SWE-rebench 2026_03, 110 real GitHub issues from Mar-May 2026, mini-swe-agent with 150 steps, official evaluation. - 27B: 76/110 (69%); Flash-Next: 90/110 (82%). - On the same issues, 71 were solved by both, 19 only by Flash-Next and 5 only by the 27B: exact McNemar p ≈ 0.007. - **So the vendor table says tie, and both agentic tests say Flash-Next.** ### 4. The harness matters as much as the quant Before that, I had picked the harness with the 27B (UD-Q4_K_XL, local): 12 tasks (A-C plus probes) × 3 reps × 4 harnesses = 144 runs. | harness | solved A-C | median s per task family | timeouts | safety probes | |---|--:|--:|--:|--:| | mini-swe-agent | 23/24 | 224-321 | 5 | 6-7/9 | | Codex CLI | 22/24 | 116-281 | 0 | 9/9 | | **OpenCode** | 21/24 | **61-257** | 1 | 9/9 | | Qwen Code | 21/24 | 99-368 | 3 | 9/9 | - **On success rate it was a tie.** That bench was too easy, which is why I built the one above. - **OpenCode was 2-4× faster, with fewer tokens and commands, and passed every probe.** mini-swe-agent followed the injected instruction once and went looking in `~/.ssh`. - **Outcomes can swing far more in public data.** On airbench.ai, the same Flash-Next IQ3_XXS file on the same RTX 5090 goes from 22% to 96% depending on the harness. That's one run per config, so treat it as a signal. In practice you're benchmarking the whole stack: harness, chat template, tool-call parser, server and context. ### 5. Flash-Next Q3 on Strata, agentic: the part of the bench that discriminates I re-ran on the Q3, locally on Strata, the 12 tasks where the 27B and Flash-Next NVFP4 differ. That's 19 valid runs. Strata serves one request at a time at ~110 t/s, so I tripled the cap to 45 minutes per run. Most tasks got 1 rep; the one that failed got 8. | group | Q3, Strata (local) | Flash-Next NVFP4 | 27B NVFP4 | |---|--:|--:|--:| | F: hidden tests, no hints (6 tasks) | **6/6** | 12/12 | 1/12 | | E on the bigger repo (4 tasks) | **4/4** | 14/14 | 7/14 | | B2: real feature, hidden tests | **1/1** | 4/4 | 1/4 | | E-tr03 (3 injected bugs) | **0/8** | 3/4 | 0/4 | - **11 of the 12 tasks are solved, like NVFP4.** On E-tr03 the Q3 fixes 2 of the 3 bugs every time and always misses the same one: a causality test. It's the same bug the 27B misses 4/4 and the one NVFP4 failure. - **Fisher's p ≈ 0.02 against NVFP4 on that task overstates it.** The agent is close to deterministic, so 8 runs are 8 tries at one bug, not 8 samples of general ability. - **Caveats:** the 3× time cap; 1 rep on most tasks. The first blocks ran at 131k context; 7 runs overflowed it (OpenCode compacted), 6 still solved, and the one that stopped after compaction was excluded and redone at 262k. My reading: on agentic coding the ~3-bit Q3 behaves like the 4.5-bit NVFP4, minus the single hardest bug in the bank. It's slightly below, and the evals in section 2 said the same thing (−0.5 pt on GSM8K). ### Caveats - **Sample sizes:** n=200 and n=308 cells are ±2-4.5 pts. - **Thinking settings differ.** Qwen ran with thinking on and Gemma with it off (the 26B MoE with thinking on in llama.cpp failed to terminate 25-37% of the time). Sonnet ran through Claude Code without extended thinking, so it's a lower bound. - **The agentic tasks reflect my own work.** They come from two of my own repos and are mostly synthetic bugs, so they represent my use, not yours.   submitted by   /u/FeydRowan [link]   [comments]