The gap is smaller than they told you: local 27B nearly matches frontier on real code tests

Wait 5 sec.

I did not want to guess, so I pulled the benchmark’s own published results for the same task. The numbers that matter: On this exact task, the frontier cloud models get an average of 96.6% of the hidden tests right. This local model got 98.0%. The frontier models pass this task outright, with zero missed tests, only about two times out of three. Missing it is normal, not a scandal. Across the full 113-task benchmark, the best cloud models score about 70 to 74%. Read those together. On this one task, a four-bit 27B running alone on a single consumer graphics card was right about as often as the frontier average, and the frontier missed the same task just as easily as it did. ------------------------------------------------------------------------------------------------------------------------------------ The model that produced the DeepSWE result is the IQ4_XS at 196K (the entry is named Q4-Quality but the per-boot settings differ; both are below). Model: Qwen3.8-27B (unsloth dynamic quant) Repo: https://huggingface.co/unsloth/Qwen3.8-27B-GGUF File: Qwen3.8-27B-UD-IQ4_XS.gguf (14.25 GB)   submitted by