I asked 3 Qwen 3.8 models to solve recent AtCoder problems at 4 thinking efforts

Wait 5 sec.

I was looking into finalizing my daily driver and it was not clear which one to chose. I have a 5060ti 16GB so I picked the following contenders, which seem to be working at a reasonable speed on my hardware: Swift 1.5 IQ2_XS GSQ-RCO under Strata Flash Next IQ3_S GSQ-RCO under Strata 27B IQS_3 GSQ-RCO under llama.cpp I also wanted to understand how thinking effort impacts model performance, so I also included all 4 available thinking efforts (none, low, medium, xhigh) as part of the sweep. I ended up building my own mini-benchmark using AtCoder's problems as test data and LiveCodeBench as a basis for a test harness. I picked only the problems from the past 30 days for test set to make sure that the models did not see them as part of their training. I ended up with 14 problems of varying difficulty as part of the test set. All tests were ran with thinking budget capped to 32k tokens. Hardware: AMD 7945HX CPU, 64GB of DDR5 5200 RAM and 5060ti 16GB. Software: OS - Arch Linux. llama.cpp: build 11149, 140k context, Strata - 0.1.38 with 128k context. It took almost 11 hours to run all tests. CAP+DNF burn is the time spent on problems which were not solved. https://preview.redd.it/rf85frufznuh1.png?width=893&format=png&auto=webp&s=7c03164f7b605e9973a6aaae7f22229dbdf15764 Here is the summary table: https://preview.redd.it/wojizpqeunuh1.png?width=845&format=png&auto=webp&s=e732e20d69d2abd6cfe47159b4cff3ef92a3f082 Legend: problem - problem id at AtCoder pts - difficulty points, given to each problem by atcoder low, med, xhi, non - efforts CAP - model ran into 32K thinking cap DNF - no code was produced ERR - code was produced but did not return expected result Rest of the cells show model runtime in seconds in case if expected result was returned Result interpretation: Thinking does not buy much on this benchmark. All three models peak within ±2 problems of their low score, and no model improves monotonically from low → med → xhi. Flash peaks at med (11), Swift at xhi (10) but with 3 ERRs at med, 27b is flat at 10 everywhere. No thinking effort is the only setting that reliably degrades — and it degrades in a specific way: DNFs appear only at none (4 of them, all flash/swift), plus ERRs jump (27b 0→2, flash 0→2, swift 0→4). Without a thinking channel the models either skip the code fence or write wrong code. Raising effort mostly buys CAP-outs, not solutions. Flash goes 4→3→5 CAPs, Swift 3→3→4, 27b 3→4→4, while time rises 33–57% from low to xhi. The extra thinking tokens go to problems the model was going to fail anyway. Final verdict: Swift is fast but sloppy, I may use it if speed is more important than accuracy. 27B is reliable but very slow. Flash was able to solve more problems and was faster, so I do not see reasons for using 27B. Flash IQ3_S seems to be be a great compromise in terms of speed and accuracy. Furthermore, it solved one more problem at medium effort than 27B. I also noticed during my daily use that at Q3 quant Flash seems to be smarter than 27B and sometimes reports insights which 27B missed. In terms of thinking effort, I will continue using medium and switch to xhigh if medium is stuck. I may also try turning off thinking for simple tasks. I do understand the limitations of this test - there were not too many data points and model behavior has some randomness to it so the numbers should be taken with a grain of salt. I may continue with this testing and narrow it down to more complex problems with much higher thinking budget and see what happens. Please let me know what you think and share your own observations.   submitted by   /u/ColorsOfCosmos [link]   [comments]