Which model, which harness? I have data for you.

Wait 5 sec.

I keep testing harness & model setups on real tasks: basic math, vision, computer use (reading emails, navigating stores) and coding. I created a benchmark and a leaderboard. You can see it here: https://airbench.ai/leaderboard?k=poL It measure capabilities (a % of sucess on the various tasks) and speed. For reference, Claude Code Opus 5.5 have a 100% (14mn26s). It's possible to reach the same score locally with zcode/4xRTX6k/glm-5.3-flash-NVFP4: 100% (31m 13s). Same quality, just a bit slower. If you accept just a litle bit of error you can speed things: * DSHv0.2rc2/RTXPRO6000WS/qwen3.8-flash-next-NVFP4: 98% (23m 30s) * qwen3.8-flash-next-iq3_xxs-strata is the speed pick: 96% (7m 41s) on opencode and 94% in 11m 31s on omp. Yes faster that Claude Code!!!! Other findings: 1) On local hardware, the harness matters as much as the model. The same strata quant on the same 5090 scores anywhere from 22% to 96% depending on the harness. 2) Local can now match proprietary models. two example 3) Best model (single RTX 5090) - swift-1.5-qwen3.8-27b-q6_k is the most robust. It scored 96 / 94 / 92% on pi / omp / opencode and averages 82% across 5 harnesses, the best of any model tested on several. - qwen3.8-flash-next-iq3_xxs-strata is the speed pick: 96% in 7m 41s on opencode and 94% in 11m 31s on omp. - qwen3.8-27b-nvfp4 can reach 96%, but it takes 1h 40m to 1h 50m and depends heavily on the harness (37% to 96%). - Things that hurt: the MTP variants lose ground every time (nvfp4-mtp averages 52% vs 70% without it; swift on pi drops from 96% to 55% with MTP). A 65k context also hurts (45–61%). Gemma-4-26b is fast but tops out at 47%. 4) Best harness To compare fairly, I used the three models that all five harnesses ran on the same 5090 (swift q6_k, flash-next-strata, 27b-nvfp4): opencode: 94% average (92 / 96 / 94) omp: 91% (94 / 94 / 86) pi: 71% (96 / 80 / 37) hermes: 67% (82 / 22 / 96) openclaw: 56% (45 / 53 / 69) Opencode and omp are the only harnesses that stay above 85% whichever model you give them. Pi is very good on some models and unreliable on others. Hermes can score well but is slow: most of its local runs take 1h 20m+ and several hit the 2-hour cap, so its scores are partly answers that arrived too late. The cloud runs show the same pattern. With deepseek-v4.1-flash, omp, pi and opencode all score 98%, while hermes gets 82%. If you have one 5090 today: use opencode or omp with swift-1.5-qwen3.8-27b-q6_k for reliability, or with qwen3.8-flash-next-strata for speed. Ok if you want to read more detailed analys like this one, you can contribute as well! https://airbench.ai/ My website allow everyone to benchmark their setup and contribute to the leaderboard. It's extremly easy to test your local agent: just copy a prompt the website will generate for you. My hope is that we can test much more config on many various hardware. (1) The website requires a login, sorry for that, but it helps keeping false submissions away (2) The website don't ask enough details about the config, so please your the notes field to document your setup in details Let me know what you think.   submitted by   /u/dh7net [link]   [comments]