TL;DR: For my use case, mradermacher/Signal-3.8-27B-Terse-Coder.i1-Q4_K_M proved out. Ymmv based on your domain-specific tests. Time taken to solve problems compared to frontier models is massive; especially if you, like me, run a potato. My current feasible model's KPI over the full eval set is 28x slower. Time gap might be significantly more forgiving for folks with better hardware. About a week or so ago, I shared some prelim tests comparing Qwen3.8-27B fine-tunes. I've since ran a multi-day comprehensive 469 domain-specific eval against some of these fine-tunes. For consistency, I ran them all with llama.cpp and same thinking settings. I also ran the full eval set on Opus 5.5 and Astra as well as partially against Qwen3.8-Flash-Next via OpenRouter. Some providers disclose quantization and others do not. Would be nice if OpenRouter made it mandatory to do so. Others: Astra had the lowest token usage - although it appears the provider masks reasoning. No model got 100% accuracy. Opus 5.5 got close and topped the list at 99.6%. Of my local models, mradermacher/Signal-3.8-27B-Terse-Coder.i1-Q4_K_M had the lower token usage and time to completion of tasks while achieving higher accuracy. A Q4_K_M fine-tune performing better than other L or XL is a nice find. It overthought to cut off only once and passed more tests than others. I read somewhere that the Signal-Terse-Coder is a combo of AgentionAI's Signal-3.8-27B and Shockem's Terse-Coder LoRA. I don't understand the mechanics here but some sort of magic must be going on under the hood. I wouldn't read too much into TTFTs of models on OpenRouter. They cache and I can't do really much about it to influence it. Here's what I'm calling the PTA index. This will differ by model, eval, and hardware on hand. But can be part of a grounding KPI to measure one's progress by. https://preview.redd.it/y27wkaoha5uh1.png?width=2966&format=png&auto=webp&s=0848931754587ba911d5eac3795ab081bce45c02 Median tokens and more interesting info here. The number of questions each model overthought on is represented in "Cut off" column (my setting at max 16,384 tokens per question): 1 violation by Signal-Terse-Coder, 4 by Dirk, and 4 by Unsloth. 19 minutes vs 9 hours is wild; counterpoint: as wild as 0 privacy for frontier is when compared with ~100 for local. Better hardware is the normalizer. https://preview.redd.it/iucw41rzc5uh1.png?width=2984&format=png&auto=webp&s=742905e5344c85f78abd14c2028f729e86c42995 I suspect Astra is appearing to get fewest tokens for 456 questions out of 469 by hiding its reasoning. The same question answered by Opus 5.5 and Astra shows reasoning of Astra is perhaps masked by the provider. Nevertheless, the Astra response is succinct with no code commentary and a valid pass to what was asked. Opus 5.5 https://preview.redd.it/87cicfq1f5uh1.png?width=2846&format=png&auto=webp&s=1f3687c1ed9349a74c873e51b62c46699917783d Astra https://preview.redd.it/aztvm4nkf5uh1.png?width=2858&format=png&auto=webp&s=ec8612c8cded25fa96990eb832159ad0acd192fb It's not in the list but gpt-oss-120b *shat the bed big time* on my pandas/numpy tasks. Only domain specific evals can uncover cases like that and warn you which models/fine-tunes to steer clear of for particular tasks. I also tried AgentionAI/Qwen3.8-27B-AP-Q4_K_XL and UkisAI/Swift-1.5-Qwen3.8-27B-Q4_K_L but I cut the run around 115 questions for both. A clear pattern had developed by then and didn't see a need to let the run go on to completion. https://preview.redd.it/zhiqa0bqt5uh1.png?width=2990&format=png&auto=webp&s=062fb0fb37cedbf58c1a966b8086aa6360f04850 Initally, I made the tool for myself. But have since decoupled the engine from tests/data to make it extensible to other users' choice. It's available here https://github.com/ashe-wb/tuieval I've only been testing from my M1 Max 32GB. Can't say what you may run into on your gear. If any issues or requests, open them in github. I'll put my time into it and maintain it.   submitted by   /u/norenEnmotalen [link]   [comments]