I was looking through some Qwen3.8 27b Finetunes and found this thing called VeriLoop E2 from Tsinghua SIGS Robot Lab https://huggingface.co/tsinghua-sigs-robot-lab/VeriLoop-E2 It's a 27B post-trained model based on Qwen3.8-27B and the benchmark numbers are kinda insane , probably benchmaxxed The README claims: SWE-bench Pro: 76.2 Terminal-Bench 2.1: 88.8 DeepSWE v1.1: 64.6 Terminal-Bench 3: 29.7 Terminal-Bench 4: 37.9 AIME 2026: 98.3 GPQA Diamond: 93.9 Apex 2025: 89.6 SWE-Marathon: 45.0 And then I started looking at the comparison table and that's where it got really weird On Terminal-Bench 2.1 it reports 88.8, which is the same as the listed GPT-5.6 Sol score and above Kimi K3 at 88.3 On SWE-bench Pro it's at 76.2, above Qwen3.8-Max at 67.7 and GPT-5.6 Sol at 64.6, although Claude Fable 5.1 is higher at 81.2 still tooo suspicious Terminal-Bench 3 is 29.7, above the listed GLM-5.3 at 28.3 GPQA Diamond is 93.9, basically right next to GPT-5.6 Sol at 94.1 and above Kimi K3 at 93.5 and this is apparently a fine tune of best 27B model Are these numbers actually legit / reproducible or Benchmaxxed? I'm more interested in whether there's something about the evaluation setup / Harness / benchmark protocol that makes these comparisons less straightforward than they look I'd try an quantised version of model and report soon :)   submitted by   /u/zyxciss [link]   [comments]