Unsloth, Swift1.5, Peculiar-Ragdoll, ThinkingCap - Qwen3.8-27B

Wait 5 sec.

In a previous post I shared comparison between Swift1.5 and peculiar-ragdoll's checkpoints. Added the original unsloth Q4_K_XL and ThinkingCap Q4_K_M (they don't offer L or XL) to the comparison. Here are the results over a 69 set of eval questions. All tests are now run at same "medium" reasoning effort. https://preview.redd.it/avvtcevqcvsh1.png?width=2956&format=png&auto=webp&s=dbeea6d9156165d64169eebc1b5284cf7c0416c6   submitted by   /u/norenEnmotalen [link]   [comments]