−2 logit bias on Bonsai 2 27B: 44/50 → 43/50 on MATH-500, +3% tokens

Wait 5 sec.

A recent post here reported that a −2 logit bias on "wait", "maybe", and "perhaps" made Qwen3.5-4B more accurate and shorter on 50 MATH-500 questions. I tried it on Ternary Bonsai 2 27B (PTQ1_0, 5.53 GiB) on an RTX 5060 Laptop 8 GB under Windows, Prism llama.cpp build adfffbe41. It went the other way. Run Correct Avg tokens tok/s Truncations A baseline 44/50 845.2 29.27 2 B −2 bias 43/50 872.3 29.28 2 Both runs used temp 0, seed 42, a 2048 reasoning budget, a 3072 token cap, and the same 50 questions. Run B biased nine token ids covering the lowercase, leading-space, and capitalized forms of each word. 47 answers matched; one truncated miss became correct, and two correct answers became misses. One deterministic pair at temp 0, so treat it as one data point, not proof either way. Everything is in the repo, including every raw reply: https://github.com/7dollarbooks/bonsai2-logit-bias-test Run by Joseph Murray Adams.   submitted by   /u/7dollarbooks_dev [link]   [comments]