Gufo performance .... 70tps Qwen 3.8 27b but you need to read the fine print.

Wait 5 sec.

So everyone has been yelling about how I should be using Gufo instead of halogen because it's open source and it's "just as good or better". Checking in on their GitHub (GitHub.com/gufo-org/gufo) got me immediately .. "Qwen 27B Q4: 70.56 tok/s single user, 123 tok/s with 8 users" on a strix halo device? Yes please ... So I broke down and tried it today ... Setup: gufo 0.4.0 from their podman image, Qwen3.8 27B UD-Q4_K_XL from Unsloth plus the DFlash2 Q4_K_M draft model, using their own benchmark script and their own settings (greedy, thinking off, 128 output tokens, prompt cache off). If you want the short version ... yeah I got 70.22 tok/s. So the number is real. But the prompt that produces it is "Write the word red exactly 1000 times". But it's also not real. In that figure all the speed comes from the speculative decoding. The draft model guesses like 7 tokens ahead, the 27b checks them in one pass and keeps what it agrees with. When the output is the same word over and over the draft is right every time. On a real prompt it's right maybe half the time. Their benchmark has a second set of nine ordinary prompts (some C++, a word problem, a summary, Italian, Chinese, JSON, a bit of fiction, a debugging checklist). On those: | | repeat-a-word prompt | normal prompts | |---|---|---| | 1 user | 70.2 tok/s | 39.4 tok/s median, anywhere from 22 to 52 depending on the prompt | | 8 users, their "aggregated" number | 122.6 | 82.4 | | 8 users, tokens actually delivered per second | 82 | 52 | About that last row. The "123 tok/s aggregated" figure is each request's decode speed added together, with prompt processing and queue time left out. If you just count tokens coming out of the box per second of wall clock it's 82, or 52 on normal. prompts. To be fair to the gufo people, none of this is hidden. Their benchmark docs have separate "mixed" and "repetitive" columns and the mixed numbers they publish match what I got. It's only the repo description and the top of the README that lead with the best case. And 39 tok/s from a 27B at Q4 on an APU is still really good. Without the draft model their docs put it around 12. The other thing I wanted to know was how it compares to halogen (peonist-ai/halogen-flash-server), which is what I normally run. Both can serve Qwen3.8 Flash-Next, so I put that on both and sent the same prompts to each. Two boxes, same hardware, same OS image. Greedy, thinking off, 256 tokens. | | halogen 0.13.8 | gufo 0.4.0 | |---|---|---| | nine normal prompts, average decode | 43.9 tok/s | 38.2 tok/s | | 4 users at once, end to end | 76.6 tok/s | 63.1 tok/s | | cold prompt processing, ~9.7k tokens | 1288 tok/s | 1495 tok/s | | the "red" prompt | 56.9 tok/s | 87.4 tok/s | So for everyday generation halogen was about 13% faster for one user and about 18% faster with four. gufo was 16% faster at chewing through a long prompt and a lot faster on the repetitive one. I'll add this just in case, because someone will ask or at least try to poke about it in the comments - I know the weights aren't the same. halogen uses its own 4-bit format, gufo uses the Unsloth GGUF. I only measured speed. I did not compare output quality at all. - They were two different machines but identical hardware and software, and my boxes have agreed within 1% on other benchmarks, but it's still two machines. - One run each was done for the head to head. The reproduction of their numbers was 3 reps. - gufo has shipped four releases over the last four day, so this could all be stale by next week. There was quite a lot of stuff that I liked about gufo that isn't performance related. It takes plain GGUFs, it's MIT, the 27B loads in about 3 seconds (Flash-Next in 13), the per-request log line tells you draft acceptance and cache hits, and it does 8 batched sessions. It also has ASR, TTS and image models that I haven't touched. Their benchmark hashes the output with and without the draft model and it was identical every time, so the speculative path isn't changing what the model says. One thing to keep in mind if you try it out is that it reserves memory per session up front. Flash-Next with 4 sessions at 64k context took 94GB. So it isn't smoke and mirrors exactly. Everything I checked reproduced. Just know that the 70 is a ceiling you'll only hit if your workload is incredibly predictable text, and plan around the 30s for the 27B on normal stuff. I kept this all setup to tinker with on actual output quality over the next few days, I'm happy to run other prompts or try it with different settings if anyone wants to see something specific.   submitted by   /u/brainchillzZ [link]   [comments]