Finally, the NPU is being useful in my pi coding agent. Halogen shipped the endpoints, I wired them in expecting a gimmick, and kept four tools. tldr: Qwen3.8 Flash-Next, a 125B MoE, on a 70W tablet. Same bug fix: 13.6 min with NPU search vs 18.7 without. Receipts in the repo. The payoff: it replaces about 95% of my cloud calls. The hardest few percent still goes to the top models, GLM 5.3 or Opus. I still can't believe it. Opus 4.8-class intelligence on my tablet, unlimited tokens. Flash-Next decodes at 64 tok/s and prefills ~1,500 tok/s. First token lands in ~0.03s, about 43x faster than a cloud call measured side by side, and still 7x with a second agent hammering the server. Rate the taste, not the throughput: pasted a real timeshift error from my system log to seven runs. All seven said healthy, nothing to fix. The difference is what they proved. In pi: Flash-Next, 2m55s: proved it with a journalctl trace to a racing notify-send, a pacman.log check, and the upstream PR found. glm-5.3-flashx, 2m41s: the most precise answer, spotting that the snapshot mount got unmounted under the script's last line. No PR. GLM 5.3 on max, 8m30s: the deepest answer of all, source-level forensics down to the function names and the one-second race window. No PR. glm-5.3-flash, 9m09s: proved it with a live reproduction of the status file. No PR. In opencode: flash in 1m8s with the right verdict and the wrong mechanism, flashx in 1m30s correct and corroborated, and the full 753B GLM 5.3 in 6m15s correct with the PR missed. Same pi harness, same task, 125B at medium effort against 320B and 753B tiers at max. First to the full answer: 2m55s. When I had GLM 5.3 flashx rate both results, it picked qwen too. All seven answers side by side: local vs cloud model comparison. Search. The agent stops guessing paths and finds the right file first try. ~0.1s per lookup, beat ripgrep 15/20 vs 9/20 on realistic queries. Dup scan. Catches copied and renamed files git never shows you. Found 45 pairs across 4 repos in 8.4s, one renamed file at exactly 1.000 cosine. Decisions. Yes/no branching stops eating full turns of the big model. A 0.8b handles it in 120ms, 78% accurate. Screening. Prompt injection gets flagged before the agent acts on it. 0.7s a message, zero false alarms, fails open. 42% recall, so a smoke detector, not a safe. A working day claws back about half an hour over bare pi: faster bug fixes, faster compaction, faster lookups and routing, no oversized tool dumps in context. Against a cloud setup it's more, since every turn there pays the network wait. On bug fix heavy days it grows. Honest part: the GPU still does the thinking. The NPU didn't make it faster, it changed what tokens got spent on. ~7% iGPU cost only when they overlap. Compaction: my 194k session, sidecar summary in ~50s vs 166 on the main model. 97% cache hit. The official halogen launch is a 24-flag docker command. Mine is one command, uninstall undoes it. Fully local: 262k context, code never leaves the box. Anyone else using the NPU for something real? I found nothing. repo | halogen 0.17.1 | benchmarks   submitted by   /u/stereohype [link]   [comments]