Your Strix Halo NPU is sitting idle. My coding agent uses it for search, dedup, decisions and screening tldr: Halogen 0.16+ exposes your idle XDNA2 NPU to OpenAI-style endpoints - I wired four of its models into a real coding agent and couldn't find anyone else doing that. I measured everything, then cut what scored weak. The stack survived a 2-hour gauntlet that crossed the compaction trigger, and halogen 0.17.1 dropped NPU latency 30x and decode ~40%. NPU tool what it does measured codebase_search semantic file search (embed + rerank) 15/20 vs 9/20 for grep, 70-130ms on 0.17.1 dedup_scan find near-duplicate files across repos 45 pairs in 8.4s, catches renames decide constrained binary decision (0.8b) 78% on realistic prompts, ~120ms auto-guard injection screening, fail-open tripwire 42% recall, 0% false positives (30-prompt canary) /rag-index build a search index for a repo ~5,700 tok/s The retrieval is real but only for questions without keywords. "Where is the timeout applied when sending a new request instance?" - semantic search finds it; "fix the failing test_query_encoding" - grep is faster, because the test name tells you what to search for. I proved the effect was real with a placebo control: shuffled-vector index vs real index, both arms solved in 1-2 min. The NPU earns its keep on the questions grep can't name. The NPU only accelerates the "where to look" step. The other 95% of task time is reading code and writing the fix - all GPU-bound. Measured with a keyword-poor task: 13.6 min mean with NPU vs 18.7 mean without (the NPU arm's slowest run barely used the tool) - the GPU does the thinking, the NPU does the pointing. Stack comparison (this box, same model): stack / backend model prefill (t/s) decode (t/s) takeaway halogen 0.17.1 + NPU tools flash-next v2.hgn 1,311 64 MTP / 38 serial fastest, most instrumented halogen 0.16.2 (before) same ~1,300 42-46 MTP NPU calls 30x slower llama.cpp rocm + qwen4exp MTP UD-IQ4_XS GGUF n/a 47 first working MTP on ROCm (community) Moderation scored 67% in my first eval and I cut it. Then I re-tested it as a passive tripwire on a 30-prompt canary set: 42% recall, 0% false positives in 18 benign prompts including pentest-flavored ones. It ships now as a fail-open warning label, not a gate - "low-noise screening, not a security boundary." benchmark numbers, methodology & what I removed after testing Rig: ASUS ROG Flow Z13 (Ryzen AI Max+ 395, Radeon 8060S, 128 GB unified, Linux), benches at 70 W sustained. NPU driver needs IOMMU on. Halogen 0.17.1, ling3.0-tiny sidecar on llama.cpp 0.7.6.1-era build. NPU microbench (0.17.1, four models resident): Single-query embed: 70ms; rerank: 130ms; guard: 70-130ms (on 0.16.2 every NPU call carried ~4s of fixed overhead - 0.16.3's batching fix removed it. If you're on an older release, your NPU numbers are 30x worse than they should be.) Decider (binary): 118ms, 78% accuracy (18 realistic prompts) Decider (3-option triage): 925ms, 40% accuracy (removed - too biased toward "keep") Moderation: 42% recall / 0% fp on 30-prompt canary (ships as tripwire) Text gen (qwen3.5-2b): 16.6 tok/s decode (removed - slower than the sidecar) LLM decode while NPU works: unchanged (separate silicon) Capability eval (20 intent queries, 2 repos, 3.8M tokens indexed): NPU semantic (src-only index): 15/20 top-3 correct file Ripgrep keyword baseline: 9/20 top-3 Grep latency: ~2-33s per query Near-duplicate scan: 38 files across 4 repo copies, embedded in 8.4s; found 45 pairs with cosine > 0.90; caught every known duplicate plus a renamed file at 1.000 cosine. Keyword-poor task (implement --fail-fast in a smoke script): NPU arm: 6.2 / 11.9 / 9.7 / 26.6 min (mean 13.6, all 4 completed) Baseline arm: 24.8 / 16.6 / 14.8 min + 1 failed (25% failure rate) Saving: 5.1 min per task, with the caveat that the slowest NPU run barely used the tool The invalid experiment I almost published: my first "controlled experiment" showed a 2x speedup - but the NPU arm never had a working index (path bug), and the task was a public httpx fix the model had memorized (92s, zero search calls). Per-call telemetry caught it, and every tool call now logs its index and query. Compaction offload, measured: the sidecar summarizes a 113k-token session in ~50-60s vs 166s for main-model compaction (whose request restructures the conversation, so the prompt cache misses). Rule-retention canary: 10/10 with thinking off. Net: ~80-95s saved per compaction event. Two-hour live gauntlet: one continuous session, 119 model turns, 7 phases of real repo work (docs audit, a 4-provider dry-run harness, a mail inspector, an adversarial input canary, a subagent code review, vision verification of a rendered dashboard). 430k fresh tokens, 14.1M cached (97% hit), five commits, all verified after the fact: the dry-run passes 4/4 providers, the inspector's tests pass, and the vision phase found three real bugs in a feature the docs called done - including a chart that silently stretched 35 seconds of samples across a "24 hours" axis. The session also crossed the compaction trigger at turn 100, which is what surfaced act 2 above. The harness measures its own engine: a /speedtest template runs a 9-probe battery (MTP vs serial decode, joint speculation, NPU latency, cache-cold prefill, sidecar summarize) and grades itself against expected bands. First run: 8/9 in-band; the miss was a real 1.3s tool overhead, measured and documented - not the NPU's fault. What's in the harness now (10 extensions): - Ling-tiny suite (4): compaction, commit messages, repo maps (freshness-cached), branch summaries - all on the sidecar, thinking off - npu-retrieval: codebase_search + dedup_scan + decide (3 tools that earned their place; auto-index refuses oversized roots) - progress-tracker: crash recovery - kill the agent mid-task, the next session picks up the exact task - auto-guard: the fail-open tripwire above - harness-tune: /tune for live config knobs - turn-timer + the subagent fleet Removed after measurement: moderate-as-gate (67%), npu_write (16.6 tok/s), 3-option triage (40%), tool-result triage and two more extensions that logged zero fires in real sessions, plan-mode (3 uses in 569 sessions). Shipping tools that measured weak would be worse than not shipping them. 0.17.1 bonus: joint speculative decoding pushes two concurrent streams to 72.7 tok/s aggregate, serial decode hit 59.6 tok/s (n-gram read-ahead), prefill unchanged at ~1,300 tok/s, and the NPU batching fix took search latency from ~4s to 0.1s. I also pinned the container image version - floating :latest means your receipts measure a moving target. install & ops Install - fully automated, one prompt total. Clone github.com/aic0d3r/qwen38-strix-halo-harness (tagged v1.0.0, MIT, changelog in the repo) and run ./setup.sh --halogen auto. It installs pi via npm if missing, pulls the pinned halogen image, and if you don't have the model files it downloads them for you: ~124 GB resumable (that's the one prompt you'll see - it checks your free space first), then the 35 MB portable llama-server for the sidecar, the tiny gguf sha256-verified against Hugging Face, ten extensions, prompt templates, skills, theme, and the systemd keep-alive unit - and your own model entries and pi settings are never touched. ./setup.sh --doctor verifies every piece with live probes; ./setup.sh --uninstall backs everything up and returns pi to stock in one command. Docker is the one prerequisite we won't auto-install (system package, needs sudo). First boot pins the 47.7 GiB lookup table from disk, so a few minutes of startup is normal. The launcher pre-flights allocator fragmentation before committing to 262k context and fails in seconds with the exact fix instead of wedging for half an hour, and a 15-min systemd compaction timer keeps the allocator healthy between sessions. Container image is version-pinned; scripts/session-audit.sh runs post-session forensics (cache hit, compaction events, sidecar rejections, NPU index anomalies) and checks the serving version against the pin. Anyone else found good uses for the Strix Halo NPU besides what halogen ships - and did you measure recall on your own prompt set before trusting the guard model? Links: halogen 0.17.1 | my harness | the game-ladder benchmark | quants Benchmark, scorer, and toolkit are my own open-source repos. Ran with AI assistance. Every number comes from logged runs on my machine, and the two claims I retracted are still in my post history - that's the point.   submitted by   /u/stereohype [link]   [comments]