Follow-up to my post from a few weeks ago (14 architectures, full PEFT). This update is about a portability bug that was hiding behind its own correctness, because it's the most interesting thing I've fixed since. The bug: the fix for one machine broke six fixtures on another Back when I tuned the backend for Intel Gen9, I baked those kernel shapes into the portable path. That was wrong, but not for the reason you'd guess. Two of the reductions in the saved-module path aren't really compared against "PyTorch in general" — they're compared against the PyTorch CPU library on the machine running the oracle. And ATen dispatches its vectorized CPU kernels by instruction set at run time. An AVX2 host gets 8-wide kernels; an AVX-512 host gets 16-wide ones, and the reduction shape changes with that dispatch. So my "portable" AVX2-shaped kernels were exactly right on my AVX2-only laptop and one ulp off on my AMD ROG Ally (Ryzen Z1 Extreme, which is an AVX-512 part). One ulp doesn't sound like much until it gets amplified through every lower norm on the gradient path: the Gemma 4 saved-stage model.embed_tokens adjoint went from 7.45e-9 to 3.22e-6, and six previously green PEFT saved-module fixtures (gemma3, gemma4, minimax_m2, minimax_m3, smollm3, qwen2_5_sliding_tied) crossed the 2e-7 gate. Neither shape is wrong — only one matches a given machine, and baking in either one breaks the other. The fix: probe the host, not the vendor Kernel variants are still selected by GPU vendor. Those two reductions are now selected by host CPU capability instead: capability is probed once per process and cached, then the matching module pair is dispatched (linear_forward_lane2 / linear_forward_lane4, and the 8-lane / 16-lane transformer_cross_entropy builds). HIERARCHOS_ATEN_VECTOR_WIDTH=8|16 pins the shape for qualification when a reference wheel's kernels disagree with the CPU's own capability. Host GPU CPU dispatch Status Intel i5-6200U / HD Graphics 520 (2016 Skylake-U) Intel Gen9 AVX2 only, no avx512f 32/32 LoRA, 32/32 switching, 32/32 saved AMD Ryzen Z1 Extreme RDNA 3 AVX-512 32/32 LoRA, 32/32 switching, 32/32 saved Same 2e-7 gate, unchanged. No tolerance was loosened to get there. What I verified on each side On the Intel machine, the post-change matrix is bit-identical, field for field, to its pre-change report across all 32 families — peft, gradient, two-step AdamW, frozen base, resume, lifecycle — which is how I know the AMD fix didn't quietly cost the Gen9 path anything. Also 693 passed / 0 failed / 9 ignored on the Rust lib suite and a clean strict headline forward run. On the AMD side, the fix was re-qualified end to end: 32/32 on all three stages, provenance clean. The harness fingerprints the pinned Transformers source alongside the shaders and binaries, and on the Intel side I re-derived the whole fingerprint from the pushed tree myself: 3951 inputs, zero changed, zero missing. So "green" refers to one frozen set of reference math, not whatever happened to be on disk. Same caveats as always This is deterministic FP32 tiny-model correctness against a reference implementation, not a claim about arbitrary checkpoint sizes, dtypes, or hyperparameters. "Supported text graph" ≠ "the whole multimodal package works natively." The AVX-512 dispatch is only qualified on the AMD machine, since it's the only host I have that can execute it natively. The 16-lane module also doesn't rebuild byte-identically with the glslang version on my Intel box (one extra type/id, one difference in +inf materialization), so I've left it as the committed AMD-built module and documented that rather than swapping it without re-qualifying both hosts. I'd rather report that than pretend it's clean. NVIDIA and other GPUs are genuinely unqualified — the path is raw Vulkan, so they're untested rather than excluded. What I'd love from you Last time several people asked about hardware other than mine, so that's the ask again: if you build it on an AVX-512 laptop, an AVX2-only machine, or an NVIDIA/Intel GPU, I want to know what you get. The two reductions above are the ones most likely to behave differently on your CPU, and knowing your host's vector width is now part of the answer. The new cross-platform section in the README documents the whole thing, including which host classes are measured and which aren't. Repo: https://github.com/necat101/Hierarchos-Native Compatibility/parity record: https://github.com/necat101/Hierarchos-Native/blob/main/hierarchos-vulkan/COMPATIBILITY.md Regression audit: https://github.com/necat101/Hierarchos-Native/blob/main/AMD_REGRESSION_AUDIT.md Per-host tuning and measurements: https://github.com/necat101/Hierarchos-Native/blob/main/hierarchos-vulkan/VENDOR_TUNING.md   submitted by   /u/PhysicsDisastrous462 [link]   [comments]