Follow-up to my post from about two weeks ago. A lot has changed since then, so I wanted to post an update on where the backend is now. Where the green architectures stand When I posted last time, 7 architectures had verified full training support. Everything is now held to the same strict harness: a pinned local Hugging Face Transformers source tree as the oracle, forward logits + gradients + two full AdamW steps compared, and every named parameter checked again after export. The hard ceiling is 2e-7 absolute error. No loosening tolerances and no rounding numbers afterward to make the README look better. 14 architectures pass that gate today, led by the one I'm probably proudest of: Architecture Scope Falcon H1 / H1R parallel GQA/RoPE attention + Mamba2 in every layer; full training, full fine-tuning, LoRA, saved modules DeepSeek V4 causal LM Phi-4 Multimodal text backbone Phi-3 causal LM Kimi K2.5 text backbone Kimi K3 / KimiLinear hybrid KDA + MLA GPT-OSS causal LM incl. router bias SmolLM3 mixed RoPE/NoPE + YaRN Qwen2.5 / Qwen3.5 / Qwen4-Exp dense, DeltaNet, QSA, PLE, MoE Mistral 4, MiniMax M3, Gemma 3/4, MiniMax M2 causal LM Worst observed two-step AdamW parameter error across all of them: 1.19e-7. Best: 2.6e-8. For hardware context, all of the local Vulkan validation I've been reporting was run on my ASUS ROG Ally Z1 Extreme, using its AMD RDNA 3 integrated GPU. So the RDNA 3 results here are from that specific machine rather than testing across several different AMD systems. The bigger news: PEFT actually works now In the last post, "LoRA/PEFT-style fine-tuning" was basically one line in a feature list. It's a real workflow now, and I've verified the full lifecycle: LoRA fine-tuning with HF-compatible adapter export (adapter_config.json / adapter_model.safetensors), so adapters can round-trip with the PEFT ecosystem modules_to_save — full trainable replacements for Linears, RMSNorm/LayerNorm, lm_head, and input embeddings, including named-adapter switching and bank isolation. Adapter A leaking into adapter B is explicitly tested for. Exact resume — adapter weights + AdamW moments + step + dropout RNG state restore bit-identically against an uninterrupted run Merge/unmerge, disable-adapter base restoration, and multi-adapter loading A parameter-budget flag that automatically chooses the largest LoRA rank that fits within a requested percentage of the base model The CLI fails closed if you try to use saved modules on an architecture that hasn't passed its corresponding gate 32 architecture surfaces across 20 families pass all three PEFT stages — LoRA, saved modules, and adapter switching — under the same 2e-7 gate, with frozen-base drift exactly 0.0. The validation harness also fingerprints the pinned Transformers source alongside my shaders and binaries now, so a qualification run can't silently end up testing against different reference math. A small note on the last couple weeks I didn't get quite as many working days out of the last two weeks as I normally would have. Partway through this I got covid, then when that started to go away, it became a secondary nasty ear infection that ended up perforating my eardrum. I was running fevers around 104°F at one point and eventually went to the hospital, so I lost a few days to that and I'm on antibiotics now. I'm doing better, though, and still managed to get most of what I wanted finished. There are still things I want to clean up and expand, but I figured this was a good point to get the current work in front of people rather than holding the update back. Same caveats as before This is deterministic FP32 tiny-model correctness against a reference implementation, which should in theory ensure total mathematical parity for training and finetuning larger models with this backend, however, things are currently bound to FP32 training runs still, I eventually plan to work on MXFP4 weight tying to reduce memory footprints while fine-tuning (plastic parameters will still be trained in FP32 with PEFT in this config) "supported text graph" ≠ "the entire multimodal package works natively." Unsupported functionality is supposed to fail closed rather than silently falling back to an approximation. Repo https://github.com/necat101/Hierarchos-Native Architecture inventory: hierarchos-vulkan/README_ARCHITECTURES.md Compatibility/parity record: hierarchos-vulkan/COMPATIBILITY.md PEFT qualification evidence: PROGRESS_PEFT_AUDIT.md CLI PEFT guide: hierarchos-native-cli/README.md The hardware I've personally validated this on is an ASUS ROG Ally Z1 Extreme with its AMD RDNA 3 GPU. I'm very interested in criticism, compatibility reports, and especially results from people trying it on other hardware — NVIDIA, Intel, or other AMD GPUs. I'd also love people to stress-test the PEFT resume/merge paths specifically. That's some of the newest code in the project, so it's probably the most useful area to try to break right now.   submitted by   /u/PhysicsDisastrous462 [link]   [comments]