Closed-form trajectory weight surgery: transferring 4B capabilities into 0.8B without billions of tokens (why middle layers break, and how 4 anchor blocks fixed it)

Wait 5 sec.

Standard distillation usually means burning weeks of compute and billions of tokens hoping the student model eventually mimics the teacher. We wanted to see what happens if you skip backprop entirely and treat transfer as a closed-form trajectory matching problem between layers. The idea is straightforward: feed a small batch of calibration prompts through both models, capture layer-to-layer hidden state trajectories, and solve for weight updates directly in the student's MLP blocks using regularized least squares and spectral projection. We tested this across two architectures: Qwen 3.5 (transferring from 4B down to 0.8B) and old GPT-2 small just to see if it would instantly disintegrate into gibberish like it usually does when you touch its weights. Both stayed coherent, but the initial Qwen test hit a wall: Editing all 24 layers of Qwen 0.8B completely melted the model (+64.78% NLL loss explosion). When we checked singular value entropy across the network, layers 1 to 22 turned out to be a chaotic polysemantic soup with entropy over 0.90. If you try to force raw trajectories through those middle layers, you basically scramble the model's internal memory knots. The fix was restricting the surgery to 4 anchor points (layers 0, 7, 15, and 23) where representations actually maintain clean linear structure. Once we did that: Held-out NLL dropped by 10.8% across 30 diverse benchmarks (-23.8% in biomedicine, -14.6% in math and logic). Zero-shot 400-task HellaSwag went from 54.75% to 55.25% (+0.50%), verified locally in Vulkan llama.cpp. Base 0.8B originally failed binary tree inversion by spitting out dead commented pseudo-code. The edited checkpoint wrote clean recursive Python on the first try. On Russian logic paradoxes, it even started firing reasoning tags spontaneously, which was wild to see on a raw base model with zero chat template applied. Best of all: we don't have an H100 cluster or even a 4090. All trajectory extraction and weight solving was done locally on a crusty 8GB RX 580 using layer-by-layer VRAM streaming and a DirectML patch to stop Qwen's Gated DeltaNet attention from throwing driver errors. Everything is open source if you want to inspect or replicate: Code and math writeup: https://github.com/dsadawq3/DynamicTune Transferred base model weights: https://huggingface.co/F-Labs/Qwen3.5-0.8B-DynamicTune-Base If anyone here has a 24GB-32GB card (4090, 5090, or server silicon) and wants to push this further, here is what would be interesting to test: Transplanting reasoning trajectories from 27B models down to 9B, 4B, or 2B. Squeezing larger models (like Gemma) into mobile sizes without weeks of retraining. Transplanting refusal-ablation vectors directly from uncensored models without fine-tuning. Using pre-trained Sparse Autoencoders (SAEs) to unknot layers 1-22 so we don't have to skip them. Happy to answer questions or dig into the failure modes in the comments.   submitted by   /u/AdventurousTwo6445 [link]   [comments]