I was running Qwen3.8-Flash-Next (base and later Swift-1.5 NVFP4) on a paid pod seat for about a month, and it started answering with short strings that echoed its own earlier reasoning. I assumed the context was getting too big for it. But it wasn't. Across almost 13,000 turns the echo appeared 266 times, and the more copies of it already in context, the likelier the next turn produced it. The culprit was `preserve_thinking: true` feeding prior reasoning back into every prompt. Removing the echo removed the problem. Once I trusted Flash-Next again, I decided to check out the latest buzz of Flash-Next on Strata but had to see if it would work on AMD. 176B total parameters, 6B active per token, 66.4 GB of weights on a 23.8 GiB resident in VRAM, 31.6 GB of expert weights and a 28.8 GB PLE table in system RAM. Results: Decode is 105–160 tok/s. Context is 500,000, and I measured it actually holding 506,849 tokens. Quality held between 81% and 84% with thinking off and 88–94% with it on across 4k, 64k and 256k, and nothing truncated. Decode fell 18% over that range — slower at depth, not worse. An incredible improvement over my llama.cpp and Swift 1.5 27b dense. Set to 500k context in pi and now my daily driver. This has been another huge leap forward for me locally. Full writeup https://t0mj.github.io/holler/posts/flashnext-q2-strata-amd-daily-driver/ I plan to do more writeups about my experiences with flash-next and strata on AMD for local development on a gaming rig. Disclaimer: My writing is a mix of hand written with ai guidance, pushback, probing, and evidence gathering. (Yes, I wrote this statement.)   submitted by   /u/human_in_the_looop [link]   [comments]