In this paper by Meta and University of Washington researchers, they threw out compaction. Instead, they allowed the model to edit its own context window on every "turn," using its shell skills to avoid re-generating the whole thing, and deciding on its own what to keep. The model was also given the option to "just append the output" in a given turn. Benchmark results went up. And they used smaller models in their research, like Qwen 3.5 9b and Qwen 3.6 27b for instance. So these techniques should prove effective on models we're running at home, and ought to be easy to implement in harnesses like Pi or OpenCode. It should work on any model large enough to be competent in bash. I look forward to seeing the community's results with this!   submitted by   /u/boutell [link]   [comments]