For quantized models, the quantization error propagates across the context size. I feel like 128k is the sweet spot preserving enough information for the current task in hand, and avoiding quality degradation due to quant error accumulation. I argue a 128k with compaction beats raw 256k/ 1M context size. That is valid for quantized models (around Q4), providing a good compaction (e.g. PI agent compaction). The compaction provides the additional benefit of intelligent noise filtering (getting rid of: file dumps, terminal logs, outdated information, ..). What is your experience?   submitted by   /u/Informal-Trouble2183 [link]   [comments]