Hi, I have a question about the elephant in the room. Why is everyone using harness-side context management? Why do these tools have to send huge blobs of text to the server every single turn? You don't send your whole WhatsApp history every time you send a new message. The server already knows what conversation you're in. I feel like I'm missing some important point here, because I just don't understand why this is made in the most unoptimized way possible. With the current approach the server/harness has to: - match the full text against some existing KV slot - recover from checkpoints if something doesn't match - sometimes lose an important slot because you decided to ask one question - somehow deal with not knowing when a chat is actually disposed - jump around message reconstruction to avoid KV invalidation And then every harness has to implement its own version of context compaction and KV management. Why not just make the context a server-side object? Something like: "create_slot()" "send_message(slot, message)" "delete_slot(slot)" Maybe also: "compact(slot, ...)" "set_cache_policy(slot, ...)" Then you don't need to keep sending the server a reconstructed copy of something it already has. There are already different compaction approaches, composable context, "endless cache" ideas, etc. You could just do this on the server where the actual KV is. For agents this seems even more obvious. If there is no branching, you don't need checkpoints. Just tell the server it's a linear conversation and don't waste RAM on branches that will never be used. Same thing with SSD/RAM offloading. The server knows which KV belongs to which slot, so you could tell it something is hot, cold, persistent or scratch and let it manage where it lives. And I think there is a bigger conceptual issue here. KV cache isn't really conversation history. Conversation history is an application representation of what happened. KV is the result of actually processing that conversation. Why are we forcing the application to reconstruct the history and then asking the inference server to figure out which part of its own previous computation it can reuse? It seems like the inference server should basically behave more like a database: create state → append data → query state → compact state → delete state instead of: here is the entire conversation again, please figure out which parts you already processed. Maybe there is a very good reason why this can't work. I'm genuinely asking because I feel like I'm missing something fundamental. Why isn't persistent server-side context/slot management the normal API for inference servers?   submitted by   /u/Vasili_Sk [link]   [comments]