Real long-term memory for local LLMs: saved KV state on NVMe, reloaded byte-exact instead of recompute (I built it, free on 1 GPU)

Wait 5 sec.

A large language model can only use the text that fits in its context window, and it recomputes its internal key-value (KV) state for a prompt every time the prompt is sent. We test a memory layer, the public package galahad-kv, that saves the KV state of each block of about 16,000 tokens to encrypted local NVMe disk and loads it back later, byte-exact, without recomputing it. We ran it on 50,000,000 tokens of real public text, served through vLLM on one NVIDIA H100, with Gemma 4 12B and Gemma 4 31B https://huggingface.co/papers/2610.10845 Disclosure: I built this at Corbenic AI. Free download, free on 1 GPU for non-commercial use. Paper: https://huggingface.co/papers/2610.10845 Code: https://github.com/corbenicai/galahad More discussion in r/huggingface: https://www.reddit.com/r/huggingface/comments/1x175wf/   submitted by   /u/MindPsychological140 [link]   [comments]