I was going to use the following setup but that gets me just a measly 4864 tokens of context (I forgot to consider llama.cpp's compute buffers). llama serve -hf bartowski/Ornith-1.5-9B-GGUF:IQ4_XS --fit on --cache-type-k q8_0 --cache-type-v q4_1 --temp 0.7 Are there any models that use much less KV cache? I think the Qwens won't work because of the SSM-based attention having a comparatively large state but IDK.   submitted by   /u/Aggravating-Push-207 [link]   [comments]