On OpenCode and Pi, I was fed up with the slow prefill and the SSD kicking in despite having plenty of RAM. With a few targeted tweaks, I can now achieve speeds that I can't even get on Linux. I'm running Windows 10 here. What impresses me most is that performance doesn't degrade as the context increases. I'll see if it's possible to adapt this engine to other LLMs.   submitted by   /u/LegacyRemaster [link]   [comments]