Is there a better option than llama.cpp for 4GB VRAM for Higher tokens/sec?

Wait 5 sec.

I love the idea of running local models on consumer hardware. I currently use `llama.cpp`, but I’ve been looking for a "better" alternative for a while now. I want something that utilizes my system resources more efficiently—for instance, by managing my RTX with 4GB VRAM more intelligently—and delivers higher tokens-per-second, all without the burden of heavy dependencies like PyTorch. I’ve done my research, and honestly, I haven't found a solution yet. I came across plenty of options, but unfortunately, many were built on Python and heavy libraries like PyTorch, which would essentially exhaust my limited VRAM before the model even loaded. Others were designed for running unquantized models, were outdated (based on models from two years ago), targeted a completely different class of devices (like MNN), or were merely "cool-looking" papers with no actual implementation.   submitted by   /u/your_real_Fathe_ [link]   [comments]