My Jetson Orin–optimized engine, little-gemma V1.0, substantially outperforms llama.cpp. Even after exhausting every practical GGUF option, however, it still falls short of the ideal performance level for Gemma E4B on the Jetson Orin Nano Super 8GB. I forked exllamav3 and ported key code from little-gemma V1.0. This delivered the prefill improvement I needed. Although the decoding rate took a hit, overall voice-chat latency is lower. EXL3 on the Jetson Orin Nano isn’t for everyone, but if you need the superior prefill rate, this is it. Fork: https://github.com/cortexist/exllamav3 Weights: https://huggingface.co/cortexist/gemma-4-E4B-it-qat-exl3-4.0bpw   submitted by   /u/cortexist [link]   [comments]