Performance tune for gemma4-26b-a4b flash attention shape. by frobnitzem · Pull Request #28450 · ggml-org/llama.cpp · GitHub

Wait 5 sec.

Gemma 26B A4B speedup   submitted by   /u/jacek2023 [link]   [comments]