github-actions[bot] Bot released b10181 of ggml-org/llama.cpp · July 29, 2026 08:09 ggml-org / /ggml-org/llama.cpp b10181 ggml-cuda : disable MMQ on devices with less than 48 KiB shared memory (#26141)ggml_cuda_should_use_mmq() selects MMQ purely from the quantizationtype. The current MMQ configurations are designe… Read more