CUDA: support state size 96 for ssm_scan op for Nemotron 3 Puzzle by anavp-nvidia · Pull Request #28717 · ggml-org/llama.cpp
Read post on reddit.com
Nice speedup, maybe it's time to try this model on your setup?   submitted by   /u/jacek2023 [link]   [comments]