I like tinkering with stuff and I think that now I understand at least on surface-level how a dense model like 27B works and what needs to be finetuned for a specific configuration. Then, like everyone else, I discovered Strata and running a large moe locally became possible. The problem is that I cannot get anywhere close to Strata's speeds with vanilla llama server. I have no idea what most of the moe-specific flags actually do to my machine and it all feels like a huge black box without feedback. It bothers me to not understand what's happening inside and what I need to improve or upgrade to make it better. I'm trying to run an IQ3_XXS on vanilla llama.cpp for which Strata gives about 35 tps on thinking and 50 tps code. My best run so far on llama is 15 tps. Hardware: 2 x RTX3060 12GB + RAM 64GB. I cannot find the specific 800MB mtp q2_0 drafter used by Strata on HF. The only one I found is the mtp-Qwen3.8-Flash-Next-Q4_K_M.gguf, which is 3.5x the size. It makes decode considerably slower. For now I'm testing without mtp at all. Tensor split works with dense models, doesn't work with Flash Next (allocating 23087.12 MiB on device 0: cudaMalloc failed). It feels like llama treats it like a 24GB vram card, not 2x12. Anything I should adjust here? Is anyone running on a similar setup? If you could share your working llama recipe for it, I would appreciate it. Thank you   submitted by   /u/Dreeew84 [link]   [comments]