This is a real question about a hard and annoying limitation, which is why I assume I'm about to get lambasted by 20 ppl telling me to go buy a better GPU or use 27b or something. specs: 5060 Ti 16 GB, 32 GB DDR5, mid-range SSD. Here it is: I love the idea of Strata, but the "Q2 must use mmap and immediately gets hella slow" issue is a super frustrating limitation. I'm really hoping someone has an inference engine or knows the flags or otherwise has ideas for working around this — to get an mmap-like setup working, with experts swapped in, without being garbage slow. Before you bombard me with "dynamically switching experts in and out is slow by nature" — I'm simply wondering: are there any implementations of this that work better than mmap? It seems silly to me that just a few more gigs means total failure to utilize the GPU — are there more intelligent switching methods or something? The Q1 version of Flash-Next is braintarded, which is sad, because it actually fits my RAM... I would love to run it effectively at Q2 or Q3. 27B is an amazing model but I crave more power — I hunger as a young acolyte wizard wishes to attain the power of his betters... Bonus query: can the Ista DasLab Q2 version run on Strata? And is that going to beat a decent quant of 27B at world knowledge? I assume it'll lose on code, but that's not always what life is about. Thanks for engaging/helping.   submitted by   /u/ironicstatistic [link]   [comments]