what are experimental approaches to run 30b dense models on like 8 gb vram

Wait 5 sec.

i cant rlly find any papers, but there have to be some ways to get it running without being lobotomised or offloaded, right?   submitted by   /u/Aggravating-Push-207 [link]   [comments]