How much space and memory would the next generation ngram models take?

Wait 5 sec.

I undertsand that they can offload a significant chunk of their knowledge to SSDs and RAM. This could free up VRAM memory significantly. But, then wouldn't the bottleneck become the RAM and SSD storage? Considering just how much information these models are trained upon, wouldn't even the small 27B models take up hundreds of GB RAM or terabytes of storage? Look at Qwen3.8 Flash Next. Its 2-bit quantisation is still way bigger than Qwen3.8 27B 4-bit quantisation.   submitted by   /u/accelerate_to_asi [link]   [comments]