Apple's new Mac Studio ships today with up to 512GB of unified memory, and Reuters says Apple demoed Kimi K2.6, a 1T parameter open-weight model, across four of them plugged into a single wall outlet. Thunderbolt 5 + RDMA lets the machines run as a cluster, and Apple says four Mac Studios can deliver up to 3x the inference performance of one. Cloud inference obviously still wins on scale, but we're getting weirdly close to "buy the compute once and stop paying per token" making sense for some teams. Local AI used to mostly be about privacy and tinkering. Cost is starting to become a real argument too. I'm just wondering what workload would you actually pull out of the cloud first if the hardware economics keep moving this way?   submitted by   /u/Novel-Lifeguard6491 [link]   [comments]