M5 Ultra 80Core GLM-5.3-Flash on DwarfStar Speeds

Wait 5 sec.

I've been playing around with various models on the M5 Ultra 256GB 80-core Mac Studio. These are the results over many rounds of agentic inferencing. I'm happy with the performance. Glad to have the large amount of RAM. But it does feel like the GPU is underpowered for this amount of RAM. I'm wondering if a 512GB unit for AI inference makes sense at all - because the GPU will be the clear bottleneck.   submitted by   /u/dreamingwell [link]   [comments]