DeepSeek-V4.1-Flash's prompt state is only 0.9 KB/token, so you can split it at layer 20 across CUDA/Metal. All you need is 1/10GbE network. FYI. https://tacos8me.github.io/m5-ultra/split/   submitted by   /u/harrythunder [link]   [comments]