Running inference across three machines - an AMD box (Ryzen 9 9950X + RX 7900 XTX), a smaller NVIDIA box (i5-10400 + RTX 3050), and a MacBook Pro M3 - all running LM Studio/Ollama. The serving/monitoring side is where I lose the most time: no single place to see what model/version is loaded where, token throughput, VRAM vs unified-memory pressure, etc., without checking each machine by hand. What are you all actually using for: Multi-node / multi-GPU serving + routing? Observability that is not "wire up Prometheus on every box"? Keeping track of model versions across nodes? Happy to share my current setup if it is useful.   submitted by   /u/ziyaulhuk12 [link]   [comments]