She runs Debian 13 with vLLM Docker containers, serving Qwen/Qwen3.8-27B on a single GPU with a 512K context window for long-running agent workloads, alongside a second Qwen/Qwen3.8-27B deployment using its native 262K context window for agent delegation (--max-num-seqs 8). Full build here: https://pcpartpicker.com/b/YRH2FT. And here are the vLLM Docker container files, if anybody is interested in those: GPU 1: gist.github.com/e8dba21e4bd520fa31996c5b2e75e1a0 GPU 2: gist.github.com/6b55123eddc44aa495e1671c8c5ae231   submitted by   /u/Smart-Skin-2346 [link]   [comments]