Qwen 3.8 27B on a single 3090: 114 min solo, 43 min as a worker under a GPT 6.1 SOL orchestrator

Wait 5 sec.

Quick one for anyone wondering what a single 24 GB card is good for in an agent setup. I had Qwen 3.8 27B (Q4, llama.cpp) on one RTX 3090 do all the actual coding, and GPT-6.1 Sol in the cloud act as the orchestrator: it breaks the job into pieces, hands them out, and checks what comes back. The job was three small 3D games: pool, bowling, foosball. How it came out: Game 3090 alone 3090 + Sol giving orders Sol alone Pool 43.1 min · $0 (attempt) 18.6 min · $0.05 2.9 min · $0.39 Bowling 34.7 min · $0 13.7 min · $0.06 1.7 min · $0.14 Foosball 36.4 min · $0 11.1 min · $0.06 2.0 min · $0.22 Total 114.2 min · $0 43.4 min · $0.17 6.6 min · $0.75 So the card on its own is free but slow and gets lost on the harder one. With something smarter doing the planning it finishes more and finishes faster, and the cloud bill stays small because the cloud model barely writes any code. What I'd tell someone before trying it: you still wait a lot longer than with a cloud model alone, one run per game is not a benchmark, and the $0.17 doesn't count your power bill. I ran it in Atomic Agent, the mode is called Fusion (disclaimer: I work on it). You can do the same split in any tool that lets you pick a separate model for planning and for coding. What card are you on? Would you trade the wait for a smaller bill, or is speed the whole point for you?   submitted by   /u/GapNew4766 [link]   [comments]