I've been experimenting with different coding agents lately, and I'm curious how people here are combining local models with hosted ones. For example, I'm thinking about workflows like: Claude for complex architecture or core implementation A local Qwen/Gemma model for tests, smaller fixes, or repetitive tasks Another agent for reviewing or trying an alternative implementation The part I'm still trying to figure out is how to manage the work between them. Do you run them separately in different terminals/worktrees, or are you using some kind of orchestration layer? And when a local model and a stronger hosted model both work on the same task, how do you decide which result to keep? I'm actually working on an open source project called AX Code around this problem. The idea is to provide a runtime where different coding agents can work in isolated environments and have their results tested and compared. But I'm not sure yet how much infrastructure is actually necessary. Git/worktrees already solve a lot, and tools like Claude Code and OpenCode are getting better at running multiple agents. So I'm more interested in how people are doing this today. If you're using local models as part of a real coding workflow, what's working well for you and what's still painful? Thank you!   submitted by   /u/No-Wait-7495 [link]   [comments]