Cloud-AI Cold-Turkey: Real Dev Work with Local AI (Ornith 1.5 35b-a3b and Qwen 3.8-27b; 8GB VRAM vs 32GB VRAM)

Wait 5 sec.

Spent several weeks on an 'as-much-Local-AI-as-possible' regime and have been - mostly - impressed. Yes, Qwen 3.8 27b is the (rightful) star of the show (obligatory one-shot-mario-build-reddit-comment here!). But don't underestimate what models with more modest hardware requirements can do. With Qwen forsaking us 30b-a3b-enjoyers in their latest releases, I figured I'd share my experiences with Ornith 1.5 35b-a3b - which offers that size and plays well agentically. I picked it over the similarly sized "Qwen 3.6 35b-a3b" MoE model because that Qwen MoE has no thinking levels (beyond 'On' or 'Off') and hence tends to underthink and thus undercook its answers. I had some luck queuing a 'double-check your results' follow-ups with it; but more thinking off-the-bat would be first prize. Ornith (and a few other similar models) addresses this through additional training that leaves if feeling like the equivalent of a High reasoning mode for the Qwen MoE (I don't know how much additional knowledge it's acquired - but the fact that it seems to work harder legitimately improves the results in my experience). Harness Started with Pi; found that whilst Qwen 3.8 27b was fine within it, Ornith battled with file-writes/edits frequently failing/retrying. Swapped over to OpenCode (full OpenCode-via-llama model config is below) which resolved this at the expense of a higher base context (about 11k tokens in my own setup - which includes a few optional plugins) I already use OpenCode across the board for all my cloud AI uses (mainly OpenAI/GLM) so I've been happy with this consolidation move overall; swapping to a cloud model mid-chat when things get complicated works well for me; this is where OpenCode excels. Hardware and Performance 30tps [100k-140k-context] up to 40tps [   submitted by