I went way deeper down the rabbit hole than I originally expected with local AI 😂 4-node Aether build, custom memory/reasoning stack + local VideoForge — looking for people solving the same bottlenecks

Wait 5 sec.

I figured I’d do a deeper followup because the more I read this sub the more I realize there are people here solving pieces of the exact problems I’ve been beating my head against for the last couple years 😂 What started as “I want to run an LLM locally” turned into a whole system I call Aether. At this point I’ve pretty much stopped thinking of the LLM as the assistant. The models are workers. The actual “system” is all the stuff around them: routing, memory, context, truth/state, tools, computer use, validation, recovery, media generation, etc. Current hardware: Primary: - Ryzen 9 5900XT, 16c/32t - RTX 3090 24GB - 64GB DDR4 - 2TB SSD - 2TB HDD - MSI B550 Tomahawk MAX WiFi Beta: - Ryzen 5 5600XT - 32GB RAM - integrated graphics - 1.5TB SSD Those two are connected as part of the same Aether environment rather than just being two unrelated PCs. Beta is useful for the stuff that doesn’t deserve 3090 time — CPU/RAM work, telemetry, testing, lighter workers, storage/background jobs etc — and I’m moving more toward routing workloads based on what hardware they actually need. I also have M3 and M4 Apple nodes in the setup. M3 is already SSH/script controllable from the tower and I’ve used it for remote macOS/Safari/Apple-side work. I’m gradually turning the Apple machines into useful satellite nodes instead of just having Macs sitting around doing nothing. So the direction is basically different machines + different models + different jobs instead of trying to make one machine/model do literally everything. The part I’m most interested in comparing with people here is the system above the models. I have a reasoning/orchestration layer I call PRS — Primary Reasoning Stack. Under/around that is MOC — Mini Operating Core, which deals with stuff like runtime state, context assembly, memory, truth/source tracking, permissions, lane routing, health/recovery, etc. I’ve built a pretty massive memory/cognition system at this point 😂 The idea isn’t just vector RAG. There are structured memory blocks, working/recent context, long-term memory, source-truth packs, bounded context selection, truth arbitration, grounded-response gates, cognition systems, writeback, confidence/retrieval logic, etc. The problem I’m still solving is getting the whole damn thing to behave as one reliable live reasoning path every turn. Basically: user request → PRS/MOC → relevant memory/source truth → cognition/reasoning → correct model/tool → validate what happened → recover if needed → write useful new state back Parts of that work. Parts are proof-backed individually. But I’m not going to pretend the whole end-to-end loop is perfect because it isn’t. And honestly this is probably my biggest question for the serious agent-system people here: how are you deciding what context/memory actually reaches the reasoning model every turn without either starving it or burying it under garbage? Especially once you have a LOT of memory. I’m increasingly convinced “more context” is not the same thing as “better context.” I also have a computer-use layer I call Hands & Eyes. Vision/OCR, UI state, targeting, action planning, verification, retries etc. One of the biggest lessons from that was “command returned success” does NOT mean the action actually happened. That ended up changing the whole architecture. Now I care way more about observed post-state than whether some agent says “done.” That eventually turned into a failure/governance system where meaningful failures get preserved and turned into guards/regressions instead of just patched once and forgotten. Basically: failure → evidence → cause → guard → regression That has saved me from repeating a ridiculous number of stupid mistakes. The other giant rabbit hole is VideoForge. I wanted one pipeline that can eventually go from script → shot planning → image/video generation → character/environment consistency → motion → dialogue → lip sync → SFX/music → edit/composite → QC → package/deliver. Locally I’ve been working with ComfyUI + WAN/VACE and a bunch of custom control/source-conditioning workflows. I’ve got actual local render pipelines running, recipe tracking, source-conditioned experiments, QC gates, retry logic, etc. One experiment I’m particularly interested in is what I call teacher-to-local parity. I have paid accounts with Runway, ElevenLabs and PixVerse. Instead of treating those as “the production system,” my goal is to use them as temporary teachers/benchmarks. Generate something good externally → preserve prompt/settings/reference/result → reproduce it locally → compare → learn the local recipe → eventually stop paying for that class of generation. If local can get close enough, external providers get pushed back to specialist/edge-case roles instead of being the default. I’ve already got an ElevenLabs learning/parity setup built around that idea. One of the local VideoForge tests technically generated correctly but the character walked the wrong direction 😂 Another spawned a modern car in a mythic warrior scene. Those became negative examples/guards instead of getting thrown away. That’s basically the whole philosophy of the project at this point. Where I really want advice: Memory / cognition routing How do you decide which long-term memories, manifests, source docs, past failures etc actually get injected for a task? Do you use vectors, graph retrieval, classifiers, hierarchical summaries, learned routing, a planner, some hybrid? And what did you find broke once the memory corpus got huge? Multi-model / multi-node routing Has anyone here actually built and lived with a router that considers task type, model quality, VRAM/RAM requirement, context, current node load, latency, power/cost and historical task success instead of just hardcoding “coding goes to X, vision goes to Y”? That’s where I want PRS to end up. Local video parity Anyone here training/fine-tuning local video/image systems from your own benchmark/reference set? Not trying to clone somebody’s proprietary model — I mean using outputs you’re allowed to use as comparative targets and improving the local pipeline until it can handle the same class of work. I’m especially interested in LoRA strategy, dataset size/quality, motion consistency, character identity, camera motion, prompt/conditioning transfer and automatic quality scoring. Paying per generation gets expensive REAL fast when you’re iterating constantly. Long-running agent reliability What has actually worked for you when a task is 10, 20, 50+ steps? My experience so far is that reasoning quality matters, but state awareness + recovery matters just as much. One malformed JSON response, stale file, wrong browser target or failed tool call can derail the whole thing if the system doesn’t know how to recover. And yes… I have a hardware problem 😂 Short term, when Primary eventually gets replaced, the 3090 moves into Beta so it becomes a real GPU worker instead of throwing perfectly good compute away. I’m also seriously looking at GB10 / DGX Spark-style large-memory nodes for the giant-model tier after seeing what people here are doing with distributed big-model inference. The ridiculous end-state I’m working toward for the main Aether core is: Threadripper PRO 9995WX — 96 cores 3× RTX PRO 6000 Blackwell — 96GB each 768GB RAM 4×4TB NVMe Then keep the 3090 node(s), Apple nodes and potentially Spark nodes as separate compute classes around it. So eventually something like: big fast core + 3090 workers + large-memory model nodes + Apple/macOS nodes + external providers only when they actually earn their cost Then PRS decides where work goes. That’s the dream anyway 😂 I know there are probably people here who have already hit some of the walls I’m heading toward, so I’m genuinely looking for criticism. I’m less interested in “what should work” than in what actually survived contact with production. If you saw this architecture and had to tell me “you’re overengineering THIS part” or “you’re going to regret THIS decision later” what would it be? Those answers are probably more useful to me than another benchmark chart.   submitted by   /u/Budget_One_8784 [link]   [comments]