WTF is a small overview check-in remedy (attempt) for lying agents. And more

Wait 5 sec.

I've been driving myself slightly crazy with agentic work. I mean, it's awesome and all, but this one thing: no matter how many test harnesses I put in the development loop, it still breaks shit. While happily claiming everything is green. Made me think: I need an overview I can control, a sanity check for both the agents and me. When I looped in agents to do this, it felt like it did a better job. I'm actively working on getting that as good as it can get, but the original use case is this: A "this is what really happened" reality check at the end of the message from agents. It's local, fast, free, and absolutely under active development (even though I've done 3 safety & security iterations). Test it only if you're comfortable with early dev tools. FAQ (only one so far) Why not just Git bro? Git + filesystem + Compiler + tests = software-reality evidence 👉 WTF Git is a huge part of what makes up software-reality, but it's not all, and agents get fucking confused when they need to find shit out; it takes lots of burned tokens, and it's just not optimal. CURRENT development I'm trying to fix the deterministic reality so agents don't need to waste smarts on that; we're not quite there yet but experiments are kind of promising. But among many other things I need to figure out why 14b is failing: https://preview.redd.it/q499wyg67uqh1.png?width=1418&format=png&auto=webp&s=8417063cb3e9c001142557227ba67567450b0ab0 So the main hypothesis is this: WTF isn’t supposed to make the model clever. It’s supposed to stop wasting cleverness on deterministic reality. If you're having dejavu reading this, it's because I posted about it yesterday, but it was not 100% LLM-free text so I had to go back and make a new post. Hope this stays up -   submitted by   /u/Disrupt-Linus [link]   [comments]