A few days ago I released Jeff-Qwen3.5-0.8B, a small "System 1" model that picks between options you define and returns a calibrated probability for each, in one forward pass. Speed was great on my M4 Max and RTX PRO 6000, but as a general zero-shot classifier it trailed the big models. Then it occurred to me that most decisions an agent makes in front of a local model aren't open-ended. They fall into a handful of recurring kinds: is this a prompt injection, which tool to call, how urgent is this ticket, is this answer grounded in the sources. So I trained 9 LoRA adapters, one per job, and you pick the ones you need. The server loads the base once plus whichever adapters you choose (about 40 MB each), and every request either names an adapter or goes to plain Jeff. That means you keep both: the base model stays untouched, so you still get Jeff's general zero-shot ability for anything new, and the adapters give you near-perfect accuracy in the domains you care about. Each adapter was also trained with 10% of the base model's own training data mixed in, to help it keep its general skills. Everything is on jeffhub.ai: the adapters, the results, the docs. Code on GitHub, models on Hugging Face, and you can try all nine adapters in your browser. The headline: I let Jeff + adapters answer first and pass only the queries it's unsure about to Qwen3.8-27B. Same test rows both ways, on an M4 Max: Measure Qwen3.8-27B alone Jeff + adapters, 27B only when unsure Accuracy (mean of 8 adapters*) 86.6% 95.3% Time per decision (mean) 8.1 s 0.25 s (38× faster) Wrong answers 13.4% 4.7% Memory 28.6 GB under 2 GB for Jeff, even with all 9 adapters loaded (+6.9%) On the five decisions an inbox agent makes for every message (guard, triage, support intent, tool choice, grounding) alone: 87.7% → 95.7%, 39× faster. Jeff wins outright on 8 of the nine adapters and ties on grounding (96.3% vs 96.7%, at 20× the speed). On their full held-out test sets, six of the nine adapters score 97–98%. On a GPU, a decision takes about 30 ms, whether you load one adapter or all nine. *Emotion is left out of the averages: picking the single strongest of 27 emotions (or neutral) in short Reddit comments is hard even for people, and the human labels often disagree. Jeff + adapter scores 60.6% there against the 27B's 35.6%, at 42× the speed. Including it, the average across all nine adapters is 91.4% for Jeff + adapters against 80.9% for the 27B, so leaving it out makes the gain shown above smaller, not larger. Caveats, up front: the 27B ran in 8-bit with step-by-step reasoning off (with reasoning on, the speedup would be even more dramatic); each task used a fixed sample of 300 held-out rows (500 for emotion and legal-clauses); each adapter's "pass it on" threshold was chosen on separate calibration rows, before the test rows were scored. Data: 4 adapters are trained on public data sets. 5 are mostly synthetic. Every generated row records which model wrote it, and the cards give the counts. Every data set went through a shortcut check and an independent review before training, and a lot of first drafts failed: things like the answer being given away by length. What's open: weights (Apache 2.0), code (MIT), and each adapter's test and calibration sets, so you can check every number. The training data isn't published. https://jeffhub.ai: every adapter with its results, where it goes wrong, its data and its QA report, plus Python and TypeScript examples Code: https://github.com/firelex/jeff Models and adapters: https://huggingface.co/mstrasser Try it in the browser: https://huggingface.co/spaces/mstrasser/Jeff-adapters-demo Reproduce the 27B comparison: https://github.com/firelex/jeff-reference-app This is a community preview: I'd love feedback. Next: over the next ~36 hours I'll train v1.3, a long-term-support base. The fixed parts of a prompt come first, so servers can prepare them once and reuse them, which means faster decisions. I'll then retrain all nine adapters on it and keep the request format stable, so others can build and submit their own adapters. The adapter kit, with the data checks I used, is in the repo. I've got access to more hardware now, so if there's a decision you'd like an adapter for, tell me and I'll train it. The goal: when the next generation of local models lands (like everyone, I'm watching for Qwen 4), anyone running one locally should also have a tiny, fast, well-calibrated decision layer in front of it.   submitted by   /u/Usual_Maximum7673 [link]   [comments]