Hey everyone, If you run local models via Ollama in production or personal projects, you've probably run into the hallucination problem: how do you know when a model is hallucinating without burning extra VRAM or waiting 5 seconds for a heavy judge model? The standard academic approach for this is Semantic Entropy (from an Oxford team's Nature paper last year). You sample $K$ responses at temperature 0.7, run them through a secondary NLI cross-encoder like DeBERTa to cluster equivalent meanings, and measure the entropy. High entropy = model is guessing. The problem for local setups? Running 45 pairwise comparisons through a cross-encoder eats GPU memory, adds 100ms+ latency, and completely kills throughput on consumer hardware. We wanted to see: What if we strip out the neural net completely and just use deterministic string normalization + Shannon entropy on CPU? We wrote a zero-dependency Python metric (Spanda / $R_{sc}$) that runs in 1.3 microseconds on pure CPU (zero GPU usage) and benchmarked it across local and frontier model tiers on GSM8K and TriviaQA: What we found: Small models (Qwen 1.5B): AUROC ~0.58 Small models are syntactically too sloppy for string matching. Even when they know the right answer, they format it erratically across runs, breaking exact-match clustering. Mid-sized models (Mistral 7B): AUROC ~0.71 At 7B, the 1.3µs string check matched the performance of a heavy DeBERTa NLI model (0.706 vs 0.705). Internal representations become consistent enough that formatting stabilizes. Large models (Qwen 27B): AUROC ~0.89 At 27B, exact matching was dominant ($p = 1.89 \times 10^{-28}$). When the model knows an answer, it outputs the exact same tokens across independent stochastic paths. When it doesn't, it genuinely branches into diverse incorrect answers. The Frontier Trap (120B): AUROC collapsed to 0.09 Here’s the wild part: on ungrounded factual trivia, the 120B model suffered Confident Mode Collapse. When it hallucinated, it hallucinated the exact same wrong answer across all 5 runs with zero entropy. Bigger models don't just hallucinate—they hallucinate with unanimous false certainty. (And because the strings are identical, even heavy NLI fails here). The practical takeaway for Ollama users: If you are running 7B to 27B models on structured tasks (math, code, JSON extraction, SQL, discrete QA), you do not need heavy neural guardrails. Sampling 5 paths at $T=0.7$ and measuring exact-match entropy in Python gives you ~0.89 AUROC at zero GPU cost. Quick Python snippet if you want to test it on your local Ollama instance: bash pip install spnda ollama pythonimport ollama from spnda import compute_spanda prompt = "What is the capital of Australia?" # Sample 5 paths from your local model responses = [ ollama.generate(model="mistral:7b", prompt=prompt, options={"temperature": 0.7})["response"] for _ in range (5) ] # Run zero-cost entropy check on CPU (takes ~1.5 microseconds) result = compute_spanda(responses) print (f"Risk Score: {result.risk_score:.3f}") # 0 = high confidence, 1 = high uncertainty All the raw multi-path generation logs, evaluation scripts, and the full writeup are open source: Deep dive writeup: https://liquidngas.substack.com/p/i-tried-to-make-semantic-entropy GitHub: https://github.com/nayakbhupen/Spnda   submitted by   /u/More_Slide5739 [link]   [comments]