recently saw a bunch of open decision models pop out of nowhere in the last two weeks (laya, liquid's d1, cloudflare's clef-flash, interfaze's lev), so I wanted to see how far apart they actually are on the same GPU(yes, model size is a huge factor, but still isn't the only factor). all four had the same task of reading nine wikipedia articles about centipedes (9,534 words) word by word and flag every word that names a centipede. one /v1/systemone call per word, the next word goes out the second the answer comes back request for every word: {"state": "Word: \"Scolopendra\".", "questions": {"centipede": {"type": "noul", "instructions": "Does this word name a kind of centipede?"}}} model weights engine per word (p50) words in 32s accuracy centipede names caught wrong picks Laya Laya-BF16.gguf llama.cpp b11495 3.9 ms 7,980 97.4% 70% 98 d1 3B d1-3B-AD-Q4_K_M.gguf llama.cpp b11495 6.0 ms 5,306 96.5% 51% 51 Clef-Flash 9B Clef-Flash-Q8_0.gguf llama.cpp b11495 24.4 ms 1,292 97.2% 36% 2 Lev 4B interfaze-ai/lev, bf16 lev serve (PyTorch) 51.0 ms 626 98.9% 83% 4 laya and d1 gap the other models in speed, though not so much on accuracy (yes, it does say 95%, but even saying "no" counts as a correct answer, so that's where the high acc comes from). what everyone might care about more is how well each one did their respective task and lev catches the most while being 13x slower than laya, partly because it runs in its own pytorch server instead of llama.cpp (it measured 68 ms on a different 4090, so it's CPU-sensitive too). but in the end Laya is the fastest model overall, and considering how easily it can be fine-tuned for any use case I'd say that be my go to pick setup: GPU: rented RTX 4090 (driver 580.119.02, 32 vCPU) engine: llama.cpp b11495 (commit 37ac63456, CUDA 12.8 release build), -ngl 99, everything else default Laya, Clef-Flash: the ggml-org GGUFs d1: our own AD-Q4_K_M quant (atomic.chat), runs natively on /v1/systemone since the lfm2-d1 support landed in #30110 Lev: interfaze's LoRA on Qwen3.5-4B in its own lev serve, default settings (--compile never finished warming up) latency: end to end from a Python client on the same box over localhost   submitted by   /u/Fun-Meaning-6474 [link]   [comments]