Strata vs Infernix on the same model, RTX 5090 — A/B at 262k and 512k context Hi folks, another inference engine post for your feed. I was blown away by Strata, but saw someone else post about Infernix with some wild claims, and it doesn't seem so well known, so I've spent a day downloading models and A/B testing each engine. Completely stunned that either of these pieces of software run on Windows, and the setup was relatively painless too. Same-day interleaved benches, same model (Qwen3.8-Flash-Next Uncensored, orcarouter's Apache-2.0 abliterated checkpoint), same inference contract (int8 KV, MTP spec-4 + lm-head-draft, YaRN 2 for 512k). 3 decode runs per cell, medians. Hardware: RTX 5090 32 GB (Gen5 x16), 189.6 GB RAM, 9950x, Windows 11, models on NVMe, llama-swap fronting both. Engines: Strata 0.1.41, orcarouter Q4_K_S GGUF, calibrated flags (pcie-frac 0.55, pool-workers 4, adapt-every 1/160/0.97, spec-min-p 0.7) Infernix 2026.10.09.2, published NVFP4 recipe-C artifact (76.3 GB) + shared 52 GB n-gram volume on NVMe (mmap'd, never loaded to RAM — the 63.3 GiB of experts is what's pinned). Vision works (--vision, tower offloaded to pinned RAM, borrows VRAM only while encoding) Results (decode tok/s, median of 3 / cold prefill tok/s): 262k 512k (YaRN 2) Strata Q4_K_S 120.6 / 1,550 126.3 / 3,045 Infernix NVFP4 149.0 / 3,548 137.0 / 3,497 Takeaways: Infernix +23% decode at 262k, +9% at 512k — same model, same spec contract. 512k is nearly free (≤8% on either engine). On a 32 GB 5090 experts are the bottleneck, not attention KV. Prefix cache makes repeat long prompts free (~0.09 s warm re-prefill of a 39k prompt). Quality: statistically tied — teacher-forced PPL 4.772 (Infernix) vs 4.844 (Strata's quant), ΔNLL −0.015 ± 0.010 nats; the artifact runs the abliterated weights bit-exact. Vision confirmed working on both engines — Infernix needs --vision (off by default) and hard-validates the model string (--model-id to match your proxy's entry name — got me with a 404). Cold start: Infernix ~28 s; Strata ~1 min. Caveats: decode measured at ~39k effective context (deep-filled 500k is a different KV-vs-cache question); ±10-15 tok/s run noise, so treat sub-10% deltas as noise; YaRN past 262k is experimental (speed benched, not 512k quality). Verdict: Infernix NVFP4 is the daily driver — ~150 tok/s, 512k, vision, all on one 5090. Strata stays as the second engine. Using it for coding and other tasks, it's fantastic at creative writing / role play in Silly Tavern, and I've been using it instead of smaller creative finetunes. For agentic work and tasks, it has been more than capable of handling everything I've asked of it, and has corrected mistakes from Deepseek 4.1 and GLM 5.3 Flash that I had to pay them to write over api.   submitted by   /u/zipzak [link]   [comments]