I'm using Qwen3.8-Flash-Next on Strata with pi as the harness. and it was able to have a 100% score on the airbench test suite in 11mn. Claude Code, for the same task get the same score, but in 14mn. (Airbench test various task that are useful for everyday use. It's not a benchmark to discriminate frontier models) Model: https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF Strata: https://github.com/Niko1221/Strata pi: https://github.com/earendil-works/pi And to validate the setup airbench: https://airbench.ai/ MORE DETAILS: Hardware needed to reproduce: One NVIDIA GPU with 24 GB or more, about 64 GB of system RAM or more, and 90 GB of free disk. RAM is the real limit: the IQ3_S quant keeps about 50 GB of experts in RAM. On my 64 GB machine it works but is tight: about 5 GB stayed free during the runs, with a little swap use. Strata's own estimate for IQ3_S at a 256k context is about 78 GB. Model server setup (Strata) Model: ISTA-DASLab GSQ-RCO quantization of Qwen3.8-Flash-Next, IQ3_S (3.5-bit). Strata downloads it (about 84 GB) on first start. Build the image once (I used commit 99f3dbd; CUDA_ARCHITECTURES is 120 for a 5090, 86 for a 3090): git clone https://github.com/Niko1221/Strata cd Strata git checkout 99f3dbd docker build -t strata:99f3dbd --build-arg CUDA_ARCHITECTURES=120 . Run it with the IQ3_S quant and a 256k context: docker run -d --gpus all --net=host --shm-size=8gb --ulimit memlock=-1 -v /path/to/strata-data:/data -e HOST=127.0.0.1 -e PORT=9000 -e FAMILY=qwen -e MODEL=IQ3_S -e CONTEXT=262144 -e VISION=yes --name strata strata:99f3dbd It serves an OpenAI-compatible API at http://127.0.0.1:9000/v1, model id qwen3.8-flash-next-iq3_s. Everything else is Strata's default. The engine flags its setup picked: --expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5 --mtp /data/mtp/rt --max-context 262144 --kv int8 --vision --vram-reserve-mib 700 Harness setup (pi) pi version 0.73.1 (npm package u/mariozechner/pi-coding-agent), run in a node:22-bookworm-slim container, non-interactive: pi -p --mode json --model qwen3.8-flash-next-iq3_s "" pi's config directory (PI_CODING_AGENT_DIR) has two files. models.json: {"providers": {"local": {"baseUrl": "http://127.0.0.1:9000/v1", "api": "openai-completions", "apiKey": "sk-local", "compat": {"supportsDeveloperRole": false, "supportsReasoningEffort": false}, "models": [{"id": "qwen3.8-flash-next-iq3_s", "reasoning": true, "input": ["text", "image"], "contextWindow": 262144, "maxTokens": 32768}]}}} settings.json: {"defaultProvider": "local", "defaultModel": "qwen3.8-flash-next-iq3_s", "quietStartup": true, "compaction": {"enabled": true, "reserveTokens": 65536, "keepRecentTokens": 20000}} Two settings differ from pi's defaults: a 32k output cap, and a 64k compaction reserve (with pi's default 16k, one image result could push the prompt past the window). Running the benchmark Start an airbench checkup from airbench.ai, you just have to copy a prompt that contains the chalenges and wait for the results. (I used my own orchestrator script, which creates the checkup, writes the pi config, runs pi, uploads the session log and fetches the grades. It also puts a small proxy between pi and Strata: it passes requests through unchanged, except that it shrinks an output request that would overflow the context (when at least 8k tokens of room remain) and rewords overflow errors so pi recognises them and compacts. Other harnesses on the same setup Same Strata IQ3_S 256k server on the RTX 5090, one run each: omp 18.4.2: 48/49 (1 failed) in 16 min. https://airbench.ai/checkup/2d3a4f0c-0b3c-4eec-ac79-c177815c5720 opencode 1.18.29: 46/49 (3 failed) in 11 min. https://airbench.ai/checkup/a3121ebe-3221-40f4-a565-32a44e329d20 dsh (DeepSeek Harness) 0.2.0-rc.2: 47/49 (2 failed) in 11 min. https://airbench.ai/checkup/9e464769-409c-440d-a179-1587940928da   submitted by   /u/dh7net [link]   [comments]