Based on my experiments and feedback from claude, it seems extremely unlikely and even if I get it working, not practical because of minimum context. But just asking here as a last effort. I'm able to run nvfp4 brilliantly and high context with images (or full 256k without images) with the command below and this model. ```bash This works great with NVFP4 from quasar-qat PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \ MAX_JOBS=4 \ vllm serve ~/myp/models/quasar-qat \ --max-model-len 202144 \ --max-num-seqs=4 \ --kv-cache-memory=8589934592 \ --max-num-batched-tokens 4096 \ --kv-cache-dtype fp8 \ --served-model-name Qwen3.8-27B-NVFP4 \ --host=0.0.0.0 --port 8080 \ --enable-prefix-caching \ --enable-auto-tool-choice \ --tool-call-parser qwen3_coder \ --reasoning-parser=qwen3 \ --speculative-config '{"method": "mtp", "num_speculative_tokens": 3}' ``` But I am curious about the performance difference between FP8 and NVFP4. Granted I'll lose some context, but it's always good to know what I'm losing in accuracy. But the repo size is 30.9GB so I'm wondering if I'm out of luck? https://huggingface.co/Qwen/Qwen3.8-27B-FP8/tree/main   submitted by   /u/BitGreen1270 [link]   [comments]