We had a Dell B300 in the lab for a few weeks and used it to create two fine tunes of Qwen Flash Next. Victoria (coding and agents) Qwen3.8-Flash-Next cut down by 44% using a paper / technique called REAP: 512 down to 288 per layer. Retrained at 4-bit (NVFP4) afterwards, so it's trained for the format it ships in rather than just quantized after the fact. Terminal-Bench 2.1: 70.0%, averaged over 3 runs with an 8h per-task timeout. Our previous NVFP4 build scored 62.5%. HumanEval: 159/164. 48.0 GiB of weights, including the draft head. The 95.4 GiB n-gram table is separate and not counted in that number. 280 tok/s single stream on one B300 with the draft head, versus 135 without it. GGUF Q4_K_M is 49.17 GiB. It scored 75.3% on Terminal-Bench (a single run, so treat it as noisy) and 93.2% on HumanEval (averaged over 5 runs). Uses 35% fewer output tokens than our previous build. Maple (Canadian questions) Most models answer questions about taxes, benefits and regulations as if you live in the US. Maple is fine-tuned to default to Canada. On 600 held-out questions, with search: Cites an official Canadian source: 6.0% before fine-tuning, 62.9% after. Fully correct answers: 6.6% before, 21.8% after. "No answer" responses: 47.2% before, 23.7% after. It pushes Canada onto people who said they live somewhere else less often: 2.9% before, 1.0% after. Coding holds up: 157/164 on HumanEval. Grading was done by an AI judge panel; human review hasn't happened yet. Links: https://huggingface.co/rmonsurate/Victoria https://huggingface.co/rmonsurate/Maple Happy to answer questions about running them. Edit: llama.cpp users. The GGUF carries our draft head, and mainline llama.cpp doesn't know about it yet, so it fails with "expected 1256, got 1224". Your download is fine. For now, build from our fork: github.com/rmonsurate/llama.cpp, branch qwen4exp-mtp. Prebuilt binaries are on the way. Thanks to the reader who caught this.   submitted by   /u/rmonsurate [link]   [comments]