Running the uncensored Qwen3.8-27B (HauhauCS) on a 4090 at 262K context and ~130 tok/s

Wait 5 sec.

HauhauCS ships their uncensored Qwen3.8-27B as GGUF only. NInfer, a C++/CUDA wanted its own format. Now the same model that ran at 91.6 tok/s / 131K under llama.cpp does: 262K context (the model's full native window) ~130 tok/s decode with MTP3, 70.8% acceptance 3,591 tok/s prefill on a 9K prompt Perplexity within 1.3% of the official artifact, so the conversion is clean Vision and tool calls still work Converter + writeup here: https://github.com/T-Crypt/ninfer-4090/pull/4 Questions welcome. Repo: ninfer-uncensored   submitted by   /u/Distinct-Pie2389 [link]   [comments]