Hi everyone :) The amazing Swift finetunes of Qwen3.8 27B generate much fewer tokens at mostly similar benchmark performance to the original model, while repos like HyperQwen (formerly syv-ai/qwen38-27b-rtx3090) deliver insane TPS on an RTX 3090. To get the best of both worlds, I adapted Swift 1 and Swift 1.5 for HyperQwen and benchmarked them against HyperQwen’s specialized W4A16 AutoRound fast quant for speed and quality. Performance Model Average time/task ↓ Average output tokens/task ↓ Decode tok/s ↑ Qwen - HyperQwen fast quant 108.1 s 8,985 112.1 Swift 1.0 + HyperQwen 66.2 s 5,245 105.9 Swift 1.5 + HyperQwen, INT8 heads 72.2 s 5,751 104.0 Swift 1.5 + HyperQwen INT4 heads 68.2 s 5,669 107.2 All models ran on one RTX 3090 24GB with FP8 KV cache and 150k configured context. Average task time covers about 630 tasks from different benchmarks listed below. Swift-1 seems to use the fewest tokens and has fastest task completion time. Swift-1.5-INT4 achieves roughly 37% lower average time per task compared to the HyperQwen Qwen fast model. Despite slightly lower TPS (due to lower draft acceptance) it finishes sooner because it generates fewer tokens. Quality There are some minor quality and performance tradeoffs between the models: Test Qwen HyperQwen fast Swift 1.0 Swift 1.5 INT8 heads Swift 1.5 INT4 heads GSM8K, 200 questions 97.5% 98.0% 98.0% 97.5% IFBench, 300 prompts, strict 74.0% 73.3% 73.7% 72.3% LiveCodeBench, (100-problem subset) 90% 89% 89% 91% Custom tool-call/JSON eval 29/30 28/30 30/30 30/30 English/Python perplexity ↓ 6.551 6.605 6.643 6.679 Applied changes to Swift models to adapt for HyperQwen: Changes to Swift models: Kept the upstream AWQ INT4 model weights and converted embeddings to INT8. Swift 1.0 and Swift 1.5 INT8-head variants: quantized the output head and MTP (multi-token prediction) linear layers to INT8 and added HyperQwen’s reference draft vocabulary for speculative decoding. Swift 1.5 INT4-heads: quantized the output head and MTP linear layers to GPTQ INT4 instead, and built a Swift-specific 65,536-token draft vocabulary. Setup If you want to try it yourself, point your coding agent at these setup instructions and ask it to set up Swift 1.5 + HyperQwen on your machine. All three models can be found here: https://huggingface.co/collections/daavidhauser/swift-for-hyperqwen-rtx-3090-benchmarks Big shoutout to UkisAI for Swift and syv-ai for HyperQwen! It's genuinely insane to be able to run these models on an RTX3090 at those speeds!   submitted by   /u/KingGongzilla [link]   [comments]