The Qwen3.8-27B Variant Built to Stop Overthinking

Wait 5 sec.

Swift-1.5-Qwen3.8-27B-GGUF is a set of GGUF quantizations of Swift 1.5 Qwen3.8-27B, a text-generation and reasoning model derived from Qwen3.8-27B. UkisAI (maintainer profile) trained the Swift model to reduce pathological overthinking: its reported evaluation uses 58.5% fewer mean thinking tokens than the base on GPQA-Diamond, with a 0.31 percentage-point higher score, and the release reports a 9.18× speed-up on several tasks. That speed-up is a reported result, not a general guarantee for every prompt, quantization, or machine. The model has 27B parameters according to its name; the supplied material does not specify its architecture details, hardware requirements, or a context limit for this GGUF release. Run it with a current llama.cpp-compatible runtime such as llama-server. The key decision is whether lower reasoning-token use and local GGUF deployment suit your workload: Swift improves some coding and agent benchmarks, but trails the base on several reasoning and math scores.Best use casesCoding tasks that benefit from extended reasoning. Swift 1.5 scored 81.71% on LiveCodeBench v6, compared with 76.76% for Qwen3.8-27B, while using 24.5% fewer mean reasoning tokens. This makes it a promising candidate for code generation and problem-solving workloads where you can test outputs against a compiler or test suite. The benchmark does not establish performance on every programming language or repository-scale task.Tool-using and multi-step agent tasks. On Terminal-Bench 2.1, Swift scored 72.13% versus 69.21% for the base, with mean reasoning tokens across agent calls falling from 52,265 to 43,733. The training description says the post-training focused on long-horizon and agentic tasks. Evaluate it with your own tools and task limits before relying on it in production.General reasoning with a token budget. On GPQA-Diamond, Swift scored 88.59% versus 88.28% for the base, while mean tokens fell from 15,014 to 8,717. The result supports testing Swift where reasoning-token consumption matters and a small score gain on this benchmark is useful.Local inference with selectable memory and quality tradeoffs. The repository offers GGUF tiers from Q8_0 at 29.0 GB down to IQ2_XXS at 8.9 GB. That range lets practitioners choose a file size and quantization tradeoff for their llama.cpp setup. The provided data measures quantization divergence against the BF16 source, not end-to-end task quality for every tier.LimitationsSwift does not win every benchmark. It trails Qwen3.8-27B on IFBench (72.07% vs. 73.53%), ERQA (65.40% vs. 67.45%), AIME 2026 (96.00% vs. 98.67%), and HMMT November 2025 (97.33% vs. 99.33%). If your workload prioritizes those tasks, compare both models on representative prompts rather than assuming fewer tokens preserve accuracy.The reported 9.18× speed-up applies to “several tasks”; the README does not provide a general throughput figure, tokens per second, latency distribution, or hardware details for that claim. The game-building demo is one example: the base took 104.6 minutes and Swift 11.39 minutes. It is not a controlled estimate of performance for other tasks or systems.The GGUF file sizes range from 8.9 GB to 29.0 GB, but the README gives no VRAM or RAM requirements. File size alone does not establish the memory needed to run a model, especially with a long context or KV cache. The evaluation used context 262,144 for the main BF16 serving setup and 131,072 for Terminal-Bench; those are evaluation settings, not a stated maximum context for this GGUF release.Quantization can change behavior. The supplied quantized benchmark is a single-seed evaluation on three datasets, and its authors caution that it does not establish quality parity or replace the broader multi-seed results. At the lowest GGUF tiers, the reported divergence from BF16 rises: IQ2_XXS has KLD 0.2769 at 32k and a 78.48% top-p figure, compared with Q8_0 at 0.0006 and 98.85%. These are distribution-comparison metrics, not direct task-accuracy scores.The model card lists the license as “other” but does not provide the terms in the supplied material. Commercial-use rights therefore cannot be confirmed here. The material also does not specify safety evaluations, known bias findings, fine-tuning instructions, or a maintenance schedule.How it comparesSwift-1.5-Qwen3.8-27B-GSQ-RCO-GGUF: Choose this standard GGUF when you want the listed llama.cpp quantization tiers and their reported file sizes and divergence metrics. Choose the GSQ-RCO variant if its compact mixed-precision quantizations better fit your deployment constraints; the supplied comparison does not include a controlled quality, speed, or cost comparison between the two.Swift-Qwen3.8-27B-GGUF: Choose Swift 1.5 when coding and agentic performance matter: its release reports higher LiveCodeBench v6 and Terminal-Bench 2.1 scores than the base and describes it as an upgrade over Swift 1.0. Choose Swift 1.0 if you specifically need that earlier checkpoint; its description reports 58.3% fewer thinking tokens, under 1% performance loss, and a 1.95× speed-up on several tasks, but those figures are not a direct matched comparison with Swift 1.5.Swift-Qwen3.8-27B: This is the earlier Swift model rather than the GGUF package covered here. Pick the current GGUF when you want Swift 1.5’s reported coding and agent benchmark gains in a llama.cpp-compatible format; pick the earlier model if your workflow depends on that specific checkpoint. The supplied information does not provide a controlled speed or quality comparison between these two Swift versions.ukisai_Swift-Qwen3.8-27b-GGUF: This is a third-party imatrix quantization of the earlier Swift-Qwen3.8-27B, using llama.cpp release b10896 according to the supplied description. Choose the model in this guide for Swift 1.5 and its listed quantization tiers; choose the third-party package if you need that earlier model’s imatrix quantizations. No matched quality, speed, or cost results are provided.Qwen3.8-27B-GGUF: Choose Swift 1.5 when its lower reasoning-token use and stronger reported coding and agent scores fit your workload. Choose this Qwen GGUF if you want the base model in Unsloth’s Dynamic V3.0 preview quantization. The supplied material does not provide a controlled comparison of their GGUF quality, speed, or hardware cost.Technical specificationsSwift-1.5-Qwen3.8-27B-GGUF is derived directly from Swift 1.5 Qwen3.8-27B, itself a reasoning-efficient derivative of Qwen3.8-27B. The training approach identifies tokens associated with pathological overthinking, penalizes them without directly targeting reasoning length, then uses reinforcement learning (RL) and outcome-based preference optimization (OPD) to regain accuracy. Swift 1.5 scales up post-training methods used for Swift 1.0, with emphasis on long-horizon, agentic, and coding tasks. The linked training dataset is described as a source for resampling and constructing RL environments; it is not used out of the box. No dataset size, training-step count, or compute budget is supplied.The main evaluation reports final aggregate percentages from five repeats under matched protocols. Serving used BF16, vLLM 0.27.1, a Qwen3 parser, context 262,144, and xhigh thinking. Sampling used temperature 1.0, top-p 0.95, top-k 20, min-p 0, presence penalty 0, and repetition penalty 1. Benchmarks averaged five seeds (04) per model; Terminal-Bench used five trials per task, with both models served at context 131,072 on the same Harbor build. Output caps were 100,000 tokens for GPQA-Diamond, 16,384 for C-Eval, 81,920 for IFBench, 100,000 for ERQA, 250,000 for AIME 2026, 250,000 for HMMT November 2025, and 32,768 for LiveCodeBench v6. The supplied table gives no numeric output cap for Terminal-Bench.Reported base-to-Swift scores and mean-token counts:BenchmarkQwen3.8-27BSwift 1.5Mean tokens: base → SwiftGPQA-Diamond88.28%88.59%15,014 → 8,717C-Eval90.00%90.92%1,492 → 819IFBench73.53%72.07%8,052 → 4,955ERQA67.45%65.40%4,137 → 1,906AIME 202698.67%96.00%22,014 → 13,203HMMT November 202599.33%97.33%22,032 → 14,957LiveCodeBench v676.76%81.71%11,184 → 8,448Terminal-Bench 2.169.21%72.13%52,265 → 43,733The README also reports reasoning-effort results: at xhigh, GPQA-Diamond scores were 88.28% for the base and 88.59% for Swift, with 41.9% mean thinking-token reduction; at medium, scores were 84.14% and 82.22%, with 24.8% reduction; at low, scores were 84.04% and 84.85%, with 28.7% reduction.Available quantization formats include GGUF for llama.cpp; GSQ-RCO GGUF, described as compact 23-bit; AWQ INT4 (W4A16) and AWQ + GPTQ INT4 (W4A16) for vLLM compressed-tensors; AutoRound INT4 (W4A16) for vLLM auto-round; NVFP4 for NVIDIA Blackwell; AMD Quark FP8 (W8A8); and MLX 5-bit, 4-bit, and 3-bit text-only variants for Apple MLX. The GGUF file-size and divergence table is:GGUF tierSizeKLD wikitext @512KLD wikitext @32k99% KLD @32kTop-p @32kQ8_029.0 GB0.00080.00060.00598.85%Q6_K_L25.0 GB0.00150.00140.01098.16%Q6_K23.9 GB0.00180.00160.01498.30%Q6_K_S22.9 GB0.00200.00160.01498.24%Q5_K_M20.9 GB0.00520.00610.05096.92%Q5_K_S19.6 GB0.00600.00690.05896.91%Q4_K_L18.8 GB0.01060.01030.10595.79%Q4_K_M17.4 GB0.01370.01340.16395.03%IQ4_NL17.4 GB0.01520.01400.17595.39%Q4_K_S16.4 GB0.01640.01540.17594.84%IQ4_XS15.5 GB0.01790.01730.18794.96%IQ3_M14.9 GB0.04100.03800.40991.83%Q3_K_L14.1 GB0.04420.04100.41591.63%Q3_K_M13.4 GB0.05700.05620.61490.30%IQ3_XS12.8 GB0.05830.08851.13088.54%Q3_K_S12.7 GB0.06480.06580.71289.53%IQ3_XXS12.3 GB0.07420.08440.99688.70%Q2_K10.8 GB0.16550.15461.72884.00%IQ2_M10.5 GB0.15230.14931.56884.17%IQ2_S9.7 GB0.20950.25893.02480.58%IQ2_XS9.1 GB0.24330.26222.95179.99%IQ2_XXS8.9 GB0.28660.27692.90278.48%The KLD figures compare each quantization with the Swift 1.5 BF16 source. The wikitext @512 measure uses 100 windows of 512 tokens from the wikitext-2 test set. The supplied description begins to define wikitext @32k but does not include the rest of that definition.A separate quantized evaluation used vLLM 0.29.0, tensor parallelism 1, eager execution, BF16 activations, context 131,072, template-default thinking, and seed 0. Sampling used temperature 1, top-p 0.95, top-k 20, min-p 0, presence penalty 0, and repetition penalty 1. It tested 198 GPQA-Diamond questions, 300 IFBench prompts, and 30 AIME 2026 problems, with one sample per prompt and zero request errors. This was a single-seed evaluation, not a replacement for the five-repeat BF16 results. The AMD Quark INT4 and FP8 exports have separate sanity evaluations; completed results on these three reasoning benchmarks are not available.The model is tagged Text-to-Text and has an image-text-to-text pipeline tag, but the supplied README does not document image input handling for this GGUF package. Do not assume multimodal support from the tag alone. The model card lists the license as other. The repository reports 21,524 downloads.Model inputs and outputsInputsText prompts for text generation and reasoning. The supplied material does not specify a required chat template or a maximum context for this GGUF release.The model has an image-text-to-text pipeline tag, but the README does not document image formats, image preprocessing, or image support in the GGUF runtime.OutputsGenerated text, including reasoning and responses for coding, general reasoning, mathematics, and agent tasks.No structured output schema, image output, or post-processing requirements are specified.Getting startedThe README recommends a current llama.cpp-compatible runtime such as llama-server, but it does not provide a model-loading command, Python API example, or inference code. A Python snippet would require assumptions about the runtime and chat template, so no copy-pasteable Python example can be confirmed from the