Underdog, the on-device assistant from Conway Research, has released Saluki 27B under Apache 2.0. Underdog Saluki 27B is a 2-bit GGUF of Qwen3.8-27B that fits in 7.89 GB. The full BF16 model needs 54 GB. Underdog tuned the compression to protect tool calling, the skill that turns a chat model into an agent. For developers, that means a 27B-class agent model that runs in stock llama.cpp.
TL;DR
Size: 27B dense parameters. 7.89 GB GGUF versus 54 GB for BF16.
Runs on: stock llama.cpp and apps built on it, with full GPU offload. Optional 629 MB or 928 MB vision add-on.
Performance: 96% average retention across 9 benchmarks versus full Qwen3.8-27B.
Best: Parallel tool calls, 42 versus 35 for the full model (120% retention).
Worst: AIME 2025, 79.2 versus 96.7 (about 82% retention).
Bottom line:
Best: beats the 54 GB original at tool calling in a sub-8 GB file.
Worst: competition math and multi-step reasoning drop 12 to 18 points.
What is Underdog Saluki 27B?
Saluki 27B is a 2-bit, mixed-precision GGUF of Qwen3.8-27B built for local agents. It stacks 3 layers of work:
The base is Qwen3.8-27B, a dense 27B model from the Qwen team. It has 64 layers, mixes Gated DeltaNet linear attention with gated attention, and supports 262,144 tokens natively.
The second layer is ISTA-DASLab’s Qwen3.8-27B-GSQ-RCO-GGUF. GSQ learns accurate low-bit scalar grids per tensor. RCO assigns a quantization type to each tensor under a fixed size budget. ISTA’s smallest file, IQ2_XS, is 8.4 GB at 2.50 bits per weight.
The third layer is Underdog’s own pass. It shrank the file to 7.89 GB and targeted tool calling. The file is named IQ2-mix and carries an imatrix tag. Underdog has not published the full recipe for this pass.
How does Saluki perform on benchmarks?
Underdog splits its results into 2 groups.
The first group ran both models in the same harness:
Underdog Bench: 120 tasks from BFCL v4, frozen before testing. Thinking off, temperature 0. Saluki scores 88, the full model 84, and PrismML’s Bonsai 2 scores 70.
Parallel tool calls: 100 BFCL v4 parallel tasks with the official checker. Saluki 42, full model 35.
SWE-bench Verified: 50 issues. Saluki fixes 30, the full model 33.
The second group compares Saluki with public full-size scores:
Benchmark
Saluki 27B
Qwen3.8-27B (public)
IFEval (prompt-loose)
93.5
91.5
IFBench (prompt-loose)
72.7
71.0
MBPP+
78.0
83.9
MuSR
67.5
79.6
AIME 2025 (avg@4)
79.2
96.7
AIME 2026 (avg@4)
80.0
94.6
How does Saluki compare with other compact Qwen3.8-27B builds?
Feature
Underdog Saluki 27B
Qwen3.8-27B (BF16)
ISTA GSQ-RCO IQ2_XS
PrismML Bonsai 2 27B (PTQ1_0)
Org
Underdog (Conway Research)
Qwen team
ISTA-DASLab
PrismML
Parameters
27B
27B
27B
27.36B
File size
7.89 GB
54 GB
8.4 GB
5.95 GB
Bits per weight
Not disclosed (2-bit mix)
16
2.50
1.75
Context
Not disclosed
262,144 native
Not disclosed
262K
Vision
Add-on, 629 or 928 MB
Native
0.9 GB mmproj
Optional 0.63 GB
Runtime
Stock llama.cpp
Transformers, vLLM, SGLang
Stock llama.cpp, Ollama, LM Studio
PrismML llama.cpp fork
Underdog Bench (of 120)
88
84
Not disclosed
70
Vendor headline
96% avg retention, 9 benchmarks
Baseline
100.3% zero-shot recovery
98.2% of FP16, 14 benchmarks
License
Apache 2.0
Apache 2.0
Apache 2.0
Apache 2.0
Bonsai 2 is smaller and reports 98.2% retention across 14 thinking-mode benchmarks. It also posts stronger math, including 95.00 on AIME25. However, it needs PrismML’s llama.cpp fork, since stock llama.cpp rejects its packing formats. Saluki runs on stock llama.cpp. Each vendor uses its own harness, so cross-vendor scores are not directly comparable.
Key Takeaways
Saluki 27B fits Qwen3.8-27B into 7.89 GB, down from 54 GB.
It beats the full model on tool calling: 88 versus 84.
Parallel tool calls rise to 42 from 35.
Math and reasoning take the biggest hit, down to about 82 to 85%.
It runs in stock llama.cpp under Apache 2.0.
Check out the Model Card on Hugging Face, the Underdog launch page and the announcement on X. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
The post Meet the Underdog Saluki 27B: A 2-bit Qwen3.8-27B That Beats the Original at Tool Calling appeared first on MarkTechPost.