Qwen3.8-27B Tops Hugging Face With 7 Million Downloads
A.I. / news
Qwen3.8-27B Tops Hugging Face With 7 Million Downloads
Alibaba's Apache 2.0 vision-language model, released Aug. 14, beats Claude Opus 4.6 Max on SWE-bench Pro in Alibaba's own table and trails it on eight other rows.

Alibaba's Qwen team has the most-liked model on Hugging Face's trending list: Qwen3.8-27B, with 16,524 likes and 7,020,239 downloads as of Tuesday, according to Hugging Face's model API. The weights are not gated and carry an Apache 2.0 licence, which permits commercial use.
Qwen released the model on Aug. 14, The Decoder reported. The same release included a larger variant, Qwen3.8-2.4T-A95B, which the outlet described as the model for Max-level work.
What the model card says about the 27B
Qwen3.8-27B is a dense model with 27 billion parameters, so every parameter is used for every token. Its layers alternate three Gated DeltaNet blocks with one gated-attention block, repeated 16 times. It accepts images and video as well as text.
The native context window is 262,144 tokens, extensible to 1,000,000 using RoPE scaling. Reasoning is on by default and can be switched off per request. The card recommends SGLang, vLLM or TokenSpeed for serving.
The card gives no memory requirement and no statement of what data the model was trained on. It refers to pre-training and post-training stages without detail.

The benchmark table is Alibaba's
Every figure below comes from Qwen's own model card, not from an independent lab. The card says SWE-bench Pro was run with the Claude Code harness at temperature 1.0 and a 256K context window.
- Qwen3.8-27B61.7 %
- Qwen3.7-Plus57.6 %
- Qwen3.6-27B53.5 %
- Claude Opus 4.6 Max53.4 %
Source: Qwen3.8-27B model card on Hugging Face, accessed 2026-09-29
The same table is less flattering elsewhere. On Terminal Bench 2.1, Qwen3.8-27B scores 73.0 against 78.2 for Opus 4.6 Max. On GPQA Diamond it scores 89.2 against 91.3, and on the HLE reasoning test 30.8 against 40.0.
| Benchmark | Qwen3.8-27B | Qwen3.7-Plus | Opus 4.6 Max |
|---|---|---|---|
| Terminal Bench 2.1 | 73.0 | 64.0 | 78.2 |
| GPQA Diamond | 89.2 | 90.3 | 91.3 |
| HLE | 30.8 | 34.7 | 40.0 |
| OSWorld-Verified | 84.3 | 73.3 | 72.7 |
The card also lists Qwen3.6-27B, the previous 27B release. On Terminal Bench 2.1 the new model gains 9.6 points, from 63.4 to 73.0. On HLE, Qwen3.8-27B (30.8) sits below the larger Qwen3.7-Plus (34.7), despite Qwen's claim that the 27B outperforms it in coding and office tasks.
Where it sits among open-weight releases
Prism ML's Ternary-Bonsai-2-27B, a 1.72-bit compression of this model, sits on the same trending list. For another open-weights release with a permissive licence, see Xiaomi's MiMo V2.6 under an MIT licence.
Download counts on Hugging Face measure fetches, not deployments. A count of 7 million includes automated pulls, quantization pipelines and repeat downloads.
Qwen said a hosted version with a 1,000,000-token context is coming to Qwen Cloud. It did not give a date. For how a closed frontier model is tested by an outside body, see the AI Security Institute's evaluation of GPT-6 Astra; no such independent evaluation of Qwen3.8-27B appears in the sources reviewed.
Sources
More in A.I.
- 01Amazon Releases Strands Decider 2B, an Apache 2.0 Decision Model Built on Qwen3.5-2BAWS's Strands Labs scores 72.3 percent on JevBench at a 106 ms median on an RTX 3090. TypeSafe's CEO calls the current crop of rivals less serious than his own team.
- 02OpenAI and Synopsys Sign Chip-Design Model Deal With No Customers or Benchmarks NamedGPT-Synopsys will run Synopsys EDA tools on OpenAI-hosted infrastructure under a revenue-sharing agreement. The Sept. 30 announcement gives no dollar figure and no ship date.
- 03FTC Confirms Investigation of OpenAI and Anthropic Over Consumer Risks, Names No Other CompaniesAn FTC spokesperson confirmed the Sept. 30 probe, which follows a White House accord signed Sept. 29. Neither company would comment.
- 04Ai2 Releases Olmo-core 3 With 2.7x MoE Training Speedup, Measured With Random RoutingThe open training stack reports 52,000 tokens per second per GPU on a 47B-parameter mixture of experts. The 1.2-trillion-parameter result is a benchmark configuration, not a trained model.