Qwen3.8-27B Has 6.78M Downloads and an Apache 2.0 Licence, but Its Card Gives Training Data One Line
A.I. / news
Qwen3.8-27B Has 6.78M Downloads and an Apache 2.0 Licence, but Its Card Gives Training Data One Line
Alibaba's 27B dense model scores 61.7 on SWE-bench Pro by its own table. The card does not say what it was trained on or what hardware it needs.

Alibaba's Qwen3.8-27B is the most-downloaded model on this week's Hugging Face trending list, with 17,311 likes and 6,783,589 downloads as of October 9. The weights carry an Apache 2.0 licence, which permits commercial use, according to the model card.
The release dates from August. eWeek reported it on August 17, saying the weights went to Hugging Face and ModelScope the previous Friday.
What the card specifies
The model is dense, not a mixture of experts. It has 27B parameters, which Hugging Face's metadata counts as 28B in BF16. It has 64 layers and a hidden dimension of 5,120.
The layers alternate Gated DeltaNet and Gated Attention blocks. The card also lists multi-token prediction and a vision encoder, so the model reads text, images and video.
The native context window is 262,144 tokens. It extends to 1,000,000 tokens with YaRN rope scaling, and the card warns that static YaRN can hurt performance on shorter texts.
Thinking mode is on by default. The card says it can be switched off per request, and a reasoning_effort setting takes the values xhigh, medium and low.

What the card leaves out
The card's only statement on training data is the phrase "Pre-training & Post-training." It gives no sources, composition or volume.
It also gives no GPU memory figure. Its Docker example uses --gpus all and --shm-size 32g, and it recommends SGLang, vLLM and TokenSpeed for serving. Alibaba says the model can run on consumer hardware and, once quantised, possibly on a laptop, per eWeek. Independent coverage of that claim was not found in this review.
BF16 stores each parameter in 2 bytes, so 28B parameters is roughly 56 GB before any context cache. That is arithmetic from the card's parameter count, not a figure Alibaba published.
Benchmarks, all from Alibaba
The scores below are in the card and are vendor-supplied. The card compares the model with others, but this report did not check those rows against an independent run.
| Benchmark | Qwen3.8-27B |
|---|---|
| GPQA Diamond | 89.2 |
| LiveCodeBench v6 | 90.3 |
| Terminal Bench 2.1 (Terminus) | 73.0 |
| SWE-bench Pro | 61.7 |
| OSWorld-Verified | 84.3 |
| Humanity's Last Exam | 30.8 |
- OSWorld-Verified84.3 points
- Terminal Bench 2.173 points
- SWE-bench Pro61.7 points
- WebArena-Verified64.8 points
Source: Qwen3.8-27B model card, Hugging Face, accessed 2026-10-09; vendor-supplied
MLQ.ai, cited by eWeek, said that the release materials lack a like-for-like comparison against earlier Qwen models and rivals using the same prompts, harnesses and dates. eWeek added that strong coding scores may not carry over to long tool-call chains or shifting instructions.
The licence split
The 27B is Apache 2.0. Alibaba released a second model the same week, Qwen3.8-2.4T-A95B, a mixture of experts with 2.4 trillion total and 95 billion active parameters.
That larger model carries a custom Qwen3.8-Max licence. eWeek, citing the South China Morning Post, reports that a business running a qualifying model-as-a-service or AI work assistant needs a separate licence above $50 million in revenue over any 12 consecutive months. The requirement does not apply to internal use.
The two releases therefore do not share commercial terms. A developer who tests on the 27B and moves to the larger model should read the second licence first.
What is still open
Alibaba says it has released more than 460 Qwen-family models, with over 3 billion downloads and over 300,000 derivatives, per eWeek. Those are company figures.
For another 27B Apache 2.0 release this season, see Cloudflare's Clef. For the closed-weight side of small-model pricing, see Claude Haiku 5.5. No independent evaluation of the 27B's agent scores was found as of October 9.
Sources
More in A.I.
- 01OpenAI Fires Three Safety Staff Nine Days After Publishing Outside-Audit PrinciplesThe company confirmed the dismissals on Oct. 1 but has not said what information moved, who received it, or whether the three first raised concerns internally.
- 02Google's EmbeddingGemma 2 Adds Images, Video and Audio to a 740M EmbedderThe Apache 2.0 weights lift the code-retrieval score by 9.9 points but move the multilingual text score by only 0.21.
- 03Cloudflare's Clef Beats TypeSafe's Jev on Three Tests and Loses on OneThe Apache 2.0 decision model costs nearly six times as much per million tokens as Jev, and its smaller sibling trails badly on one intent benchmark.
- 04ChatGPT's Intelligent UI Skips the Pro Thinking Level and Older Desktop Apps, and Publishes No Usage FiguresOpenAI put GPT-6 into the Chat tab on October 7 with generated buttons, forms and charts. The announcement gives one speed number and no accuracy figure.