Qwen3.8-27B Ships Under Apache 2.0 and Fits in 17GB, but Spends 160 Million Tokens Where the Median Spends 43 Million
A.I. / news
Qwen3.8-27B Ships Under Apache 2.0 and Fits in 17GB, but Spends 160 Million Tokens Where the Median Spends 43 Million
Alibaba's open-weight model scores 52 on Artificial Analysis's Intelligence Index. Its own benchmark figures are vendor-supplied, and users report slow runs.

Alibaba's Qwen team has released Qwen3.8-27B, a 27-billion-parameter open-weight model under the Apache 2.0 licence, with weights downloadable from Hugging Face without a gating form.
The model card lists a native context window of 262,144 tokens, extensible to 1,000,000. It carries no named author beyond "Qwen Team" and no detailed account of training data.
Qwen3.8-27B licence, architecture and hardware
The card describes a causal language model with a vision encoder and 64 layers in a hybrid pattern: sixteen blocks of three Gated DeltaNet layers followed by one Gated Attention layer. Hidden dimension is 5,120.
VentureBeat reported the memory needs. Full 16-bit precision takes about 56GB of GPU memory, FP8 about 28GB and 4-bit quantisation about 17GB. The outlet said it ran on an M5 Max MacBook Pro and an Nvidia DGX Spark. The card lists 1,293 quantised variants, including builds for llama.cpp, LM Studio, Jan and Ollama.
The card does not state hardware requirements. It recommends SGLang, vLLM or TokenSpeed for serving.
Vendor benchmarks and one independent score
The Qwen team ran its own evaluations using a Claude Code harness and in-house tests. Those results are vendor-supplied.
| Benchmark | Score | Who ran it |
|---|---|---|
| SWE-bench Pro | 61.7 | Qwen team |
| GPQA Diamond | 89.2 | Qwen team |
| LiveCodeBench v6 | 90.3 | Qwen team |
| Intelligence Index | 52 | Artificial Analysis |
Artificial Analysis, the independent benchmarking firm, put the model at 52 on its Intelligence Index and 51 on its Agentic Index, VentureBeat reported.

The cost is in reasoning tokens
The same testing found the model used 160 million tokens across the evaluation, against a median of 43 million, VentureBeat reported. One task took 21 minutes at default settings.
Throughput in standard configurations ran at 15 to 30 tokens per second. VentureBeat put the model at about 30 times slower than DeepSeek V4 Flash with reasoning enabled.
- Qwen3.8-27B160 M tokens
- Median43 M tokens
Source: VentureBeat, Qwen3.8-27B coverage, accessed 2026-10-02
VentureBeat named the developer Simon Willison among those who tested the model, and quoted the coding-agent maker Cline as calling it the first local model to reach frontier capability. That is a claim about one vendor's benchmark mix, not a general ranking.
The larger Qwen3.8-Max stays out of the picture
Alibaba announced Qwen3.8-Max on Aug. 3. Latent Space listed it at 2.4 trillion total parameters with about 95 billion active per token, and said Alibaba had promised open weights "next week" alongside the 27B model.
Latent Space also reported licensing questions about geographic restrictions in the USA, EU, UK and Korea. It said Alibaba had not clarified them in the coverage. The Hugging Face card for the 27B model shows plain Apache 2.0, which has no regional carve-outs.
Two earlier open-weight stories on this site cover smaller releases: Amazon's Strands Decider 2B, also built on a Qwen base, and Ai2's Olmo-core 3. The next check is whether independent harness runs reproduce the Qwen team's 61.7 on SWE-bench Pro.
Sources
More in A.I.
- 01OpenAI Ties Moonshot AI to a July Campaign That Replayed Encrypted Reasoning, Offers No Evidence PubliclyOpenAI says 16,000 requests from more than 4,000 accounts tried to recover hidden model reasoning. Its attribution to Moonshot rests on its own assertion.
- 02Gemini 4 Argon Leads 13 of 18 Benchmarks Google Chose, but Only Cyber Defenders Can Use ItGoogle priced the model at $2 and $10 per million tokens and gave access first to its Fairwind Program, with a guardrail-free version for trusted defenders.
- 03Amazon Releases Strands Decider 2B, an Apache 2.0 Decision Model Built on Qwen3.5-2BAWS's Strands Labs scores 72.3 percent on JevBench at a 106 ms median on an RTX 3090. TypeSafe's CEO calls the current crop of rivals less serious than his own team.
- 04OpenAI and Synopsys Sign Chip-Design Model Deal With No Customers or Benchmarks NamedGPT-Synopsys will run Synopsys EDA tools on OpenAI-hosted infrastructure under a revenue-sharing agreement. The Sept. 30 announcement gives no dollar figure and no ship date.