Qwen3.8-27B Is Apache 2.0 With a 262,144-Token Window and Already Has 482 Finetunes, Including Cloudflare's Clef
A.I. / news
Qwen3.8-27B Is Apache 2.0 With a 262,144-Token Window and Already Has 482 Finetunes, Including Cloudflare's Clef
Alibaba's dense 27B model claims 61.7 on SWE-bench Pro against 53.4 for Opus 4.6 Max, and the model card says nothing about training data.

Alibaba's Qwen team released Qwen3.8-27B on Hugging Face in mid-August under Apache 2.0, and by Wednesday the model page listed 482 finetunes, 160 adapters, 21 merges and 1,347 quantized versions. One of the finetunes comes from Cloudflare.
The weights are downloadable without a gating form. The card lists the licence as apache-2.0, which eWeek reported permits commercial use, modification and redistribution subject to the licence's conditions.
What the Qwen3.8-27B model card specifies
The model is dense, not a mixture of experts, so all 27 billion parameters run on every token. The Safetensors metadata reads 28 billion parameters in BF16, which works out to roughly 56 GB of weights before any quantization. The card gives no hardware guidance. It recommends SGLang, vLLM or TokenSpeed for serving.
The network has 64 layers and a hidden dimension of 5,120. Three of every four layers use Gated DeltaNet, a linear-attention design, and every fourth uses gated attention. It also carries a vision encoder and was trained for multi-token prediction.
Context is 262,144 tokens natively, extensible to 1,000,000. Maximum output is 262,144 tokens for reasoning and 131,072 for the final response. Thinking is on by default, with a reasoning_effort setting of xhigh, medium or low.
The card describes training only as "Pre-training & Post-training". It names no data sources.
Qwen's benchmark claims against Opus 4.6 Max
The card compares the model with Qwen3.6-27B, Qwen3.7-Plus, Muse Glimmer-30B and Anthropic's Opus 4.6 Max. All figures are vendor-supplied, and the eWeek report said the release lacks a complete like-for-like comparison with earlier Qwen models.
| Benchmark | Qwen3.8-27B | Qwen3.6-27B | Opus 4.6 Max |
|---|---|---|---|
| SWE-bench Pro | 61.7 | 53.5 | 53.4 |
| Terminal Bench 2.1 | 73.0 | 63.4 | 78.2 |
| OSWorld-Verified | 84.3 | 63.9 | 72.7 |
| GPQA Diamond | 89.2 | 87.8 | 91.3 |
| HLE | 30.8 | 24.0 | 40.0 |
The model leads on the coding and computer-use rows and trails on the two knowledge rows. On Humanity's Last Exam it sits 9.2 points behind Opus 4.6 Max.
- Qwen3.8-27B61.7 points
- Qwen3.7-Plus57.6 points
- Qwen3.6-27B53.5 points
- Opus 4.6 Max53.4 points
- Muse Glimmer-30B51.2 points
Source: Qwen3.8-27B model card on Hugging Face, accessed 2026-10-07
The comparison is against Opus 4.6, not Opus 5.5, the model Anthropic's Haiku 5.5 page names as the top of its current range.

Clef and the finetune ecosystem
Cloudflare's Clef is a 27B model post-trained from Qwen3.8-27B, including its vision encoder. It is also Apache 2.0. It does not write text. It takes a state and a schema of typed questions and returns a probability for each allowed answer in one forward pass.
The Clef card reports a BFCL score of 98.5 and an MMLU score of 90.3, and tested it on a single H200 GPU. A smaller Clef-Flash scores 51.0 on GPQA Diamond against 48.0 for Clef. The card does not describe training data. The site covered OpenAI's Decisions API earlier, which targets the same decision-model idea.
The licence changes at the top of the family
The permissive licence stops at 27B. The larger Qwen3.8-2.4T-A95B, called Max, has 2.4 trillion total and 95 billion active parameters and a custom licence. Per eWeek, businesses offering a model-as-a-service or AI work assistant need a separate licence above $50 million in revenue over any 12 months. Internal use is exempt. Alibaba's API charges $2 per million input tokens.
Apache 2.0 is also the licence on Google's EmbeddingGemma 2, a much smaller 740M-parameter embedding model.
The model card gives no training-data details for the 27B model, and no hardware requirements beyond the names of serving engines.
Sources
More in A.I.
- 01LTX-2.5's Free Commercial Licence Stops at $10 Million in Group Revenue, and Its Card Asks for Contact DetailsLightricks' open-weight video and audio model is free for production use below that line, but the revenue test counts parent companies and affiliates, and the card publishes no benchmark scores.
- 02Gemini 4 Argon Costs $1.99 a Task at Promo Price Against $0.72 for GPT-6.1 Sol, and Is Not Yet on SaleGoogle's introductory $2 and $10 rates match OpenAI's Sol per token, but Artificial Analysis figures cited by eesel show Argon writing 62,000 output tokens a task where GPT-6 Astra writes 27,000.
- 03EmbeddingGemma 2 Embeds Text, Images, Audio and Video in 567MB of RAM on a Pixel 11 ProGoogle DeepMind's Apache 2.0 embedding model has 740M parameters in three modular pieces, and the only benchmark number its launch post prints is a 9.92-point gain on MTEB Code.
- 04Kolibri-1 Is Apache 2.0 and Fits on One B200, but Trails Qwen3.8 27B by 9.1 Points in GermanAleph Alpha's 78B-parameter mixture-of-experts model activates 3.46B parameters per token, and its own model card shows a larger dense Qwen ahead on every headline benchmark.