The GenAI Field Guide · Trending
Alibaba's Qwen team published downloadable weights for its Max-class 2.4 trillion parameter model and for a 27B dense model that runs on one workstation.
In August 2026 the Qwen team released two Qwen3.8 checkpoints you can download: Qwen3.8-2.4T-A95B, the model behind the hosted Qwen3.8-Max API, and Qwen3.8-27B, a dense vision-language model. The 27B is Apache 2.0 and fits on a single large GPU or a well-specified laptop once quantised; the 2.4T needs a multi-GPU cluster and comes under a custom licence with revenue thresholds. If you self-host for privacy, cost or control, the 27B is the one to test this week.
Qwen3.8 is a family, and only part of it is open. The hosted flagship, Qwen3.8-Max, launched on Qwen Cloud on 3 August 2026 with a 1M token context, vision input and built-in tools. Qwen then published two checkpoints on Hugging Face and ModelScope. Qwen3.8-2.4T-A95B is the base of Qwen3.8-Max and, per the Qwen GitHub repository, the first time a Max-class Qwen model has been released openly. Qwen3.8-27B is a smaller dense model aimed at deployment on your own hardware. A third repository, Qwen3.8-Flash-Next (125B total, 6B active), appeared on 24 August as an experimental preview of the architecture Qwen says will underpin Qwen4.
The open checkpoint is not identical to the paid API. The 2.4T model card says the hosted Qwen3.8-Max adds vision input, a non-thinking mode, 1M context by default and official built-in tools; the open 2.4T is text only and always thinks. The 27B keeps image and video input and lets you switch thinking off per request. Qwen's card also says a hosted 27B with 1M context and built-in tools is offered on Qwen Cloud.
The licences differ, and this is the detail most worth forwarding. Qwen3.8-27B is Apache 2.0, the same as earlier Qwen 27B releases. Qwen3.8-2.4T-A95B ships under a custom Qwen3.8-Max License: broad rights to use, modify, fine-tune and sell, but products above 100 million monthly active users or US$20 million monthly revenue must show the model name in the interface, and a model-as-a-service or AI coding or office assistant business with more than US$50 million revenue over twelve months needs a separate licence from Qwen. Flash-Next uses a Qwen Community License 1.0 that requires that separate licence for any such business, with no revenue floor. The Decoder's launch report described the release as Apache 2.0 across the board; the licence files on Hugging Face say otherwise for the two larger models.
Both open models use the hybrid layout Qwen introduced with Qwen3.5: most layers use Gated DeltaNet, a linear attention variant whose memory does not grow with sequence length, and every fourth layer uses gated full attention. The 27B card lists 64 layers in 16 blocks of three DeltaNet layers and one attention layer, with only 4 key-value heads in the attention layers. That mix is why it can advertise a 262,144 token native context and up to 1M tokens with YaRN scaling on a model this size.
The 2.4T is a mixture of experts: 512 experts per layer, 10 routed plus 1 shared active per token, so about 95B parameters do work on each token out of 2.4T stored. It is also trained with multi-token prediction, which serving engines can use for speculative decoding. Active parameters set the compute per token, but every expert has to sit in GPU memory, which is why the BF16 repository is about 4.9 TB and the FP8 one about 2.5 TB on Hugging Face. vLLM's recipe asks for 16 GPUs across two nodes for FP8, or one 8 GPU B300 node for a community NVFP4 quantisation.
Reasoning is on by default. Responses start with a think block, controlled through the chat template: enable_thinking turns it off for the 27B, reasoning_effort picks xhigh (the default), medium or low, and preserve_thinking, also on by default, keeps earlier turns' reasoning in context for multi-turn agent work. Qwen's own card warns that lower effort can make an agent slower overall because it fails and retries more. The cost of the default shows up in independent testing: VentureBeat reports Artificial Analysis measured 160 million output tokens for the 27B across its index against a 43 million median, and Artificial Analysis's own page now lists 200 million against an 82 million median.
On quality, Qwen's 27B card reports SWE-bench Pro 61.7, Terminal Bench 2.1 73.0, GPQA Diamond 89.2 and OSWorld-Verified 84.3, mostly ahead of its own hosted Qwen3.7-Plus, though Plus still leads on GPQA and HLE. Several rows are in-house benchmarks (QwenSWEBench, CoWorkBench) run in Qwen's own harness. Independent numbers disagree with each other: VentureBeat on 17 August reported an Artificial Analysis Intelligence Index of 52 for the 27B, while the Artificial Analysis model page read on 2 October shows 34, ranked first of 142 comparable models. We could not confirm why the two figures differ, so compare scores only within one snapshot of the same index.
Anyone can download the weights from the Qwen organisation on Hugging Face or ModelScope with no gating. For the 27B you need roughly 56 GB of GPU memory at BF16, about 28 GB at FP8 or about 17 to 18 GB at 4 bit, per VentureBeat and Ollama's listing. If you only want to call it, Qwen Cloud and several OpenRouter providers host both open models.