Alibaba Releases Open Weights for Qwen3.8-Max, First Max-Class Model in Qwen Family | Model Release
Alibaba's Qwen team released open weights for Qwen3.8 2.4T A95B (Qwen3.8 Max) on August 12, 2026 via Hugging Face and ModelScope, with day 0 support on NVIDIA…
Published on MyPrivateClaw
Aug 13, 2026, 4:47 AM UTC
Coverage date
Aug 12, 2026
Last updated
Aug 13, 2026, 4:47 AM UTC
News summary
Qwen3.8 2.4T A95B is a causal language model with 2.4 trillion total parameters and 95 billion active per token, using a fine grained mixture of experts architecture (512 experts, 10 routed + 1 shared) with a hybrid of full attention and linear gated delta networks, native context length of 262,144 tokens extensible to approximately one million tokens Day 0 on NVIDIA GB300 NVL72 in FP8 precision, Qwen3.8 2.4T A95B delivers over 4K tokens per second per GPU and over 350 tokens per second per user using TensorRT LLM inference runtime Qwen3.8 2.4T A95B is officially supported on SGLang, vLLM, and TokenSpeed inference frameworks; NVIDIA Dynamo and model free NVIDIA NIM containers are also provided as serving options The model supports configurable reasoning depth through a reasoning effort parameter with three levels: xhigh (default), medium, and low Qwen3.8 Max achieves 93.0 on PaperBench…