2 min read · NVIDIA Blog
Alibaba Open-Sources Qwen3.8 2.4-Trillion Parameter Mixture of Experts Model
Alibaba released open weights for Qwen3.8-2.4T-A95B with 2.4 trillion parameters and 95B active per token. Featuring a hybrid linear/full attention architecture and configurable reasoning depth, it serves at 4K tokens/sec/GPU on NVIDIA GB300 systems.
Why it matters. You can now run frontier-scale 1M token context reasoning agent workflows on self-hosted infrastructure using standard vLLM or SGLang engines.