Alibaba’s Qianwen released Qwen3.8-Flash, a multimodal Mixture-of-Experts (MoE)
model and an early preview of the Qwen4 architecture. Production Qwen3.8-Flash
will be available via the Qwen Cloud API, priced at $0.16 per 1M input tokens
and $0.47 per 1M output tokens. The model includes 125 billion parameters plus
51 billion N-gram embedding parameters, with roughly 6 billion parameters
activated per token to boost cost efficiency.