Today Qianwen released the Qwen-Audio-3.1 series, upgrading core ASR, TTS and
Realtime interaction models and introducing new audio-creation model
Qwen-Audio-3.1-TTS-Next and audio-understanding model Qwen-Audio-3.1-ASR-Next.
Five new speech models form an end-to-end audio stack covering understanding,
generation, interaction and creation; Qianwen cut Qwen-Audio pricing across the
board—TTS down about 70%, Realtime down about 85%, ASR down about 95%—citing
lower user costs.