Alibaba launches Qwen Audio 3.1 with new models and slashes AI audio prices by up to 95 percent

Sep 23, 2026

Alibaba's AI team Qwen has released Qwen-Audio-3.1, a lineup of five models for speech recognition (ASR), text-to-speech (TTS), and real-time interaction. The ASR model improves multilingual and dialect recognition and automatically cleans up filler words and repetitions. ASR-Next adds multi-speaker identification with timestamps and detects emotions, ambient sounds, and machine noise. TTS handles multilingual synthesis with natural cross-language voice transfer. Users control emotion, speed, and style through simple text prompts like "Read this with a sharp, commanding tone, demanding respect."

TTS-Next pairs a language model with a diffusion approach to generate voice, sound effects, and background audio in a single pass. The real-time model supports simultaneous speaking and listening with instant interruption. When it detects a low mood, it responds more slowly and with more empathy, according to Qwen. Alibaba is also slashing prices. TTS drops about 70 percent, Realtime roughly 85 percent, and ASR up to 95 percent. More details on the blog and on Qwen Cloud.

Ad

Ad

AI News Without the Hype – Curated by Humans

Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.

Subscribe now

Source:

via X

← 返回资讯列表