Xiaomi released and open-sourced industrial-grade target-speaker automatic
speech recognition (ASR) large model Xiaomi-CocktailASR-1. The model targets the
cocktail-party problem, uses an end-to-end LLM architecture, and takes a short
reference audio clip of the target speaker as a voiceprint prompt to isolate and
transcribe only that speaker’s speech in multi-speaker, noisy environments.