Speech recognition
Qwen-Audio-3.0-ASR
Multilingual and dialect recognition with instruction control, hotwords, context and streaming transcription.
Alibaba · Speech & audio research
From listening and understanding to speaking and creating. Explore the Qwen-Audio series and our wider family of speech and audio projects.
Speech recognition
Multilingual and dialect recognition with instruction control, hotwords, context and streaming transcription.
Speech synthesis
Shape voices with natural-language instructions and fine-grained tags. Explore cross-lingual voice cloning, long-form speech and expressive delivery.
Speech recognition
Speech recognition for multilingual, dialect and domain-specific audio.
Speech synthesis
Explore the next generation of multilingual speech generation and voice cloning.
Voice conversation
Connect audio understanding and natural spoken interaction.
Speech synthesis
Streaming synthesis, zero-shot voice cloning and multilingual generation.
Spoken interaction
Multimodal language models for real-time, natural voice conversation.
Music generation
Explore music generation, continuation and music creation examples.
FunAudioLLM · Understanding & generation
The original FunAudioLLM project: multilingual speech understanding, emotion and event recognition, voice generation and voice interaction demos.
Explore inference, serving and application integration with FunASR. See each project for model availability and licensing.