Hear more.
Miss less.

A large-scale, instruction-controlled ASR model for multilingual speech, 16 Chinese dialectal varieties, long-context transcription, and production-grade entity recognition.

30
languages
16
Chinese dialects
3B / 35B
active / total params
Overview of Qwen-Audio-3.0-ASR and its production-oriented capabilities
Figure 1 One instruction-controlled model for multilingual speech, streaming, context, hotwords, entities, and polishing.

Reported evaluation

One model,
balanced performance.

9.40%

Macro CER on the internal 16-dialect evaluation suite.

4.57

Macro error rate on Common Voice 15 across the evaluated languages.

11/15

Industry domains where the model reports the outright highest entity recall, with one additional tie.

Interactive audio atlas

Listen across dialects and languages.

Choose a location on either map to hear the original sample and read its source-language transcript with an English translation.

30supported languages
Select a location

Japanese · Japan

日本語

Choose a marker on the map or a name below.

Flexible language prompting

Combine target languages at runtime.

language_hints = ["zh", "ja", "en"]

Browse all 30 supported languages
CER across 16 Chinese dialects
Figure 6 · Recognition accuracy across the same 16 dialect varieties.
Dialect consistency across 16 Chinese dialects
Figure 7 · ASR and AST task-consistency rates.

Fast & streaming

Listen while the transcript appears.

Play the 18-second Beijing-to-Hangzhou request and watch the transcript appear continuously in a clean side-by-side streaming interface.

Interactive streaming playbackDoubao Voice InputvsQwen-Audio-3.0-ASR
Beijing → Hangzhou travel requestVoice-timbre adjusted · original comparison replay
0:00/0:18
DOUBAO VOICE INPUTQWEN-AUDIO-3.0-ASR

Evaluation evidence

Results, where they support the story.

Each result view answers a specific performance question; the complete reported tables remain available below. Lower CER/WER is better, while higher recall and consistency are better.

CER comparison across 16 Chinese dialect test sets

Evaluation note. Dialect and industry-domain evaluations use internal test sets. Public benchmark results in Panel A are taken from official sources, while API systems in Panel B use a unified evaluation pipeline.

Table 1 Public Chinese and English ASR benchmarks
Table 2 Multilingual benchmark results
Table 3 Hotword recall with and without conditioning
Table 4 Long-audio contextual corrections

Samples from the report and launch material

See what the controls change.

Explore context corrections, entity recall, hotword gains, and native polishing without leaving the page.