QWEN-AUDIO-3.1-ASR

Beyond words.
Speakers. Sound. Meaning.

From multilingual transcription to who said what, when—and what the audio means. A model family for speech recognition, speaker diarization and audio understanding.

30 languages16 Chinese dialectsSpeaker-diarized ASR Flash / NextAudio understanding Next
Qwen-Audio-3.1-ASR capability overview
One model family. Capabilities tailored to each deployment profile. Open the vector figure ↗

Evaluation highlights

One model,
balanced performance.

10.38%

Macro CER on the internal 16-dialect evaluation suite.

4.57

Macro error rate on Common Voice 15 across the evaluated languages.

11/15

Industry domains where the model reports the outright highest entity recall, with one additional tie.

Interactive audio atlas

Listen across dialects and languages.

Choose a location on either map to hear the original sample and read its source-language transcript with an English translation.

30supported languages
Select a location

Japanese · Japan

日本語

Choose a marker on the map or a name below.

Flexible language prompting

Combine target languages at runtime.

language_hints = ["zh", "ja", "en"]

Browse all 30 supported languages
Internal dialect ASR · CER (%) ↓ · 16 dialects
Internal dialect ASR · CER (%) ↓ · 16 dialects
Internal dialect AST · Semantic sentence accuracy (%) ↑ · 11 dialects
Internal dialect AST · Semantic sentence accuracy (%) ↑ · 11 dialects
Public benchmarks · KeSpeech (AST) and WSYue (ASR) · CER (%) ↓
Public benchmarks · KeSpeech (AST) and WSYue (ASR) · CER (%) ↓

Fast & streaming

Listen while the transcript appears.

Play the 18-second Beijing-to-Hangzhou request and watch the transcript appear continuously in a clean side-by-side streaming interface.

Original 3.0 comparison replay · the audio, text and timing are preserved.

Interactive streaming playbackIndustry Voice InputvsQwen-Audio-3.1-ASR
Beijing → Hangzhou travel requestVoice-timbre adjusted · original comparison replay
0:00/0:18
INDUSTRY VOICE INPUTQWEN-AUDIO-3.1-ASR

Samples from the report and launch material

See what the controls change.

Explore context corrections, entity recall, hotword gains, and native polishing without leaving the page.

Established ASR capabilities · examples and results retained from the 3.0 report.

01 / SPEAKER-DIARIZED ASR

Who said what. Exactly when.

Five speakers, overlapping turns, one synchronized transcript. Select a segment to replay it. Available in non-streaming Flash and Next.

A five-person meeting

Recorded example · supplied model output, not live inference

Consistent speaker labelsSegment-level timestampsOverlapping speech on separate tracks

02 / GENERAL AUDIO UNDERSTANDING · NEXT

Hear the scene. Understand the event.

Speech, environmental sound and music, interpreted through instructions. Explore seven supplied examples with their original audio and outputs.

03 / EVALUATION

Measured across speech and sound.

Speaker-diarized ASR uses the updated SSE and Next results across four benchmarks. Audio understanding and industry recall retain their previously reported results.

Speaker-diarized ASR

DER / cpWER (%) ↓ · Lower is better

DER measures speaker diarization errors; cpWER measures speaker-attributed transcription errors. Bold indicates the lowest value in each column: Next leads on five metrics and SSE on three.

Audio understanding

Accuracy (%) ↑. Qwen leads the listed MMAU comparison; Gemini-3.1-Pro scores higher on MMSU.

QWEN-AUDIO-3.1-ASR

Six audio understanding tasks

Qwen-Audio-3.1-ASR industry recall comparison across 15 domains

Dialect ASR / AST figures use the updated 3.1 results. Other evaluations below retain their previously reported values.

Evaluation evidence

ASR and AST evaluation results.

Each result view answers a specific performance question; the complete reported tables remain available below. Lower CER/WER is better, while higher recall and consistency are better.

CER comparison across 16 Chinese dialect test sets

Evaluation note. Internal dialect and industry-domain evaluations use internal test sets; KeSpeech and WSYue are public dialect benchmarks. Public benchmark results in Panel A are taken from official sources, while API systems in Panel B use a unified evaluation pipeline.

For WeNetSpeech-net, all results in Panel B and our model's results in Panel A use corrected reference transcripts, following WenetSpeech issue #63. Competing-system results in Panel A are retained as reported in their original sources.

Table 1 Public Chinese and English ASR benchmarks
Table 2 Multilingual benchmark results
Table 3 Hotword recall with and without conditioning
Table 4 Long-audio contextual corrections

04 / MODEL FAMILY

Choose the capabilities you need.

Five profiles, different capability boundaries. Speaker diarization and general audio understanding are not streaming features.

Capability matrix from the supplied release material. Refer to the service documentation for availability and API details.

Model experience

Choose the deployment profile.

Open the official Model Studio documentation for short-form, file-transcription, or real-time streaming integration.

Existing 3.0 integration documentation · the 3.1 capability matrix is shown above.