INTRODUCING QWEN-AUDIO-3.1-REALTIME

Beyond conversation.
Into action.

A voice that reasons. An assistant that acts.
Bringing spoken intelligence, real-world execution, and natural
interaction into one continuous conversation.

Qwen-Audio-3.1-Realtime by Alibaba Token Foundry
Qwen-Audio-3.1-Realtime at the center of seven connected capabilities: multilingual reasoning, long-context instructions, tool execution, persona and empathy, full-duplex coordination, reliability, and persistent voice agents.

Model capabilities and system design in one view. Persistent tasks and memory are runtime extensions; matched 3.1 duplex measurements are not yet reported.

MULTI-TURN INSTRUCTIONS52.21%

Audio MultiChallenge

+5.09 percentage points vs. 3.0
SPOKEN TASK EXECUTION82.19%

τ²-Bench Audio

+3.59 percentage points vs. 3.0
MULTILINGUAL REASONING88.09%

14-language QA overall

+6.40 percentage points vs. 3.0
EMPATHETIC RESPONSE4.03/5

EchoMind

+0.33 score points vs. 3.0
UNDER THE HOOD

Technical Overview

MEASURED, NOT JUST DESCRIBED

A stronger foundation.
A more capable voice agent.

Selected results from the technical report. Explore each dimension
with exact values, metric definitions, and comparison scope.

GPT-Realtime-2 uses low effort unless stated otherwise. Results reflect the report's evaluation settings, not controlled training ablations or statistical significance. A missing result is not treated as zero. S2T results evaluate speech input with text output, not synthesized speech quality.

REAL-WORLD EXPERIENCES

See It In A Real Conversation.

BEYOND THE MODEL · A SYSTEM EXTENSION

Keep talking.
The work continues.

A persistent voice-agent runtime keeps the conversation responsive while longer tasks move forward—with explicit state, bounded memory, and results delivered at the right moment.

Persistent voice harness architecture: a foreground continuous-interaction layer with the user, realtime foreground agent, direct tools and bounded context; an orchestration runtime tracking task records and permissions; and a background durable-execution layer with a background agent, tools and environment, and separate stores.
↗

Direct or delegated

Keep short actions in the foreground. Hand multi-step work to a background executor through an explicit task contract.

◎

State you can trust

Track acceptance, execution, completion, and delivery separately. Follow-up inputs remain linked to the original task.

⊞

Memory with boundaries

Separate user preferences, long-term memory, reference knowledge, and live task state. Current instructions take precedence.

Scenario demos

QWEN-AUDIO-3.1-REALTIME

Towards reliable
agentic voice interaction.

Qwen-Audio-3.1-Realtime frames realtime interaction as Think, Act, and Speak & Coordinate — reasoning over evolving requests, executing verifiable actions, and governing how, when, and whether to speak — with evaluation mirroring each layer. Against 3.0, the reported setting gains 5.09 percentage points on Audio MultiChallenge, 6.4 points on the 14-language Big Bench Audio extension, and 3.6 points on τ²-Bench Audio, while multi-turn attack success falls by 54.50 points in Chinese and 40.00 points in English.

Explore the voice agent
Alibaba Token Foundry · Alibaba Group