Qwen3-ASR-Flash
Qwen3-ASR-Flash is QwenCloud's hosted speech-recognition model for converting audio into text. The official model ID is qwen3-asr-flash, it automatically identifies 11 languages, and the current list price is $0.000035 per second of audio.
The circulated name “Qwen-Audio-3.0-ASR-Flash” does not match the current official catalog. Using the verified qwen3-asr-flash ID prevents teams from planning around a model name that cannot be copied into the API request.
QwenCloud's live model page identifies qwen3-asr-flash as an audio-input, text-output speech-recognition model, documents automatic language detection across 11 languages, shows a DashScope request example, lists $0.000035 per audio second, and shows a 100 RPM limit. A same-day Reddit post uses the longer Qwen-Audio-3.0-ASR-Flash label and raises professional transcription questions, but it is not the naming authority.
- Transcribe multilingual meetings, interviews, calls, and recorded media.
- Estimate hosted transcription cost directly from audio duration.
- Use automatic language identification when the input language is unknown.
- Compare domain vocabulary accuracy with an existing ASR pipeline.
The production-facing identifier is qwen3-asr-flash. Treat Qwen-Audio-3.0-ASR-Flash as a circulated release label or naming mismatch unless Qwen publishes it as a separate product ID.
The model accepts audio and returns text through DashScope MultiModalConversation. QwenCloud says it can identify and transcribe 11 languages, and the example exposes optional language identification and inverse text normalization controls.
- $0.000035 per second of audio on the current model page.
- 100 requests per minute on the current model page.
- A known language can be supplied as a hint; automatic language identification can also be enabled.
Qwen describes the model as accurate and robust in complex audio, but those are vendor claims. Evaluate word error rate, speaker overlap, specialized vocabulary, timestamp needs, punctuation, latency, and failure recovery on your own recordings before replacing an existing transcription route.
Qwen3-ASR-Flash FAQ
Common questions about Qwen3-ASR-Flash.
Is Qwen-Audio-3.0-ASR-Flash the API model name?+
No current official catalog record uses that exact API name. QwenCloud lists Qwen3-ASR-Flash and instructs callers to send model="qwen3-asr-flash". Use that verified ID unless Qwen publishes a separate record.
How much does Qwen3-ASR-Flash cost?+
QwenCloud currently lists $0.000035 per second of audio. Recheck the live model page before budgeting because hosted pricing and account limits can change.
How many languages does Qwen3-ASR-Flash support?+
QwenCloud says the model can automatically identify and recognize speech in 11 languages. Test the exact accents, noise conditions, and terminology in your workload rather than treating that language count as equal quality everywhere.