Tencent Hunyuan officially released the new generation voice recognition model Hy ASR 3.0 preview.
The Hy ASR 3.0 preview is now live on the Tencent Cloud official website, providing API services that can be widely used in scenarios such as intelligent customer service, content understanding, and voice search.
On August 4, Tencent Hunyuan officially released the new generation voice recognition model Hy ASR 3.0 preview. Currently, the Hy ASR 3.0 preview has been launched on the Tencent Cloud official website providing API services to the public, which can be widely applied in scenarios such as intelligent customer service, content understanding, and voice search. Yuanbao has deeply participated in the co-research and has completed the initial launch. Users can experience enhanced capabilities such as dialect recognition, intelligent context-based error correction, and stable transcription in complex environments by simply pressing and holding the button to speak; voice input is more accurate and stable, and it's freely accessible. Products like WorkBuddy are also being integrated gradually.
Based on the language understanding capabilities of the latest generation large language model Hy3, Hy ASR 3.0 preview integrates high-precision voice recognition with deep semantic understanding capabilities. It achieves comprehensive improvements across core dimensions such as general recognition, context awareness, multi-scenario robustness, and dialect coverage, providing accurate, coherent, and intent-aligned transcription results in more complex real-world inputs, evolving from "word-for-word transcription, single-point optimization" to "understanding context, compatible with scenarios, one-click output."
The Hy ASR 3.0 preview performs leading overall on multiple open-source benchmark datasets and self-built assessment sets. In open-source evaluations, the Hy ASR 3.0 preview maintains a Word Error Rate (WER) of around 3% across multiple languages, including a Mandarin Chinese WER of 3.34%, English WER of 2.62%, and Cantonese WER of 3.12%.
From the users' perspective, the Hy ASR 3.0 preview has enhanced four types of capabilities:
More accurate general recognition: Improved recognition accuracy for general speech, dialects, and mixed Chinese-English speech scenarios, further reducing typos, missed words, and error accumulation in long audio.
Better understanding of user intent: Precise capturing of user context through context integration, intelligently correcting homophones, and eliminating semantic ambiguity.
Easier adaptation to professional scenarios: Support for hotword injection enhancement, helping the model quickly recognize brand names, personal names, and industry terms, reducing business integration and ongoing maintenance costs.
More stable in complex environments: Specialized optimization for various acoustic scenarios such as high noise, whispers, and soft talk, maintaining stable performance under complex conditions.
Related Articles
KARRIE INT'L (01050) spent HK$372,200 to repurchase 200,000 shares on August 4.

TECHTRONIC IND (00669) announced its interim results, with a profit attributable to shareholders of $738 million, an increase of 17.52% year-on-year.

Jiangsu Hengrui Pharmaceuticals (600276.SH) subsidiary has obtained the drug clinical trial approval notice.
KARRIE INT'L (01050) spent HK$372,200 to repurchase 200,000 shares on August 4.
TECHTRONIC IND (00669) announced its interim results, with a profit attributable to shareholders of $738 million, an increase of 17.52% year-on-year.

Jiangsu Hengrui Pharmaceuticals (600276.SH) subsidiary has obtained the drug clinical trial approval notice.

RECOMMEND





