UniCUE: Unified Recognition and Generation Framework for Chinese Cued Speech Video-to-Speech Generation

[Paper]   [Dataset]   [Code]

Anonymous authors

Abstract: Cued Speech (CS) enhances lipreading by incorporating hand coding, thereby assisting hearing-impaired individuals in better perceiving spoken language. The CS Video-to-Speech (CSV2S) task aims to convert CS videos directly into intelligible speech. While existing approaches typically adopt a two-stage pipeline—first performing CS Recognition (CSR) followed by text-to-speech synthesis—this strategy often suffers from error accumulation and semantic misalignment. To address these issues, we propose UniCUE, the first unified framework for directly generating speech from CS videos. UniCUE seamlessly integrates CSR to extract fine-grained visual-semantic cues that enable more accurate and fluent speech synthesis. Its core components include a pose-aware visual processor for capturing detailed lip-hand visual signals, a semantic alignment pool for precise visual-to-semantic mapping, and a VisioPhonetic adapter to facilitate cross-task representation fusion. Extensive experiments on our newly collected Chinese CS dataset demonstrate that UniCUE significantly outperforms existing methods, establishing a new state-of-the-art in the CSV2S task.

UniCUE Framework

Comparison with SOTA Methods (Normal-hearing Cuers)
GT Text
CMML
EcoCued
CSR (ours)
Lip2Speech
LipVoicer
CSV2S (ours)
UniCUE (ours)
wo men dou hen qi dai neng jian dao nin.
ni fu mu shen ti zen me yang.
zhe pi zuan jie de pin zhi cen chi bu qi.
wo men yao deng duo jiu.
wei qing wen wang lao shi zai ma.
you qu de ling hun guo ran wan li tiao yi.
Comparison with SOTA Methods (Hearing-impaied Cuers)
GT Text
CMML
EcoCued
CSR (ours)
Lip2Speech
LipVoicer
CSV2S (ours)
UniCUE (ours)
zhe ge di fang hen bu cuo.
zhe shi ta de na shou cai.
you liang ge ren bu ji ge.
hei yun ya cheng cheng yu cui.
wo xiang ni le.
miao biao huai le.
jin tian tian qi zhen hao.
suo ran wu wei.