ABSTRACT
The automatic movie dubbing model generates vivid speech
from given scripts, replicating a speaker’s timbre from a
brief timbre prompt while ensuring lip-sync with the silent
video. Existing approaches simulate a simplified workflow
where actors dub directly without preparation, overlooking
the critical director–actor interaction. In contrast, authentic workflows involve a dynamic collaboration: directors actively engage with actors, guiding them to internalize the
context cues, specifically emotion, before performance. To
address this issue, we propose a new Retrieve-Augmented
Director-Actor Interaction Learning scheme to achieve authentic movie dubbing, termed Authentic-Dubber, which
contains three novel mechanisms: (1) We construct a multimodal Reference Footage library to simulate the learning
footage provided by directors. Note that we integrate Large
Language Models (LLMs) to achieve deep comprehension
of emotional representations across multimodal signals. (2)
To emulate how actors efficiently and comprehensively internalize director-provided footage during dubbing, we propose
an Emotion-Similarity-based Retrieval-Augmentation strategy. This strategy retrieves the most relevant multimodal information that aligns with the target silent video. (3) We
develop a Progressive Graph-based speech generation approach that incrementally incorporates the retrieved multimodal emotional knowledge, thereby simulating the actor’s
final dubbing process. The above mechanisms enable the
Authentic-Dubber to faithfully replicate the authentic dubbing workflow, achieving comprehensive improvements in
emotional expressiveness. Both subjective and objective evaluations on the V2C-Animation benchmark dataset validate
the effectiveness. The source code and model checkpoints
will be released to the public. The demos are available at
https://github.com/MovieDubbing/Authentic-Dubber.
MODEL ARCHITECTURE
Figure: The proposed Authentic-Dubber consists of Multimodal Reference Footage Construction, Emotion-Similarity-based Retrieval-Augmentation, and Progressive Graph-based Speech Generation.
(* means that the node’s initial vector representation is initialized from the immediately preceding graph.)
Comparison with SOTA Dubbing Methods
Baselines:
1) FastSpeech2* (ICLR 2021): A neural text-to-speech model that incorporates video embeddings as additional input.
2) V2C-Net (CVPR 2022): Leverages a pre-trained emotion model I3D to extract and model emotional information from visual signals.
3) HPMDubbing (CVPR 2023): Performs hierarchical speech modeling by jointly utilizing a facial emotion model EmoFan (capture valence and arousal from facial expressions) and the emotion I3D model (extract scene emotional information).
4) StyleDubber (ACL 2024): Employs an Emotion Reference Transformer to model facial emotional changes at the phoneme level.
5) Speaker2Dubber (MM 2024): Utilizes cross-modal attention to bridge facial emotional features with the prosodic attributes of each phoneme.
Note: To ensure a fair comparison, our model and all baselines adopt FastSpeech2 as the speech synthesis backbone and are trained solely on the V2C dataset, which contains background noise. This choice somewhat limits the acoustic quality of the synthesized speech but does not affect the effectiveness of our model in enhancing emotional expressiveness in dubbing.
| Sample 1 (Full Video) | ||||||
| Ground-Truth | FastSpeech2* | V2C-Net | HPMDubbing | StyleDubber | Speaker2Dubber | Authentic-Dubber | Script : I am so sorry. |
|---|---|---|---|---|---|---|
| Sample2 (Full Video) | ||||||
| Ground-Truth | FastSpeech2* | V2C-Net | HPMDubbing | StyleDubber | Speaker2Dubber | Authentic-Dubber |
|---|---|---|---|---|---|---|
| Script: Great! now you can go! | ||||||
| Sample 3 (Full Video) | ||||||
| Ground-Truth | FastSpeech2* | V2C-Net | HPMDubbing | StyleDubber | Speaker2Dubber | Authentic-Dubber |
|---|---|---|---|---|---|---|
| Script: Excuse me. | ||||||
| Sample4 (Full Video) | ||||||
| Ground-Truth | FastSpeech2* | V2C-Net | HPMDubbing | StyleDubber | Speaker2Dubber | Authentic-Dubber |
|---|---|---|---|---|---|---|
| Script: I'm so sorry i broke it, ralph. | ||||||
Retrieval Result
| Utterance #1 to be Dubbed | Top 1 | Top 2 | Top 9 | Top 10 |
|---|---|---|---|---|
| Scene Retrieval | ||||
| Scene Caption: The scene conveys a warm and friendly atmosphere, with the girl's smile and the playful interaction between her and Olaf. The red and white checkered blanket adds to the cozy and inviting ambiance of the scene. | Scene Caption: The scene conveys a playful and whimsical atmosphere, with the boy's spiky hair and glasses adding to his quirky personality. The blue and white striped wallpaper and the wooden chair in the background create a cozy and nostalgic ambiance. The overall emotional tone of the scene is light-hearted and fun, with a touch of nostalgia. | Scene Caption: The scene conveys a warm and inviting atmosphere, with the girl's smile and the cozy room setting. The green and blue hues of the artwork add a touch of playfulness and creativity to the scene. Overall, it feels like a heartwarming moment between the boy and the girl, as he admires her artwork. | Scene Caption: The scene conveys a light-hearted and playful atmosphere. The characters are engaged in a friendly interaction, with the boy holding a toy and the girl looking at something in her hand. The overall color palette is muted, which adds to the whimsical and dreamlike quality of the scene. | Scene Caption: The scene conveys a sense of nostalgia and whimsy, with the girl's white dress and the colorful drawings on the wall. The light green color of the wall adds to the overall calm and peaceful atmosphere. The presence of the small figurines on the table also adds a playful touch to the scene. |
| Utterance #1 to be Dubbed | Top 1 | Top 2 | Top 9 | Top 10 |
| Face Retrieval | ||||
| Face Caption: The girl's expression starts off calm, then she smiles and appears very happy. As she looks at the snowman, her smile widens, becoming even more joyful. | Face Caption: The character's facial expression starts with a smile, then becomes more animated and excited as she continues to speak. Her eyes widen and her mouth opens slightly as she makes a funny face. Finally, she has a big smile on her face and looks happy and satisfied with herself. | Face Caption: The character starts off with a smile, then becomes more excited as he points at the crowd. He continues to be happy and enthusiastic throughout the video. | Face Caption: The character's facial expression starts with a smile, then becomes more intense as she laughs louder. Her mouth opens wider and her eyes squint in delight. Finally, she ends the conversation with a big smile and a contented look on her face. | Face Caption: The character starts off with a happy expression, then becomes more animated and excited. Finally, he ends the conversation with a smile on his face. |
| Utterance #1 to be Dubbed | Top 1 | Top 2 | Top 9 | Top 10 |
| Text Retrieval | ||||
|
Target Script : I don't worry because. Reaction Caption : I'm fine. |
Target Script : I wouldn't worry about it. Reaction Caption : I'm fine. |
Target Script : Don't worry, tim. Reaction Caption : I'm fine. |
Target Script : Don't worry, kevin. i'll save. Reaction Caption : I'll be fine. |
Target Script : I don't want to put anyone at risk again. Reaction Caption : I'm not going to do it again. |
| Utterance #2 to be Dubbed | Top 1 | Top 2 | Top 9 | Top 10 |
| Scene Retrieval | ||||
| Scene Caption: The scene conveys a somber and melancholic atmosphere. The girl's expression is sad, and the overall color palette of the scene is muted and subdued, further emphasizing the emotional weight of the moment. | Scene Caption: The scene conveys a somber and melancholic atmosphere. The girl's expression is one of sadness, and the overall color palette of the scene is muted and dark, further emphasizing the emotional tone. | Scene Caption: The scene conveys a somber and melancholic atmosphere, with the woman's tears and sobs indicating sadness or distress. The muted colors and dim lighting further emphasize the emotional weight of the moment. | Scene Caption: The scene conveys a somber and melancholic atmosphere. The girl's tears falling onto the mirror reflect her sadness, while the dark and muted colors further emphasize the emotional weight of the moment. | Scene Caption: The scene conveys a somber and melancholic atmosphere. The girl's face is hidden behind her hand, suggesting a sense of sadness or introspection. The dark and muted colors further emphasize the emotional weight of the moment. |
| Utterance #2 to be Dubbed | Top 1 | Top 2 | Top 9 | Top 10 |
| Face Retrieval | ||||
| Face Caption: The character's facial expression starts with a neutral look, then changes to a sad expression as she looks down. The expression remains the same for the duration of the video. | Face Caption: The character's facial expression starts with a neutral look, then shifts to a sad expression as she looks at the ocean. The expression remains the same for the duration of the video. | Face Caption: The character's facial expression starts with a neutral look, then changes to a sad expression as she looks at the other person. The character's eyes widen and their mouth opens slightly in sadness. Finally, the character's expression remains sad as they continue to look at the other person. | Face Caption: The character's facial expression is a mix of sadness and resignation. Her eyes are closed, and her lips are slightly parted as if she is holding back tears. The expression on her face remains consistent throughout the video. | Face Caption: The character's facial expression starts with a neutral look, then changes to a sad expression as she begins to cry. Her eyes are closed and her mouth is slightly open as she continues to cry. The video ends with the character still crying and looking down. |
| Utterance #2 to be Dubbed | Top 1 | Top 2 | Top 9 | Top 10 |
| Text Retrieval | ||||
|
Target Script : I'm sorry, blaze. it's not your fault.. Reaction Caption : I'm sorry. |
Target Script : I'm sorry, bo. Reaction Caption : I'm sorry too. |
Target Script : I'm sorry, dad. Reaction Caption : I'm sorry too. |
Target Script : I'm so sorry, hiccup. Reaction Caption : I'm sorry too. |
Target Script : I'm so sorry i broke it, ralph. Reaction Caption : I'm sorry too. |
| Utterance #3 to be Dubbed | Top 1 | Top 2 | Top 9 | Top 10 |
| Scene Retrieval | ||||
| Scene Caption: The scene conveys a tense and serious atmosphere, with the characters' expressions and body language suggesting that they are engaged in a high-stakes situation. The muted color palette further contribute to this sense of tension and urgency. | Scene Caption: The scene conveys a tense and dramatic atmosphere. The characters' expressions, body language, and the overall setting suggest that something significant is about to happen or has just happened. The dark background and the characters' intense gazes further emphasize the gravity of the situation. | Scene Caption: The scene conveys a tense and serious atmosphere, with the cowboy character looking at the horse with a concerned expression. The dark lighting and muted colors further emphasize the gravity of the situation. | Scene Caption: The scene conveys a tense and serious atmosphere. The characters are engaged in a conversation, with the woman wearing a black mask and the man wearing a black shirt. The lighting is dim, which adds to the intensity of the scene. | Scene Caption: The scene conveys a tense and serious atmosphere, with the characters' expressions and body language suggesting that they are engaged in a high-stakes situation. The muted color palette further contribute to this sense of tension and urgency. |
| Utterance #3 to be Dubbed | Top 1 | Top 2 | Top 9 | Top 10 |
| Face Retrieval | ||||
| Face Caption: The character starts off with a neutral expression, then his expression changes to a frown as he speaks. He then looks very angry with his eyes glaring. | Face Caption: The character starts off with a neutral expression, then becomes angry. He then continues to speak with an angry expression on his face. | Face Caption: The character starts off with a neutral expression, then becomes angry and shouts. He looks very angry with his eyes glaring. | Face Caption: The character's facial expression changes from neutral to angry as he speaks on the phone. His eyebrows raise, his eyes narrow, and his lips purse in a tight line. He maintains this angry expression throughout the conversation. | Face Caption: The character starts off with a neutral expression, then becomes angry as he speaks. His mouth opens wider and his eyes narrow during this transition. Finally, he finishes speaking with an intense look on his face. |
| Utterance #3 to be Dubbed | Top 1 | Top 2 | Top 9 | Top 10 |
| Text Retrieval | ||||
|
Target Script : She must be stopped! you have to go after her. Reaction Caption : X gets angry. |
Target Script : Stop them! Reaction Caption : X gets angry. |
Target Script : we must attack the dragon riders' nest at once! Reaction Caption : X gets angry. |
Target Script : Find that glitch! destroy that cart! Reaction Caption : X gets angry. |
Target Script : Do not give them that chance. Reaction Caption : X gets angry. |