Fine-grained Video Dubbing Duration Alignment with Segment Supervised Preference Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cui, Chaoqun, Huang, Liangbin, Wang, Shijing, Tong, Zhe, Huang, Zhaolong, Zeng, Xiao, Liu, Xiaofeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911102317625344
author Cui, Chaoqun
Huang, Liangbin
Wang, Shijing
Tong, Zhe
Huang, Zhaolong
Zeng, Xiao
Liu, Xiaofeng
author_facet Cui, Chaoqun
Huang, Liangbin
Wang, Shijing
Tong, Zhe
Huang, Zhaolong
Zeng, Xiao
Liu, Xiaofeng
contents Video dubbing aims to translate original speech in visual media programs from the source language to the target language, relying on neural machine translation and text-to-speech technologies. Due to varying information densities across languages, target speech often mismatches the source speech duration, causing audio-video synchronization issues that significantly impact viewer experience. In this study, we approach duration alignment in LLM-based video dubbing machine translation as a preference optimization problem. We propose the Segment Supervised Preference Optimization (SSPO) method, which employs a segment-wise sampling strategy and fine-grained loss to mitigate duration mismatches between source and target lines. Experimental results demonstrate that SSPO achieves superior performance in duration alignment tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2508_08550
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Fine-grained Video Dubbing Duration Alignment with Segment Supervised Preference Optimization
Cui, Chaoqun
Huang, Liangbin
Wang, Shijing
Tong, Zhe
Huang, Zhaolong
Zeng, Xiao
Liu, Xiaofeng
Sound
Computation and Language
Video dubbing aims to translate original speech in visual media programs from the source language to the target language, relying on neural machine translation and text-to-speech technologies. Due to varying information densities across languages, target speech often mismatches the source speech duration, causing audio-video synchronization issues that significantly impact viewer experience. In this study, we approach duration alignment in LLM-based video dubbing machine translation as a preference optimization problem. We propose the Segment Supervised Preference Optimization (SSPO) method, which employs a segment-wise sampling strategy and fine-grained loss to mitigate duration mismatches between source and target lines. Experimental results demonstrate that SSPO achieves superior performance in duration alignment tasks.
title Fine-grained Video Dubbing Duration Alignment with Segment Supervised Preference Optimization
topic Sound
Computation and Language
url https://arxiv.org/abs/2508.08550