Latent Temporal Discrepancy as Motion Prior: A Loss-Weighting Strategy for Dynamic Fidelity in T2V

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Meiqi, Song, Bingze, Lin, Ruimin, Zhu, Chen, Feng, Xiaokun, Wu, Jiahong, Chu, Xiangxiang, Huang, Kaiqi
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908793403604992
author Wu, Meiqi
Song, Bingze
Lin, Ruimin
Zhu, Chen
Feng, Xiaokun
Wu, Jiahong
Chu, Xiangxiang
Huang, Kaiqi
author_facet Wu, Meiqi
Song, Bingze
Lin, Ruimin
Zhu, Chen
Feng, Xiaokun
Wu, Jiahong
Chu, Xiangxiang
Huang, Kaiqi
contents Video generation models have achieved notable progress in static scenarios, yet their performance in motion video generation remains limited, with quality degrading under drastic dynamic changes. This is due to noise disrupting temporal coherence and increasing the difficulty of learning dynamic regions. {Unfortunately, existing diffusion models rely on static loss for all scenarios, constraining their ability to capture complex dynamics.} To address this issue, we introduce Latent Temporal Discrepancy (LTD) as a motion prior to guide loss weighting. LTD measures frame-to-frame variation in the latent space, assigning larger penalties to regions with higher discrepancy while maintaining regular optimization for stable regions. This motion-aware strategy stabilizes training and enables the model to better reconstruct high-frequency dynamics. Extensive experiments on the general benchmark VBench and the motion-focused VMBench show consistent gains, with our method outperforming strong baselines by 3.31% on VBench and 3.58% on VMBench, achieving significant improvements in motion quality.
format Preprint
id arxiv_https___arxiv_org_abs_2601_20504
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Latent Temporal Discrepancy as Motion Prior: A Loss-Weighting Strategy for Dynamic Fidelity in T2V
Wu, Meiqi
Song, Bingze
Lin, Ruimin
Zhu, Chen
Feng, Xiaokun
Wu, Jiahong
Chu, Xiangxiang
Huang, Kaiqi
Computer Vision and Pattern Recognition
Video generation models have achieved notable progress in static scenarios, yet their performance in motion video generation remains limited, with quality degrading under drastic dynamic changes. This is due to noise disrupting temporal coherence and increasing the difficulty of learning dynamic regions. {Unfortunately, existing diffusion models rely on static loss for all scenarios, constraining their ability to capture complex dynamics.} To address this issue, we introduce Latent Temporal Discrepancy (LTD) as a motion prior to guide loss weighting. LTD measures frame-to-frame variation in the latent space, assigning larger penalties to regions with higher discrepancy while maintaining regular optimization for stable regions. This motion-aware strategy stabilizes training and enables the model to better reconstruct high-frequency dynamics. Extensive experiments on the general benchmark VBench and the motion-focused VMBench show consistent gains, with our method outperforming strong baselines by 3.31% on VBench and 3.58% on VMBench, achieving significant improvements in motion quality.
title Latent Temporal Discrepancy as Motion Prior: A Loss-Weighting Strategy for Dynamic Fidelity in T2V
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.20504