Latent Temporal Discrepancy as Motion Prior: A Loss-Weighting Strategy for Dynamic Fidelity in T2V
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Meiqi, Song, Bingze, Lin, Ruimin, Zhu, Chen, Feng, Xiaokun, Wu, Jiahong, Chu, Xiangxiang, Huang, Kaiqi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ImagerySearch: Adaptive Test-Time Search for Video Generation Beyond Semantic Dependency Constraints
by: Wu, Meiqi, et al.
Published: (2025)
by: Wu, Meiqi, et al.
Published: (2025)
Artifact-Aware Evaluation for High-Quality Video Generation
by: Zhu, Chen, et al.
Published: (2026)
by: Zhu, Chen, et al.
Published: (2026)
Taming Hallucinations: Boosting MLLMs' Video Understanding via Counterfactual Video Generation
by: Huang, Zhe, et al.
Published: (2025)
by: Huang, Zhe, et al.
Published: (2025)
VMBench: A Benchmark for Perception-Aligned Video Motion Generation
by: Ling, Xinran, et al.
Published: (2025)
by: Ling, Xinran, et al.
Published: (2025)
Omni-WorldBench: Towards a Comprehensive Interaction-Centric Evaluation for World Models
by: Wu, Meiqi, et al.
Published: (2026)
by: Wu, Meiqi, et al.
Published: (2026)
Stochastic Self-Guidance for Training-Free Enhancement of Diffusion Models
by: Chen, Chubin, et al.
Published: (2025)
by: Chen, Chubin, et al.
Published: (2025)
Revealing the Dark Secrets of Extremely Large Kernel ConvNets on Robustness
by: Chen, Honghao, et al.
Published: (2024)
by: Chen, Honghao, et al.
Published: (2024)
Omni-Effects: Unified and Spatially-Controllable Visual Effects Generation
by: Mao, Fangyuan, et al.
Published: (2025)
by: Mao, Fangyuan, et al.
Published: (2025)
How Texts Help? A Fine-grained Evaluation to Reveal the Role of Language in Vision-Language Tracking
by: Li, Xuchen, et al.
Published: (2024)
by: Li, Xuchen, et al.
Published: (2024)
DTLLM-VLT: Diverse Text Generation for Visual Language Tracking Based on LLM
by: Li, Xuchen, et al.
Published: (2024)
by: Li, Xuchen, et al.
Published: (2024)
Visual Language Tracking with Multi-modal Interaction: A Robust Benchmark
by: Li, Xuchen, et al.
Published: (2024)
by: Li, Xuchen, et al.
Published: (2024)
DTVLT: A Multi-modal Diverse Text Benchmark for Visual Language Tracking Based on LLM
by: Li, Xuchen, et al.
Published: (2024)
by: Li, Xuchen, et al.
Published: (2024)
Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation
by: Zhao, Chenxi, et al.
Published: (2026)
by: Zhao, Chenxi, et al.
Published: (2026)
Ranking-aware Reinforcement Learning for Ordinal Ranking
by: Hao, Aiming, et al.
Published: (2026)
by: Hao, Aiming, et al.
Published: (2026)
Taming Preference Mode Collapse via Directional Decoupling Alignment in Diffusion Reinforcement Learning
by: Chen, Chubin, et al.
Published: (2025)
by: Chen, Chubin, et al.
Published: (2025)
Embedding-perturbed Exploration Preference Optimization for Flow Models
by: Hu, Sujie, et al.
Published: (2026)
by: Hu, Sujie, et al.
Published: (2026)
PeLK: Parameter-efficient Large Kernel ConvNets with Peripheral Convolution
by: Chen, Honghao, et al.
Published: (2024)
by: Chen, Honghao, et al.
Published: (2024)
MACE-Dance: Motion-Appearance Cascaded Experts for Music-Driven Dance Video Generation
by: Yang, Kaixing, et al.
Published: (2025)
by: Yang, Kaixing, et al.
Published: (2025)
FlowDreamer: Exploring High Fidelity Text-to-3D Generation via Rectified Flow
by: Li, Hangyu, et al.
Published: (2024)
by: Li, Hangyu, et al.
Published: (2024)
Spherical Latent Motion Prior for Physics-Based Simulated Humanoid Control
by: Tan, Jing, et al.
Published: (2026)
by: Tan, Jing, et al.
Published: (2026)
Unveiling Chain of Step Reasoning for Vision-Language Models with Fine-grained Rewards
by: Chen, Honghao, et al.
Published: (2025)
by: Chen, Honghao, et al.
Published: (2025)
Finger in Camera Speaks Everything: Unconstrained Air-Writing for Real-World
by: Wu, Meiqi, et al.
Published: (2024)
by: Wu, Meiqi, et al.
Published: (2024)
MAR-GRPO: Stabilized GRPO for AR-diffusion Hybrid Image Generation
by: Ma, Xiaoxiao, et al.
Published: (2026)
by: Ma, Xiaoxiao, et al.
Published: (2026)
There is No VAE: End-to-End Pixel-Space Generative Modeling via Self-Supervised Pre-training
by: Lei, Jiachen, et al.
Published: (2025)
by: Lei, Jiachen, et al.
Published: (2025)
An Efficient Watermarking Method for Latent Diffusion Models via Low-Rank Adaptation and Dynamic Loss Weighting
by: Lin, Dongdong, et al.
Published: (2024)
by: Lin, Dongdong, et al.
Published: (2024)
Learning Dynamic Representations via An Optimally-Weighted Maximum Mean Discrepancy Optimization Framework for Continual Learning
by: Huang, KaiHui, et al.
Published: (2025)
by: Huang, KaiHui, et al.
Published: (2025)
A Temporal Modeling Framework for Video Pre-Training on Video Instance Segmentation
by: Zhong, Qing, et al.
Published: (2025)
by: Zhong, Qing, et al.
Published: (2025)
FlowLoss: Dynamic Flow-Conditioned Loss Strategy for Video Diffusion Models
by: Wu, Kuanting, et al.
Published: (2025)
by: Wu, Kuanting, et al.
Published: (2025)
The Structural Origin of Attention Sink: Variance Discrepancy, Super Neurons, and Dimension Disparity
by: Li, Siquan, et al.
Published: (2026)
by: Li, Siquan, et al.
Published: (2026)
Hyperbolic Space Learning Method Leveraging Temporal Motion Priors for Human Mesh Recovery
by: Zhang, Xiang, et al.
Published: (2025)
by: Zhang, Xiang, et al.
Published: (2025)
Temporal Continual Learning with Prior Compensation for Human Motion Prediction
by: Tang, Jianwei, et al.
Published: (2025)
by: Tang, Jianwei, et al.
Published: (2025)
Latent Conditioned Loco-Manipulation Using Motion Priors
by: Stępień, Maciej, et al.
Published: (2025)
by: Stępień, Maciej, et al.
Published: (2025)
Multi-Floor Exploration for Ground Robots via an Incremental Reachable Graph and Structural Priors
by: Zhu, Zhiwen, et al.
Published: (2026)
by: Zhu, Zhiwen, et al.
Published: (2026)
LD-RPS: Zero-Shot Unified Image Restoration via Latent Diffusion Recurrent Posterior Sampling
by: Li, Huaqiu, et al.
Published: (2025)
by: Li, Huaqiu, et al.
Published: (2025)
Extraction and Recovery of Spatio-Temporal Structure in Latent Dynamics Alignment with Diffusion Models
by: Wang, Yule, et al.
Published: (2023)
by: Wang, Yule, et al.
Published: (2023)
Semantic and Temporal Integration in Latent Diffusion Space for High-Fidelity Video Super-Resolution
by: Wang, Yiwen, et al.
Published: (2025)
by: Wang, Yiwen, et al.
Published: (2025)
VS-LLM: Visual-Semantic Depression Assessment based on LLM for Drawing Projection Test
by: Wu, Meiqi, et al.
Published: (2025)
by: Wu, Meiqi, et al.
Published: (2025)
iPay: Integrated Payment Action Recognition via Multimodal Networks and Adaptive Spatial Prior Learning
by: Huang, Kaicong, et al.
Published: (2026)
by: Huang, Kaicong, et al.
Published: (2026)
TeleGate: Whole-Body Humanoid Teleoperation via Gated Expert Selection with Motion Prior
by: Li, Jie, et al.
Published: (2026)
by: Li, Jie, et al.
Published: (2026)
Semantic Ensemble Loss and Latent Refinement for High-Fidelity Neural Image Compression
by: Li, Daxin, et al.
Published: (2024)
by: Li, Daxin, et al.
Published: (2024)
Similar Items
-
ImagerySearch: Adaptive Test-Time Search for Video Generation Beyond Semantic Dependency Constraints
by: Wu, Meiqi, et al.
Published: (2025) -
Artifact-Aware Evaluation for High-Quality Video Generation
by: Zhu, Chen, et al.
Published: (2026) -
Taming Hallucinations: Boosting MLLMs' Video Understanding via Counterfactual Video Generation
by: Huang, Zhe, et al.
Published: (2025) -
VMBench: A Benchmark for Perception-Aligned Video Motion Generation
by: Ling, Xinran, et al.
Published: (2025) -
Omni-WorldBench: Towards a Comprehensive Interaction-Centric Evaluation for World Models
by: Wu, Meiqi, et al.
Published: (2026)