VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhang, Xiangdong, Liao, Jiaqi, Zhang, Shaofeng, Meng, Fanqing, Wan, Xiangpeng, Yan, Junchi, Cheng, Yu |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
PCP-MAE: Learning to Predict Centers for Point Masked Autoencoders
par: Zhang, Xiangdong, et autres
Publié: (2024)
par: Zhang, Xiangdong, et autres
Publié: (2024)
Towards More Diverse and Challenging Pre-training for Point Cloud Learning: Self-Supervised Cross Reconstruction with Decoupled Views
par: Zhang, Xiangdong, et autres
Publié: (2025)
par: Zhang, Xiangdong, et autres
Publié: (2025)
Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation
par: Meng, Fanqing, et autres
Publié: (2024)
par: Meng, Fanqing, et autres
Publié: (2024)
Dual-Branch Center-Surrounding Contrast: Rethinking Contrastive Learning for 3D Point Clouds
par: Zhang, Shaofeng, et autres
Publié: (2025)
par: Zhang, Shaofeng, et autres
Publié: (2025)
DreamWorld: Unified World Modeling in Video Generation
par: Tan, Boming, et autres
Publié: (2026)
par: Tan, Boming, et autres
Publié: (2026)
Reasoning via Video: The First Evaluation of Video Models' Reasoning Abilities through Maze-Solving Tasks
par: Yang, Cheng, et autres
Publié: (2025)
par: Yang, Cheng, et autres
Publié: (2025)
On the Evaluation and Refinement of Vision-Language Instruction Tuning Datasets
par: Liao, Ning, et autres
Publié: (2023)
par: Liao, Ning, et autres
Publié: (2023)
REPA Works Until It Doesn't: Early-Stopped, Holistic Alignment Supercharges Diffusion Training
par: Wang, Ziqiao, et autres
Publié: (2025)
par: Wang, Ziqiao, et autres
Publié: (2025)
EvoTok: A Unified Image Tokenizer via Residual Latent Evolution for Visual Understanding and Generation
par: Li, Yan, et autres
Publié: (2026)
par: Li, Yan, et autres
Publié: (2026)
Exploiting Unlabeled Data with Multiple Expert Teachers for Open Vocabulary Aerial Object Detection and Its Orientation Adaptation
par: Li, Yan, et autres
Publié: (2024)
par: Li, Yan, et autres
Publié: (2024)
Motion Control for Enhanced Complex Action Video Generation
par: Zhou, Qiang, et autres
Publié: (2024)
par: Zhou, Qiang, et autres
Publié: (2024)
Align and Aggregate: Compositional Reasoning with Video Alignment and Answer Aggregation for Video Question-Answering
par: Liao, Zhaohe, et autres
Publié: (2024)
par: Liao, Zhaohe, et autres
Publié: (2024)
VideoGrain: Modulating Space-Time Attention for Multi-grained Video Editing
par: Yang, Xiangpeng, et autres
Publié: (2025)
par: Yang, Xiangpeng, et autres
Publié: (2025)
VideoCoF: Unified Video Editing with Temporal Reasoner
par: Yang, Xiangpeng, et autres
Publié: (2025)
par: Yang, Xiangpeng, et autres
Publié: (2025)
Unified Batch Normalization: Identifying and Alleviating the Feature Condensation in Batch Normalization and a Unified Framework
par: Wang, Shaobo, et autres
Publié: (2023)
par: Wang, Shaobo, et autres
Publié: (2023)
RISE-Video: Can Video Generators Decode Implicit World Rules?
par: Liu, Mingxin, et autres
Publié: (2026)
par: Liu, Mingxin, et autres
Publié: (2026)
MMPhysVideo: Scaling Physical Plausibility in Video Generation via Joint Multimodal Modeling
par: Lin, Shubo, et autres
Publié: (2026)
par: Lin, Shubo, et autres
Publié: (2026)
StreamGVE: Training-Free Video Editing via Few-Step Streaming Video Generation
par: Jiao, Guanlong, et autres
Publié: (2026)
par: Jiao, Guanlong, et autres
Publié: (2026)
Evolution of Video Generative Foundations
par: Hu, Teng, et autres
Publié: (2026)
par: Hu, Teng, et autres
Publié: (2026)
PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation
par: Liu, Zhuoman, et autres
Publié: (2024)
par: Liu, Zhuoman, et autres
Publié: (2024)
Identity-GRPO: Optimizing Multi-Human Identity-preserving Video Generation via Reinforcement Learning
par: Meng, Xiangyu, et autres
Publié: (2025)
par: Meng, Xiangyu, et autres
Publié: (2025)
PhysAlign: Physics-Coherent Image-to-Video Generation through Feature and 3D Representation Alignment
par: Xiong, Zhexiao, et autres
Publié: (2026)
par: Xiong, Zhexiao, et autres
Publié: (2026)
U-REPA: Aligning Diffusion U-Nets to ViTs
par: Tian, Yuchuan, et autres
Publié: (2025)
par: Tian, Yuchuan, et autres
Publié: (2025)
Tora3: Trajectory-Guided Audio-Video Generation with Physical Coherence
par: Liao, Junchao, et autres
Publié: (2026)
par: Liao, Junchao, et autres
Publié: (2026)
Video-as-Answer: Predict and Generate Next Video Event with Joint-GRPO
par: Cheng, Junhao, et autres
Publié: (2025)
par: Cheng, Junhao, et autres
Publié: (2025)
Learning Adaptive and Temporally Causal Video Tokenization in a 1D Latent Space
par: Li, Yan, et autres
Publié: (2025)
par: Li, Yan, et autres
Publié: (2025)
InternVideo-Next: Towards General Video Foundation Models without Video-Text Supervision
par: Wang, Chenting, et autres
Publié: (2025)
par: Wang, Chenting, et autres
Publié: (2025)
SWIFT: Prompt-Adaptive Memory for Efficient Interactive Long Video Generation
par: Tan, Shanwen, et autres
Publié: (2026)
par: Tan, Shanwen, et autres
Publié: (2026)
Direct-a-Video: Customized Video Generation with User-Directed Camera Movement and Object Motion
par: Yang, Shiyuan, et autres
Publié: (2024)
par: Yang, Shiyuan, et autres
Publié: (2024)
Pathwise Test-Time Correction for Autoregressive Long Video Generation
par: Xiang, Xunzhi, et autres
Publié: (2026)
par: Xiang, Xunzhi, et autres
Publié: (2026)
Repulsor: Accelerating Generative Modeling with a Contrastive Memory Bank
par: Zhang, Shaofeng, et autres
Publié: (2025)
par: Zhang, Shaofeng, et autres
Publié: (2025)
FullDiT: Multi-Task Video Generative Foundation Model with Full Attention
par: Ju, Xuan, et autres
Publié: (2025)
par: Ju, Xuan, et autres
Publié: (2025)
UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video Generation
par: Huang, Jiehui, et autres
Publié: (2025)
par: Huang, Jiehui, et autres
Publié: (2025)
SwiftVideo: A Unified Framework for Few-Step Video Generation through Trajectory-Distribution Alignment
par: Sun, Yanxiao, et autres
Publié: (2025)
par: Sun, Yanxiao, et autres
Publié: (2025)
EVA: Zero-shot Accurate Attributes and Multi-Object Video Editing
par: Yang, Xiangpeng, et autres
Publié: (2024)
par: Yang, Xiangpeng, et autres
Publié: (2024)
DGL: Dynamic Global-Local Prompt Tuning for Text-Video Retrieval
par: Yang, Xiangpeng, et autres
Publié: (2024)
par: Yang, Xiangpeng, et autres
Publié: (2024)
Tempered Self-Similarity Alignment for Physically Plausible Video Generation
par: Kim, Manjin, et autres
Publié: (2026)
par: Kim, Manjin, et autres
Publié: (2026)
Inference-time Physics Alignment of Video Generative Models with Latent World Models
par: Yuan, Jianhao, et autres
Publié: (2026)
par: Yuan, Jianhao, et autres
Publié: (2026)
Unified Camera Positional Encoding for Controlled Video Generation
par: Zhang, Cheng, et autres
Publié: (2025)
par: Zhang, Cheng, et autres
Publié: (2025)
Content Adaptive based Motion Alignment Framework for Learned Video Compression
par: Zhang, Tiange, et autres
Publié: (2025)
par: Zhang, Tiange, et autres
Publié: (2025)
Documents similaires
-
PCP-MAE: Learning to Predict Centers for Point Masked Autoencoders
par: Zhang, Xiangdong, et autres
Publié: (2024) -
Towards More Diverse and Challenging Pre-training for Point Cloud Learning: Self-Supervised Cross Reconstruction with Decoupled Views
par: Zhang, Xiangdong, et autres
Publié: (2025) -
Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation
par: Meng, Fanqing, et autres
Publié: (2024) -
Dual-Branch Center-Surrounding Contrast: Rethinking Contrastive Learning for 3D Point Clouds
par: Zhang, Shaofeng, et autres
Publié: (2025) -
DreamWorld: Unified World Modeling in Video Generation
par: Tan, Boming, et autres
Publié: (2026)