Improving the Physics of Video Generation with VJEPA-2 Reward Signal
Fuente:
arXiv
Saved in:
| Main Authors: | Yuan, Jianhao, Zhang, Xiaofeng, Friedrich, Felix, Beltran-Velez, Nicolas, Hall, Melissa, Askari-Hemmat, Reyhane, Han, Xiaochuang, Ballas, Nicolas, Drozdzal, Michal, Romero-Soriano, Adriana |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Inference-time Physics Alignment of Video Generative Models with Latent World Models
by: Yuan, Jianhao, et al.
Published: (2026)
by: Yuan, Jianhao, et al.
Published: (2026)
Improving Geo-diversity of Generated Images with Contextualized Vendi Score Guidance
by: Hemmat, Reyhane Askari, et al.
Published: (2024)
by: Hemmat, Reyhane Askari, et al.
Published: (2024)
Increasing the Utility of Synthetic Images through Chamfer Guidance
by: Dall'Asen, Nicola, et al.
Published: (2025)
by: Dall'Asen, Nicola, et al.
Published: (2025)
Feedback-guided Data Synthesis for Imbalanced Classification
by: Hemmat, Reyhane Askari, et al.
Published: (2023)
by: Hemmat, Reyhane Askari, et al.
Published: (2023)
Multi-Modal Language Models as Text-to-Image Model Evaluators
by: Chen, Jiahui, et al.
Published: (2025)
by: Chen, Jiahui, et al.
Published: (2025)
Improving the Scaling Laws of Synthetic Data with Deliberate Practice
by: Askari-Hemmat, Reyhane, et al.
Published: (2025)
by: Askari-Hemmat, Reyhane, et al.
Published: (2025)
On Improved Conditioning Mechanisms and Pre-training Strategies for Diffusion Models
by: Ifriqi, Tariq Berrada, et al.
Published: (2024)
by: Ifriqi, Tariq Berrada, et al.
Published: (2024)
Multimodal RewardBench 2: Evaluating Omni Reward Models for Interleaved Text and Image
by: Hu, Yushi, et al.
Published: (2025)
by: Hu, Yushi, et al.
Published: (2025)
DIG In: Evaluating Disparities in Image Generations with Indicators for Geographic Diversity
by: Hall, Melissa, et al.
Published: (2023)
by: Hall, Melissa, et al.
Published: (2023)
Unified Text-Image Generation with Weakness-Targeted Post-Training
by: Chen, Jiahui, et al.
Published: (2026)
by: Chen, Jiahui, et al.
Published: (2026)
Why Less is More (Sometimes): A Theory of Data Curation
by: Dohmatob, Elvis, et al.
Published: (2025)
by: Dohmatob, Elvis, et al.
Published: (2025)
Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models
by: Wang, Shuting, et al.
Published: (2025)
by: Wang, Shuting, et al.
Published: (2025)
Non-Euclidean Sliced Optimal Transport Sampling
by: Genest, Baptiste, et al.
Published: (2024)
by: Genest, Baptiste, et al.
Published: (2024)
Implementing a Machine Learning Deformer for CG Crowds: Our Journey
by: Arcelin, Bastien, et al.
Published: (2024)
by: Arcelin, Bastien, et al.
Published: (2024)
NePHIM: A Neural Physics-Based Head-Hand Interaction Model
by: Wagner, Nicolas, et al.
Published: (2024)
by: Wagner, Nicolas, et al.
Published: (2024)
The Intricate Dance of Prompt Complexity, Quality, Diversity, and Consistency in T2I Models
by: Xiaofeng, Zhang, et al.
Published: (2025)
by: Xiaofeng, Zhang, et al.
Published: (2025)
Dr.Jit: A Just-In-Time Compiler for Differentiable Rendering
by: Jakob, Wenzel, et al.
Published: (2022)
by: Jakob, Wenzel, et al.
Published: (2022)
Audio2Rig: Artist-oriented deep learning tool for facial animation
by: Arcelin, Bastien, et al.
Published: (2024)
by: Arcelin, Bastien, et al.
Published: (2024)
Learning Human-like Locomotion Based on Biological Actuation and Rewards
by: Kim, Minkwan, et al.
Published: (2024)
by: Kim, Minkwan, et al.
Published: (2024)
Deformation Recovery: Localized Learning for Detail-Preserving Deformations
by: Sundararaman, Ramana, et al.
Published: (2024)
by: Sundararaman, Ramana, et al.
Published: (2024)
CAD-Coder: Text-to-CAD Generation with Chain-of-Thought and Geometric Reward
by: Guan, Yandong, et al.
Published: (2025)
by: Guan, Yandong, et al.
Published: (2025)
Shap-MeD
by: Laverde, Nicolás, et al.
Published: (2025)
by: Laverde, Nicolás, et al.
Published: (2025)
PosterReward: Unlocking Accurate Evaluation for High-Quality Graphic Design Generation
by: Lai, Jianyu, et al.
Published: (2026)
by: Lai, Jianyu, et al.
Published: (2026)
R3-RECON: Radiance-Field-Free Active Reconstruction via Renderability
by: Jin, Xiaofeng, et al.
Published: (2026)
by: Jin, Xiaofeng, et al.
Published: (2026)
Automatic Inbetweening for Stroke‐Based Painterly Animation
by: Nicolas Barroso, et al.
Published: (2024)
by: Nicolas Barroso, et al.
Published: (2024)
Non‐Euclidean Sliced Optimal Transport Sampling
by: Baptiste Genest, et al.
Published: (2024)
by: Baptiste Genest, et al.
Published: (2024)
VideoFrom3D: 3D Scene Video Generation via Complementary Image and Video Diffusion Models
by: Kim, Geonung, et al.
Published: (2025)
by: Kim, Geonung, et al.
Published: (2025)
TransVDM: Motion-Constrained Video Diffusion Model for Transparent Video Synthesis
by: Li, Menghao, et al.
Published: (2025)
by: Li, Menghao, et al.
Published: (2025)
Nearest Neighbor Classification for Classical Image Upsampling
by: Matthews, Evan, et al.
Published: (2024)
by: Matthews, Evan, et al.
Published: (2024)
SURF: Signature-Retained Fast Video Generation
by: Ding, Kaixin, et al.
Published: (2025)
by: Ding, Kaixin, et al.
Published: (2025)
FreqPrior: Improving Video Diffusion Models with Frequency Filtering Gaussian Noise
by: Yuan, Yunlong, et al.
Published: (2025)
by: Yuan, Yunlong, et al.
Published: (2025)
Holographic Parallax Improves 3D Perceptual Realism
by: Kim, Dongyeon, et al.
Published: (2024)
by: Kim, Dongyeon, et al.
Published: (2024)
T2Bs: Text-to-Character Blendshapes via Video Generation
by: Luo, Jiahao, et al.
Published: (2025)
by: Luo, Jiahao, et al.
Published: (2025)
Artist‐Inator: Text‐based, Gloss‐aware Non‐photorealistic Stylization
by: J. Daniel Subias, et al.
Published: (2025)
by: J. Daniel Subias, et al.
Published: (2025)
Data Visualization for Improving Financial Literacy: A Systematic Review
by: Du, Meng, et al.
Published: (2025)
by: Du, Meng, et al.
Published: (2025)
AnimaMimic: Imitating 3D Animation from Video Priors
by: Xie, Tianyi, et al.
Published: (2025)
by: Xie, Tianyi, et al.
Published: (2025)
HIL: Hybrid Imitation Learning of Diverse Parkour Skills from Videos
by: Wang, Jiashun, et al.
Published: (2025)
by: Wang, Jiashun, et al.
Published: (2025)
Topology-Aware Optimization of Gaussian Primitives for Human-Centric Volumetric Videos
by: Jiang, Yuheng, et al.
Published: (2025)
by: Jiang, Yuheng, et al.
Published: (2025)
Improving Sparse IMU-based Motion Capture with Motion Label Smoothing
by: Meng, Zhaorui, et al.
Published: (2025)
by: Meng, Zhaorui, et al.
Published: (2025)
Visualization of Escher‐like Kaleidoscopic Spherical Patterns of Regular Polyhedron Symmetry
by: Krzysztof Gdawiec, et al.
Published: (2024)
by: Krzysztof Gdawiec, et al.
Published: (2024)
Similar Items
-
Inference-time Physics Alignment of Video Generative Models with Latent World Models
by: Yuan, Jianhao, et al.
Published: (2026) -
Improving Geo-diversity of Generated Images with Contextualized Vendi Score Guidance
by: Hemmat, Reyhane Askari, et al.
Published: (2024) -
Increasing the Utility of Synthetic Images through Chamfer Guidance
by: Dall'Asen, Nicola, et al.
Published: (2025) -
Feedback-guided Data Synthesis for Imbalanced Classification
by: Hemmat, Reyhane Askari, et al.
Published: (2023) -
Multi-Modal Language Models as Text-to-Image Model Evaluators
by: Chen, Jiahui, et al.
Published: (2025)