Video Generation Models Are Good Latent Reward Models
Fuente:
arXiv
Saved in:
| Main Authors: | Mi, Xiaoyue, Yu, Wenqing, Lian, Jiesong, Jie, Shibo, Zhong, Ruizhe, Liu, Zijun, Zhang, Guozhen, Zhou, Zixiang, Xu, Zhiyong, Zhou, Yuan, Lu, Qinglin, Tang, Fan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SoliReward: Mitigating Susceptibility to Reward Hacking and Annotation Noise in Video Generation Reward Models
by: Lian, Jiesong, et al.
Published: (2025)
by: Lian, Jiesong, et al.
Published: (2025)
Euphonium: Steering Video Flow Matching via Process Reward Gradient Guided Stochastic Dynamics
by: Zhong, Ruizhe, et al.
Published: (2026)
by: Zhong, Ruizhe, et al.
Published: (2026)
SARA: Semantically Adaptive Relational Alignment for Video Diffusion Models
by: Lian, Jiesong, et al.
Published: (2026)
by: Lian, Jiesong, et al.
Published: (2026)
Pack and Force Your Memory: Long-form and Consistent Video Generation
by: Wu, Xiaofei, et al.
Published: (2025)
by: Wu, Xiaofei, et al.
Published: (2025)
UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal Interactions
by: Zhang, Guozhen, et al.
Published: (2025)
by: Zhang, Guozhen, et al.
Published: (2025)
Arbitrary Generative Video Interpolation
by: Zhang, Guozhen, et al.
Published: (2025)
by: Zhang, Guozhen, et al.
Published: (2025)
HunyuanVideo-HOMA: Generic Human-Object Interaction in Multimodal Driven Human Animation
by: Huang, Ziyao, et al.
Published: (2025)
by: Huang, Ziyao, et al.
Published: (2025)
USV: Unified Sparsification for Accelerating Video Diffusion Models
by: Wu, Xinjian, et al.
Published: (2025)
by: Wu, Xinjian, et al.
Published: (2025)
Audio-visual Controlled Video Diffusion with Masked Selective State Spaces Modeling for Natural Talking Head Generation
by: Hong, Fa-Ting, et al.
Published: (2025)
by: Hong, Fa-Ting, et al.
Published: (2025)
Harmony: Harmonizing Audio and Video Generation through Cross-Task Synergy
by: Hu, Teng, et al.
Published: (2025)
by: Hu, Teng, et al.
Published: (2025)
StreamAvatar: Streaming Diffusion Models for Real-Time Interactive Human Avatars
by: Sun, Zhiyao, et al.
Published: (2025)
by: Sun, Zhiyao, et al.
Published: (2025)
HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters
by: Chen, Yi, et al.
Published: (2025)
by: Chen, Yi, et al.
Published: (2025)
Hunyuan-GameCraft: High-dynamic Interactive Game Video Generation with Hybrid History Condition
by: Li, Jiaqi, et al.
Published: (2025)
by: Li, Jiaqi, et al.
Published: (2025)
ActAvatar: Temporally-Aware Precise Action Control for Talking Avatars
by: Peng, Ziqiao, et al.
Published: (2025)
by: Peng, Ziqiao, et al.
Published: (2025)
HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation
by: Hu, Teng, et al.
Published: (2025)
by: Hu, Teng, et al.
Published: (2025)
Latent-Reframe: Enabling Camera Control for Video Diffusion Model without Training
by: Zhou, Zhenghong, et al.
Published: (2024)
by: Zhou, Zhenghong, et al.
Published: (2024)
PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement
by: Hu, Teng, et al.
Published: (2025)
by: Hu, Teng, et al.
Published: (2025)
DFSAttn: Dynamic Fine-grained Sparse Attention for Efficient Video Generation
by: Hu, Jie, et al.
Published: (2026)
by: Hu, Jie, et al.
Published: (2026)
On uniformly quasiconformal Anosov diffeomorphisms with two dimensional distributions
by: Zhang, Jiesong
Published: (2023)
by: Zhang, Jiesong
Published: (2023)
Multi-modal Segment Assemblage Network for Ad Video Editing with Importance-Coherence Reward
by: Tang, Yolo Yunlong, et al.
Published: (2022)
by: Tang, Yolo Yunlong, et al.
Published: (2022)
Development of numerical methods for nonlinear hybrid stochastic functional differential equations with infinite delay
by: Li, Guozhen, et al.
Published: (2025)
by: Li, Guozhen, et al.
Published: (2025)
CaC: Advancing Video Reward Models via Hierarchical Spatiotemporal Concentrating
by: Wang, Jiyuan, et al.
Published: (2026)
by: Wang, Jiyuan, et al.
Published: (2026)
HunyuanVideo: A Systematic Framework For Large Video Generative Models
by: Kong, Weijie, et al.
Published: (2024)
by: Kong, Weijie, et al.
Published: (2024)
Interactive Visual Assessment for Text-to-Image Generation Models
by: Mi, Xiaoyue, et al.
Published: (2024)
by: Mi, Xiaoyue, et al.
Published: (2024)
Making Avatars Interact: Towards Text-Driven Human-Object Interaction for Controllable Talking Avatars
by: Zhang, Youliang, et al.
Published: (2026)
by: Zhang, Youliang, et al.
Published: (2026)
Enhancing Spatial Understanding in Image Generation via Reward Modeling
by: Tang, Zhenyu, et al.
Published: (2026)
by: Tang, Zhenyu, et al.
Published: (2026)
Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models
by: Wei, Yuancheng, et al.
Published: (2026)
by: Wei, Yuancheng, et al.
Published: (2026)
Free-GVC: Towards Training-Free Extreme Generative Video Compression with Temporal Coherence
by: Ling, Xiaoyue, et al.
Published: (2026)
by: Ling, Xiaoyue, et al.
Published: (2026)
GAS: Enhancing Reward-Cost Balance of Generative Model-assisted Offline Safe RL
by: Liu, Zifan, et al.
Published: (2026)
by: Liu, Zifan, et al.
Published: (2026)
Beyond VLM-Based Rewards: Diffusion-Native Latent Reward Modeling
by: Liu, Gongye, et al.
Published: (2026)
by: Liu, Gongye, et al.
Published: (2026)
MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?
by: Chen, Zhaorun, et al.
Published: (2024)
by: Chen, Zhaorun, et al.
Published: (2024)
Bora: Biomedical Generalist Video Generation Model
by: Sun, Weixiang, et al.
Published: (2024)
by: Sun, Weixiang, et al.
Published: (2024)
Thinking with Frames: Generative Video Distortion Evaluation via Frame Reward Model
by: Wang, Yuan, et al.
Published: (2026)
by: Wang, Yuan, et al.
Published: (2026)
Bounded cohomology of diffeomorphism groups of higher dimensional spheres
by: Zhou, Zixiang
Published: (2024)
by: Zhou, Zixiang
Published: (2024)
Towards Intrinsic Interpretability of Large Language Models:A Survey of Design Principles and Architectures
by: Gao, Yutong, et al.
Published: (2026)
by: Gao, Yutong, et al.
Published: (2026)
OD-VAE: An Omni-dimensional Video Compressor for Improving Latent Video Diffusion Model
by: Chen, Liuhan, et al.
Published: (2024)
by: Chen, Liuhan, et al.
Published: (2024)
HMARK: Radioactive Multi-Bit Semantic-Latent Watermarking for Diffusion Models
by: Li, Kexin, et al.
Published: (2025)
by: Li, Kexin, et al.
Published: (2025)
AgentRM: Enhancing Agent Generalization with Reward Modeling
by: Xia, Yu, et al.
Published: (2025)
by: Xia, Yu, et al.
Published: (2025)
Owl-1: Omni World Model for Consistent Long Video Generation
by: Huang, Yuanhui, et al.
Published: (2024)
by: Huang, Yuanhui, et al.
Published: (2024)
PRMB: Benchmarking Reward Models in Long-Horizon CBT-based Counseling Dialogue
by: Zhou, Yougen, et al.
Published: (2026)
by: Zhou, Yougen, et al.
Published: (2026)
Similar Items
-
SoliReward: Mitigating Susceptibility to Reward Hacking and Annotation Noise in Video Generation Reward Models
by: Lian, Jiesong, et al.
Published: (2025) -
Euphonium: Steering Video Flow Matching via Process Reward Gradient Guided Stochastic Dynamics
by: Zhong, Ruizhe, et al.
Published: (2026) -
SARA: Semantically Adaptive Relational Alignment for Video Diffusion Models
by: Lian, Jiesong, et al.
Published: (2026) -
Pack and Force Your Memory: Long-form and Consistent Video Generation
by: Wu, Xiaofei, et al.
Published: (2025) -
UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal Interactions
by: Zhang, Guozhen, et al.
Published: (2025)