Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Harold Haodong, Huang, Haojian, Chen, Qifeng, Yang, Harry, Lim, Ser-Nam |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Generation: Unlocking Universal Editing via Self-Supervised Fine-Tuning
by: Chen, Harold Haodong, et al.
Published: (2024)
by: Chen, Harold Haodong, et al.
Published: (2024)
Temporal Regularization Makes Your Video Generator Stronger
by: Chen, Harold Haodong, et al.
Published: (2025)
by: Chen, Harold Haodong, et al.
Published: (2025)
VideoMerge: Towards Training-free Long Video Generation
by: Zhang, Siyang, et al.
Published: (2025)
by: Zhang, Siyang, et al.
Published: (2025)
LightGen: Efficient Image Generation through Knowledge Distillation and Direct Preference Optimization
by: Wu, Xianfeng, et al.
Published: (2025)
by: Wu, Xianfeng, et al.
Published: (2025)
VideoGen-of-Thought: Step-by-step generating multi-shot video with minimal manual intervention
by: Zheng, Mingzhe, et al.
Published: (2025)
by: Zheng, Mingzhe, et al.
Published: (2025)
Zero-shot Synthetic Video Realism Enhancement via Structure-aware Denoising
by: Wang, Yifan, et al.
Published: (2025)
by: Wang, Yifan, et al.
Published: (2025)
VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models
by: Huang, Haojian, et al.
Published: (2025)
by: Huang, Haojian, et al.
Published: (2025)
Towards Chunk-Wise Generation for Long Videos
by: Zhang, Siyang, et al.
Published: (2024)
by: Zhang, Siyang, et al.
Published: (2024)
FineCLIPER: Multi-modal Fine-grained CLIP for Dynamic Facial Expression Recognition with AdaptERs
by: Chen, Haodong, et al.
Published: (2024)
by: Chen, Haodong, et al.
Published: (2024)
FineQuest: Adaptive Knowledge-Assisted Sports Video Understanding via Agent-of-Thoughts Reasoning
by: Chen, Haodong, et al.
Published: (2025)
by: Chen, Haodong, et al.
Published: (2025)
VideoGen-of-Thought: Step-by-step generating multi-shot video with minimal manual intervention
by: Zheng, Mingzhe, et al.
Published: (2024)
by: Zheng, Mingzhe, et al.
Published: (2024)
AC-Foley: Reference-Audio-Guided Video-to-Audio Synthesis with Acoustic Transfer
by: Fang, Pengjun, et al.
Published: (2026)
by: Fang, Pengjun, et al.
Published: (2026)
Fast Encoding and Decoding for Implicit Video Representation
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
Scene Co-pilot: Procedural Text to Video Generation with Human in the Loop
by: Qian, Zhaofang, et al.
Published: (2024)
by: Qian, Zhaofang, et al.
Published: (2024)
Enhancing Physical Plausibility in Video Generation by Reasoning the Implausibility
by: Hao, Yutong, et al.
Published: (2025)
by: Hao, Yutong, et al.
Published: (2025)
Video Decomposition Prior: A Methodology to Decompose Videos into Layers
by: Shrivastava, Gaurav, et al.
Published: (2024)
by: Shrivastava, Gaurav, et al.
Published: (2024)
DiReCT: Disentangled Regularization of Contrastive Trajectories for Physics-Refined Video Generation
by: Meyarian, Abolfazl, et al.
Published: (2026)
by: Meyarian, Abolfazl, et al.
Published: (2026)
CML-Bench: A Framework for Evaluating and Enhancing LLM-Powered Movie Scripts Generation
by: Zheng, Mingzhe, et al.
Published: (2025)
by: Zheng, Mingzhe, et al.
Published: (2025)
AlignVid: Training-Free Attention Scaling for Semantic Fidelity in Text-Guided Image-to-Video Generation
by: Liu, Yexin, et al.
Published: (2025)
by: Liu, Yexin, et al.
Published: (2025)
FinePhys: Fine-grained Human Action Generation by Explicitly Incorporating Physical Laws for Effective Skeletal Guidance
by: Shao, Dian, et al.
Published: (2025)
by: Shao, Dian, et al.
Published: (2025)
FSViewFusion: Few-Shots View Generation of Novel Objects
by: Hussain, Rukhshanda, et al.
Published: (2024)
by: Hussain, Rukhshanda, et al.
Published: (2024)
GaussianVTON: 3D Human Virtual Try-ON via Multi-Stage Gaussian Splatting Editing with Image Prompting
by: Chen, Haodong, et al.
Published: (2024)
by: Chen, Haodong, et al.
Published: (2024)
Model Reveals What to Cache: Profiling-Based Feature Reuse for Video Diffusion Models
by: Ma, Xuran, et al.
Published: (2025)
by: Ma, Xuran, et al.
Published: (2025)
OrthoPhys: Physically Plausible Video Generation with Orthogonal-View Geometry Guidance
by: Wang, Cong, et al.
Published: (2026)
by: Wang, Cong, et al.
Published: (2026)
DreamMask: Boosting Open-vocabulary Panoptic Segmentation with Synthetic Data
by: Tu, Yuanpeng, et al.
Published: (2025)
by: Tu, Yuanpeng, et al.
Published: (2025)
Chain of Event-Centric Causal Thought for Physically Plausible Video Generation
by: Wang, Zixuan, et al.
Published: (2026)
by: Wang, Zixuan, et al.
Published: (2026)
DetailCLIP: Detail-Oriented CLIP for Fine-Grained Tasks
by: Monsefi, Amin Karimi, et al.
Published: (2024)
by: Monsefi, Amin Karimi, et al.
Published: (2024)
Niagara: Normal-Integrated Geometric Affine Field for Scene Reconstruction from a Single View
by: Wu, Xianzu, et al.
Published: (2025)
by: Wu, Xianzu, et al.
Published: (2025)
Find, Fix, Reason: Context Repair for Video Reasoning
by: Huang, Haojian, et al.
Published: (2026)
by: Huang, Haojian, et al.
Published: (2026)
Tempered Self-Similarity Alignment for Physically Plausible Video Generation
by: Kim, Manjin, et al.
Published: (2026)
by: Kim, Manjin, et al.
Published: (2026)
Language-free Compositional Action Generation via Decoupling Refinement
by: Liu, Xiao, et al.
Published: (2023)
by: Liu, Xiao, et al.
Published: (2023)
Unleashing Video Language Models for Fine-grained HRCT Report Generation
by: Fang, Yingying, et al.
Published: (2026)
by: Fang, Yingying, et al.
Published: (2026)
Go with Your Gut: Scaling Confidence for Autoregressive Image Generation
by: Chen, Harold Haodong, et al.
Published: (2025)
by: Chen, Harold Haodong, et al.
Published: (2025)
MMPhysVideo: Scaling Physical Plausibility in Video Generation via Joint Multimodal Modeling
by: Lin, Shubo, et al.
Published: (2026)
by: Lin, Shubo, et al.
Published: (2026)
Structure-Aware Prototype Guided Trusted Multi-View Classification
by: Huang, Haojian, et al.
Published: (2025)
by: Huang, Haojian, et al.
Published: (2025)
VideoDPO: Omni-Preference Alignment for Video Diffusion Generation
by: Liu, Runtao, et al.
Published: (2024)
by: Liu, Runtao, et al.
Published: (2024)
From Generated Human Videos to Physically Plausible Robot Trajectories
by: Ni, James, et al.
Published: (2025)
by: Ni, James, et al.
Published: (2025)
Towards Unified 3D Object Detection via Algorithm and Data Unification
by: Li, Zhuoling, et al.
Published: (2024)
by: Li, Zhuoling, et al.
Published: (2024)
Composing Object Relations and Attributes for Image-Text Matching
by: Pham, Khoi, et al.
Published: (2024)
by: Pham, Khoi, et al.
Published: (2024)
Mitigating Dialogue Hallucination for Large Vision Language Models via Adversarial Instruction Tuning
by: Park, Dongmin, et al.
Published: (2024)
by: Park, Dongmin, et al.
Published: (2024)
Similar Items
-
Beyond Generation: Unlocking Universal Editing via Self-Supervised Fine-Tuning
by: Chen, Harold Haodong, et al.
Published: (2024) -
Temporal Regularization Makes Your Video Generator Stronger
by: Chen, Harold Haodong, et al.
Published: (2025) -
VideoMerge: Towards Training-free Long Video Generation
by: Zhang, Siyang, et al.
Published: (2025) -
LightGen: Efficient Image Generation through Knowledge Distillation and Direct Preference Optimization
by: Wu, Xianfeng, et al.
Published: (2025) -
VideoGen-of-Thought: Step-by-step generating multi-shot video with minimal manual intervention
by: Zheng, Mingzhe, et al.
Published: (2025)