Proprio: Latent Self-Scoring and Inference-Time Refinement for Physically Plausible Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Hassan, Mariam, Messaoud, Kaouther, Li, Wuyang, Alahi, Alexandre |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Anchored Video Generation: Decoupling Scene Construction and Temporal Synthesis in Text-to-Video Diffusion Models
by: Hassan, Mariam, et al.
Published: (2025)
by: Hassan, Mariam, et al.
Published: (2025)
Towards Generalizable Trajectory Prediction Using Dual-Level Representation Learning And Adaptive Prompting
by: Messaoud, Kaouther, et al.
Published: (2025)
by: Messaoud, Kaouther, et al.
Published: (2025)
Social-Transmotion: Promptable Human Trajectory Prediction
by: Saadatnejad, Saeed, et al.
Published: (2023)
by: Saadatnejad, Saeed, et al.
Published: (2023)
OmniTraj: Pre-Training on Heterogeneous Data for Adaptive and Zero-Shot Human Trajectory Prediction
by: Gao, Yang, et al.
Published: (2025)
by: Gao, Yang, et al.
Published: (2025)
Forecast-PEFT: Parameter-Efficient Fine-Tuning for Pre-trained Motion Forecasting Models
by: Wang, Jifeng, et al.
Published: (2024)
by: Wang, Jifeng, et al.
Published: (2024)
EverAnimate: Minute-Scale Human Animation via Latent Flow Restoration
by: Li, Wuyang, et al.
Published: (2026)
by: Li, Wuyang, et al.
Published: (2026)
Stable Video Infinity: Infinite-Length Video Generation with Error Recycling
by: Li, Wuyang, et al.
Published: (2025)
by: Li, Wuyang, et al.
Published: (2025)
UniTraj: A Unified Framework for Scalable Vehicle Trajectory Prediction
by: Feng, Lan, et al.
Published: (2024)
by: Feng, Lan, et al.
Published: (2024)
VoxDet: Rethinking 3D Semantic Occupancy Prediction as Dense Object Detection
by: Li, Wuyang, et al.
Published: (2025)
by: Li, Wuyang, et al.
Published: (2025)
LayerSync: Self-aligning Intermediate Layers
by: Haghighi, Yasaman, et al.
Published: (2025)
by: Haghighi, Yasaman, et al.
Published: (2025)
Social-Mamba: Socially-Aware Trajectory Forecasting with State-Space Models
by: Luan, Po-Chien, et al.
Published: (2026)
by: Luan, Po-Chien, et al.
Published: (2026)
From Generation to Generalization: Emergent Few-Shot Learning in Video Diffusion Models
by: Acuaviva, Pablo, et al.
Published: (2025)
by: Acuaviva, Pablo, et al.
Published: (2025)
Tempered Self-Similarity Alignment for Physically Plausible Video Generation
by: Kim, Manjin, et al.
Published: (2026)
by: Kim, Manjin, et al.
Published: (2026)
PhyPrompt: RL-based Prompt Refinement for Physically Plausible Text-to-Video Generation
by: Wu, Shang, et al.
Published: (2026)
by: Wu, Shang, et al.
Published: (2026)
Drift-Resistant Navigation World Model with Anchored Epipolar Guidance
by: Luan, Po-Chien, et al.
Published: (2026)
by: Luan, Po-Chien, et al.
Published: (2026)
SenCache: Accelerating Diffusion Model Inference via Sensitivity-Aware Caching
by: Haghighi, Yasaman, et al.
Published: (2026)
by: Haghighi, Yasaman, et al.
Published: (2026)
Enhancing Physical Plausibility in Video Generation by Reasoning the Implausibility
by: Hao, Yutong, et al.
Published: (2025)
by: Hao, Yutong, et al.
Published: (2025)
Chain of Event-Centric Causal Thought for Physically Plausible Video Generation
by: Wang, Zixuan, et al.
Published: (2026)
by: Wang, Zixuan, et al.
Published: (2026)
ScoreHOI: Physically Plausible Reconstruction of Human-Object Interaction via Score-Guided Diffusion
by: Li, Ao, et al.
Published: (2025)
by: Li, Ao, et al.
Published: (2025)
FG$^2$: Fine-Grained Cross-View Localization by Fine-Grained Feature Matching
by: Xia, Zimin, et al.
Published: (2025)
by: Xia, Zimin, et al.
Published: (2025)
Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation
by: Chen, Harold Haodong, et al.
Published: (2025)
by: Chen, Harold Haodong, et al.
Published: (2025)
MMPhysVideo: Scaling Physical Plausibility in Video Generation via Joint Multimodal Modeling
by: Lin, Shubo, et al.
Published: (2026)
by: Lin, Shubo, et al.
Published: (2026)
From Generated Human Videos to Physically Plausible Robot Trajectories
by: Ni, James, et al.
Published: (2025)
by: Ni, James, et al.
Published: (2025)
OrthoPhys: Physically Plausible Video Generation with Orthogonal-View Geometry Guidance
by: Wang, Cong, et al.
Published: (2026)
by: Wang, Cong, et al.
Published: (2026)
Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception
by: Simon, Marcel, et al.
Published: (2025)
by: Simon, Marcel, et al.
Published: (2025)
Inference-time Physics Alignment of Video Generative Models with Latent World Models
by: Yuan, Jianhao, et al.
Published: (2026)
by: Yuan, Jianhao, et al.
Published: (2026)
NAT: Learning to Attack Neurons for Enhanced Adversarial Transferability
by: Nakka, Krishna Kanth, et al.
Published: (2025)
by: Nakka, Krishna Kanth, et al.
Published: (2025)
RAP: 3D Rasterization Augmented End-to-End Planning
by: Feng, Lan, et al.
Published: (2025)
by: Feng, Lan, et al.
Published: (2025)
X-GRM: Large Gaussian Reconstruction Model for Sparse-view X-rays to Computed Tomography
by: Liu, Yifan, et al.
Published: (2025)
by: Liu, Yifan, et al.
Published: (2025)
Co-Supervised Learning: Improving Weak-to-Strong Generalization with Hierarchical Mixture of Experts
by: Liu, Yuejiang, et al.
Published: (2024)
by: Liu, Yuejiang, et al.
Published: (2024)
VLIPP: Towards Physically Plausible Video Generation with Vision and Language Informed Physical Prior
by: Yang, Xindi, et al.
Published: (2025)
by: Yang, Xindi, et al.
Published: (2025)
Bootstrapping Physics-Grounded Video Generation through VLM-Guided Iterative Self-Refinement
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
Loc$^2$: Interpretable Cross-View Localization via Depth-Lifted Local Feature Matching
by: Xia, Zimin, et al.
Published: (2025)
by: Xia, Zimin, et al.
Published: (2025)
Social-Pose: Enhancing Trajectory Prediction with Human Body Pose
by: Gao, Yang, et al.
Published: (2025)
by: Gao, Yang, et al.
Published: (2025)
GeoDistill: Geometry-Guided Self-Distillation for Weakly Supervised Cross-View Localization
by: Tong, Shaowen, et al.
Published: (2025)
by: Tong, Shaowen, et al.
Published: (2025)
Inference-Time Text-to-Video Alignment with Diffusion Latent Beam Search
by: Oshima, Yuta, et al.
Published: (2025)
by: Oshima, Yuta, et al.
Published: (2025)
LaMo: Self-Supervised Latent Motion Priors for Physical Realism in Video Generation
by: Jiang, Bo, et al.
Published: (2026)
by: Jiang, Bo, et al.
Published: (2026)
UniLumos: Fast and Unified Image and Video Relighting with Physics-Plausible Feedback
by: Liu, Ropeway, et al.
Published: (2025)
by: Liu, Ropeway, et al.
Published: (2025)
Body Part-Based Representation Learning for Occluded Person Re-Identification
by: Somers, Vladimir, et al.
Published: (2022)
by: Somers, Vladimir, et al.
Published: (2022)
Keypoint Promptable Re-Identification
by: Somers, Vladimir, et al.
Published: (2024)
by: Somers, Vladimir, et al.
Published: (2024)
Similar Items
-
Anchored Video Generation: Decoupling Scene Construction and Temporal Synthesis in Text-to-Video Diffusion Models
by: Hassan, Mariam, et al.
Published: (2025) -
Towards Generalizable Trajectory Prediction Using Dual-Level Representation Learning And Adaptive Prompting
by: Messaoud, Kaouther, et al.
Published: (2025) -
Social-Transmotion: Promptable Human Trajectory Prediction
by: Saadatnejad, Saeed, et al.
Published: (2023) -
OmniTraj: Pre-Training on Heterogeneous Data for Adaptive and Zero-Shot Human Trajectory Prediction
by: Gao, Yang, et al.
Published: (2025) -
Forecast-PEFT: Parameter-Efficient Fine-Tuning for Pre-trained Motion Forecasting Models
by: Wang, Jifeng, et al.
Published: (2024)