Enhancing Physical Plausibility in Video Generation by Reasoning the Implausibility
Fuente:
arXiv
Saved in:
| Main Authors: | Hao, Yutong, Chen, Chen, Mian, Ajmal Saeed, Xu, Chang, Liu, Daochang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Implicit Neural Representation-Based Continuous Single Image Super-Resolution: An Empirical Benchmark
by: Nasir, Tayyab, et al.
Published: (2026)
by: Nasir, Tayyab, et al.
Published: (2026)
Multi-scale Coarse-to-fine Modeling for Test-time Human Motion Control
by: Le, Nhat, et al.
Published: (2026)
by: Le, Nhat, et al.
Published: (2026)
NAIMA: Semantics Aware RGB Guided Depth Super-Resolution
by: Nasir, Tayyab, et al.
Published: (2026)
by: Nasir, Tayyab, et al.
Published: (2026)
Generative Physical AI in Vision: A Survey
by: Liu, Daochang, et al.
Published: (2025)
by: Liu, Daochang, et al.
Published: (2025)
Enhancing Privacy-Utility Trade-offs to Mitigate Memorization in Diffusion Models
by: Chen, Chen, et al.
Published: (2025)
by: Chen, Chen, et al.
Published: (2025)
Investigating Memorization in Video Diffusion Models
by: Chen, Chen, et al.
Published: (2024)
by: Chen, Chen, et al.
Published: (2024)
Towards Memorization-Free Diffusion Models
by: Chen, Chen, et al.
Published: (2024)
by: Chen, Chen, et al.
Published: (2024)
Class-Partitioned VQ-VAE and Latent Flow Matching for Point Cloud Scene Generation
by: Edirimuni, Dasith de Silva, et al.
Published: (2026)
by: Edirimuni, Dasith de Silva, et al.
Published: (2026)
Exploring Local Memorization in Diffusion Models via Bright Ending Attention
by: Chen, Chen, et al.
Published: (2024)
by: Chen, Chen, et al.
Published: (2024)
Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level
by: Deng, Andong, et al.
Published: (2024)
by: Deng, Andong, et al.
Published: (2024)
TRAVL: A Recipe for Making Video-Language Models Better Judges of Physics Implausibility
by: Motamed, Saman, et al.
Published: (2025)
by: Motamed, Saman, et al.
Published: (2025)
Faithful Extreme Image Rescaling with Learnable Reversible Transformation and Semantic Priors
by: Wei, Hao, et al.
Published: (2026)
by: Wei, Hao, et al.
Published: (2026)
Exploiting Semantic and Pixel Representations for Ultra-Low Bitrate Image Compression
by: Wei, Hao, et al.
Published: (2026)
by: Wei, Hao, et al.
Published: (2026)
Mantis: Mamba-native Tuning is Efficient for 3D Point Cloud Foundation Models
by: Guo, Zihao, et al.
Published: (2026)
by: Guo, Zihao, et al.
Published: (2026)
Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation
by: Chen, Harold Haodong, et al.
Published: (2025)
by: Chen, Harold Haodong, et al.
Published: (2025)
Chain of Event-Centric Causal Thought for Physically Plausible Video Generation
by: Wang, Zixuan, et al.
Published: (2026)
by: Wang, Zixuan, et al.
Published: (2026)
Auto-Regressive Diffusion for Generating 3D Human-Object Interactions
by: Geng, Zichen, et al.
Published: (2025)
by: Geng, Zichen, et al.
Published: (2025)
PointDiffuse: A Dual-Conditional Diffusion Model for Enhanced Point Cloud Semantic Segmentation
by: He, Yong, et al.
Published: (2025)
by: He, Yong, et al.
Published: (2025)
Sparse Points to Dense Clouds: Enhancing 3D Detection with Limited LiDAR Data
by: Kumar, Aakash, et al.
Published: (2024)
by: Kumar, Aakash, et al.
Published: (2024)
SDFA: Structure Aware Discriminative Feature Aggregation for Efficient Human Fall Detection in Video
by: Zahan, Sania, et al.
Published: (2025)
by: Zahan, Sania, et al.
Published: (2025)
OrthoPhys: Physically Plausible Video Generation with Orthogonal-View Geometry Guidance
by: Wang, Cong, et al.
Published: (2026)
by: Wang, Cong, et al.
Published: (2026)
Diffusion Masked Pretraining for Dynamic Point Cloud
by: Zhang, Zhuoyue, et al.
Published: (2026)
by: Zhang, Zhuoyue, et al.
Published: (2026)
Context-Enhanced Video Moment Retrieval with Large Language Models
by: Liu, Weijia, et al.
Published: (2024)
by: Liu, Weijia, et al.
Published: (2024)
Draft-and-Target Sampling for Video Generation Policy
by: Zhang, Qikang, et al.
Published: (2026)
by: Zhang, Qikang, et al.
Published: (2026)
Tempered Self-Similarity Alignment for Physically Plausible Video Generation
by: Kim, Manjin, et al.
Published: (2026)
by: Kim, Manjin, et al.
Published: (2026)
Temporally Consistent Referring Video Object Segmentation with Hybrid Memory
by: Miao, Bo, et al.
Published: (2024)
by: Miao, Bo, et al.
Published: (2024)
MMPhysVideo: Scaling Physical Plausibility in Video Generation via Joint Multimodal Modeling
by: Lin, Shubo, et al.
Published: (2026)
by: Lin, Shubo, et al.
Published: (2026)
UniLumos: Fast and Unified Image and Video Relighting with Physics-Plausible Feedback
by: Liu, Ropeway, et al.
Published: (2025)
by: Liu, Ropeway, et al.
Published: (2025)
ARMFlow: AutoRegressive MeanFlow for Online 3D Human Reaction Generation
by: Geng, Zichen, et al.
Published: (2025)
by: Geng, Zichen, et al.
Published: (2025)
Compress Guidance in Conditional Diffusion Sampling
by: Dinh, Anh-Dung, et al.
Published: (2024)
by: Dinh, Anh-Dung, et al.
Published: (2024)
From Generated Human Videos to Physically Plausible Robot Trajectories
by: Ni, James, et al.
Published: (2025)
by: Ni, James, et al.
Published: (2025)
Multistream Network for LiDAR and Camera-based 3D Object Detection in Outdoor Scenes
by: Ibrahim, Muhammad, et al.
Published: (2025)
by: Ibrahim, Muhammad, et al.
Published: (2025)
External Knowledge Enhanced 3D Scene Generation from Sketch
by: Wu, Zijie, et al.
Published: (2024)
by: Wu, Zijie, et al.
Published: (2024)
Deep Learning Based 3D Segmentation: A Survey
by: He, Yong, et al.
Published: (2021)
by: He, Yong, et al.
Published: (2021)
Avatar Concept Slider: Controllable Editing of Concepts in 3D Human Avatars
by: Foo, Lin Geng, et al.
Published: (2024)
by: Foo, Lin Geng, et al.
Published: (2024)
Modeling Human Skeleton Joint Dynamics for Fall Detection
by: Zahan, Sania, et al.
Published: (2025)
by: Zahan, Sania, et al.
Published: (2025)
Occlusion-aware Text-Image-Point Cloud Pretraining for Open-World 3D Object Recognition
by: Nguyen, Khanh, et al.
Published: (2025)
by: Nguyen, Khanh, et al.
Published: (2025)
Dynamic watermarks in images generated by diffusion models
by: Chen, Yunzhuo, et al.
Published: (2025)
by: Chen, Yunzhuo, et al.
Published: (2025)
Deepfake Detection with Spatio-Temporal Consistency and Attention
by: Chen, Yunzhuo, et al.
Published: (2025)
by: Chen, Yunzhuo, et al.
Published: (2025)
Simultaneous Multiple Object Detection and Pose Estimation using 3D Model Infusion with Monocular Vision
by: Li, Congliang, et al.
Published: (2022)
by: Li, Congliang, et al.
Published: (2022)
Similar Items
-
Implicit Neural Representation-Based Continuous Single Image Super-Resolution: An Empirical Benchmark
by: Nasir, Tayyab, et al.
Published: (2026) -
Multi-scale Coarse-to-fine Modeling for Test-time Human Motion Control
by: Le, Nhat, et al.
Published: (2026) -
NAIMA: Semantics Aware RGB Guided Depth Super-Resolution
by: Nasir, Tayyab, et al.
Published: (2026) -
Generative Physical AI in Vision: A Survey
by: Liu, Daochang, et al.
Published: (2025) -
Enhancing Privacy-Utility Trade-offs to Mitigate Memorization in Diffusion Models
by: Chen, Chen, et al.
Published: (2025)