LikePhys: Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood Preference
Fuente:
arXiv
Saved in:
| Main Authors: | Yuan, Jianhao, Pizzati, Fabio, Pinto, Francesco, Kunze, Lars, Laptev, Ivan, Newman, Paul, Torr, Philip, De Martini, Daniele |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PhysMoDPO: Physically-Plausible Humanoid Motion with Preference Optimization
by: Zhang, Yangsong, et al.
Published: (2026)
by: Zhang, Yangsong, et al.
Published: (2026)
Video Motion Transfer with Diffusion Transformers
by: Pondaven, Alexander, et al.
Published: (2024)
by: Pondaven, Alexander, et al.
Published: (2024)
Learning to Generate Rigid Body Interactions with Video Diffusion Models
by: Romero, David, et al.
Published: (2025)
by: Romero, David, et al.
Published: (2025)
Towards Reliable Identification of Diffusion-based Image Manipulations
by: Costanzino, Alex, et al.
Published: (2025)
by: Costanzino, Alex, et al.
Published: (2025)
Not Just Pretty Pictures: Toward Interventional Data Augmentation Using Text-to-Image Generators
by: Yuan, Jianhao, et al.
Published: (2022)
by: Yuan, Jianhao, et al.
Published: (2022)
NAR-*ICP: Neural Execution of Classical ICP-based Pointcloud Registration Algorithms
by: Panagiotaki, Efimia, et al.
Published: (2024)
by: Panagiotaki, Efimia, et al.
Published: (2024)
MinkOcc: Towards real-time label-efficient semantic occupancy prediction
by: Sze, Samuel, et al.
Published: (2025)
by: Sze, Samuel, et al.
Published: (2025)
ActCam: Zero-Shot Joint Camera and 3D Motion Control for Video Generation
by: Khalifi, Omar El, et al.
Published: (2026)
by: Khalifi, Omar El, et al.
Published: (2026)
MessyKitchens: Contact-rich object-level 3D scene reconstruction
by: Ansari, Junaid Ahmed, et al.
Published: (2026)
by: Ansari, Junaid Ahmed, et al.
Published: (2026)
Specify and Edit: Overcoming Ambiguity in Text-Based Image Editing
by: Iakovleva, Ekaterina, et al.
Published: (2024)
by: Iakovleva, Ekaterina, et al.
Published: (2024)
GraphSCENE: On-Demand Critical Scenario Generation for Autonomous Vehicles in Simulation
by: Panagiotaki, Efimia, et al.
Published: (2024)
by: Panagiotaki, Efimia, et al.
Published: (2024)
ActionParty: Multi-Subject Action Binding in Generative Video Games
by: Pondaven, Alexander, et al.
Published: (2026)
by: Pondaven, Alexander, et al.
Published: (2026)
MatchDiffusion: Training-free Generation of Match-cuts
by: Pardo, Alejandro, et al.
Published: (2024)
by: Pardo, Alejandro, et al.
Published: (2024)
RAG-Driver: Generalisable Driving Explanations with Retrieval-Augmented In-Context Learning in Multi-Modal Large Language Model
by: Yuan, Jianhao, et al.
Published: (2024)
by: Yuan, Jianhao, et al.
Published: (2024)
Ensemble of Pre-Trained Models for Long-Tailed Trajectory Prediction
by: Thuremella, Divya, et al.
Published: (2025)
by: Thuremella, Divya, et al.
Published: (2025)
PSyDUCK: Training-Free Steganography for Latent Diffusion
by: Mahfuz, Aqib, et al.
Published: (2025)
by: Mahfuz, Aqib, et al.
Published: (2025)
MALT: Improving Reasoning with Multi-Agent LLM Training
by: Motwani, Sumeet Ramesh, et al.
Published: (2024)
by: Motwani, Sumeet Ramesh, et al.
Published: (2024)
IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments
by: Bordes, Florian, et al.
Published: (2025)
by: Bordes, Florian, et al.
Published: (2025)
Latent Guard: a Safety Framework for Text-to-image Generation
by: Liu, Runtao, et al.
Published: (2024)
by: Liu, Runtao, et al.
Published: (2024)
Hidden in Plain Sight: Evaluating Abstract Shape Recognition in Vision-Language Models
by: Hemmat, Arshia, et al.
Published: (2024)
by: Hemmat, Arshia, et al.
Published: (2024)
Long Story Short: Story-level Video Understanding from 20K Short Films
by: Ghermi, Ridouane, et al.
Published: (2024)
by: Ghermi, Ridouane, et al.
Published: (2024)
Real-Fake: Effective Training Data Synthesis Through Distribution Matching
by: Yuan, Jianhao, et al.
Published: (2023)
by: Yuan, Jianhao, et al.
Published: (2023)
On Pretraining Data Diversity for Self-Supervised Learning
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
Select2Plan: Training-Free ICL-Based Planning through VQA and Memory Retrieval
by: Buoso, Davide, et al.
Published: (2024)
by: Buoso, Davide, et al.
Published: (2024)
VDNA-PR: Using General Dataset Representations for Robust Sequential Visual Place Recognition
by: Ramtoula, Benjamin, et al.
Published: (2024)
by: Ramtoula, Benjamin, et al.
Published: (2024)
That's My Point: Compact Object-centric LiDAR Pose Estimation for Large-scale Outdoor Localisation
by: Pramatarov, Georgi, et al.
Published: (2024)
by: Pramatarov, Georgi, et al.
Published: (2024)
LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs
by: Wang, Xiaodong, et al.
Published: (2025)
by: Wang, Xiaodong, et al.
Published: (2025)
BLAZER: Bootstrapping LLM-based Manipulation Agents with Zero-Shot Data Generation
by: Das, Rocktim Jyoti, et al.
Published: (2025)
by: Das, Rocktim Jyoti, et al.
Published: (2025)
VFusion3D: Learning Scalable 3D Generative Models from Video Diffusion Models
by: Han, Junlin, et al.
Published: (2024)
by: Han, Junlin, et al.
Published: (2024)
Model Merging and Safety Alignment: One Bad Model Spoils the Bunch
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
Fantastic Features and Where to Find Them: A Probing Method to combine Features from Multiple Foundation Models
by: Ramtoula, Benjamin, et al.
Published: (2025)
by: Ramtoula, Benjamin, et al.
Published: (2025)
Masked Gamma-SSL: Learning Uncertainty Estimation via Masked Image Modeling
by: Williams, David S. W., et al.
Published: (2024)
by: Williams, David S. W., et al.
Published: (2024)
Mitigating Distributional Shift in Semantic Segmentation via Uncertainty Estimation from Unlabelled Data
by: Williams, David S. W., et al.
Published: (2024)
by: Williams, David S. W., et al.
Published: (2024)
Understanding Reasoning in Thinking Language Models via Steering Vectors
by: Venhoff, Constantin, et al.
Published: (2025)
by: Venhoff, Constantin, et al.
Published: (2025)
As Firm As Their Foundations: Can open-sourced foundation models be used to create adversarial examples for downstream tasks?
by: Hu, Anjun, et al.
Published: (2024)
by: Hu, Anjun, et al.
Published: (2024)
Extracting Training Data from Document-Based VQA Models
by: Pinto, Francesco, et al.
Published: (2024)
by: Pinto, Francesco, et al.
Published: (2024)
AlignGuard: Scalable Safety Alignment for Text-to-Image Generation
by: Liu, Runtao, et al.
Published: (2024)
by: Liu, Runtao, et al.
Published: (2024)
Mixture of Distributions Matters: Dynamic Sparse Attention for Efficient Video Diffusion Transformers
by: Liu, Yuxi, et al.
Published: (2026)
by: Liu, Yuxi, et al.
Published: (2026)
Real-time 3D semantic occupancy prediction for autonomous vehicles using memory-efficient sparse convolution
by: Sze, Samuel, et al.
Published: (2024)
by: Sze, Samuel, et al.
Published: (2024)
Similar Items
-
PhysMoDPO: Physically-Plausible Humanoid Motion with Preference Optimization
by: Zhang, Yangsong, et al.
Published: (2026) -
Video Motion Transfer with Diffusion Transformers
by: Pondaven, Alexander, et al.
Published: (2024) -
Learning to Generate Rigid Body Interactions with Video Diffusion Models
by: Romero, David, et al.
Published: (2025) -
Towards Reliable Identification of Diffusion-based Image Manipulations
by: Costanzino, Alex, et al.
Published: (2025) -
Not Just Pretty Pictures: Toward Interventional Data Augmentation Using Text-to-Image Generators
by: Yuan, Jianhao, et al.
Published: (2022)