Can Your Model Separate Yolks with a Water Bottle? Benchmarking Physical Commonsense Understanding in Video Generation Models
Fuente:
arXiv
Saved in:
| Main Authors: | Sanli, Enes, Tezcan, Baris Sarper, Erdem, Aykut, Erdem, Erkut |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LAMP: Language-Assisted Motion Planning for Controllable Video Generation
by: Kizil, Muhammed Burak, et al.
Published: (2025)
by: Kizil, Muhammed Burak, et al.
Published: (2025)
Auteur: Language-Driven Cinematographic Framing for Human-Centric Video Generation
by: Kizil, Muhammed Burak, et al.
Published: (2026)
by: Kizil, Muhammed Burak, et al.
Published: (2026)
A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features
by: Karanfil, Enes, et al.
Published: (2025)
by: Karanfil, Enes, et al.
Published: (2025)
EVREAL: Towards a Comprehensive Benchmark and Analysis Suite for Event-based Video Reconstruction
by: Ercan, Burak, et al.
Published: (2023)
by: Ercan, Burak, et al.
Published: (2023)
HUE Dataset: High-Resolution Event and Frame Sequences for Low-Light Vision
by: Ercan, Burak, et al.
Published: (2024)
by: Ercan, Burak, et al.
Published: (2024)
HyperE2VID: Improving Event-Based Video Reconstruction via Hypernetworks
by: Ercan, Burak, et al.
Published: (2023)
by: Ercan, Burak, et al.
Published: (2023)
GaussianVideo: Efficient Video Representation via Hierarchical Gaussian Splatting
by: Bond, Andrew, et al.
Published: (2025)
by: Bond, Andrew, et al.
Published: (2025)
CLIPAway: Harmonizing Focused Embeddings for Removing Objects via Diffusion Models
by: Ekin, Yigit, et al.
Published: (2024)
by: Ekin, Yigit, et al.
Published: (2024)
Beyond Gaussian Bottlenecks: Topologically Aligned Encoding of Vision-Transformer Feature Spaces
by: Bond, Andrew, et al.
Published: (2026)
by: Bond, Andrew, et al.
Published: (2026)
Evaluating Linguistic Capabilities of Multimodal LLMs in the Lens of Few-Shot Learning
by: Dogan, Mustafa, et al.
Published: (2024)
by: Dogan, Mustafa, et al.
Published: (2024)
SonicDiffusion: Audio-Driven Image Generation and Editing with Pretrained Diffusion Models
by: Biner, Burak Can, et al.
Published: (2024)
by: Biner, Burak Can, et al.
Published: (2024)
Spherical Vision Transformers for Audio-Visual Saliency Prediction in 360-Degree Videos
by: Cokelek, Mert, et al.
Published: (2025)
by: Cokelek, Mert, et al.
Published: (2025)
Conflated Inverse Modeling to Generate Diverse and Temperature-Change Inducing Urban Vegetation Patterns
by: Tezcan, Baris Sarper, et al.
Published: (2026)
by: Tezcan, Baris Sarper, et al.
Published: (2026)
VidStyleODE: Disentangled Video Editing via StyleGAN and NeuralODEs
by: Ali, Moayed Haji, et al.
Published: (2023)
by: Ali, Moayed Haji, et al.
Published: (2023)
TanDiT: Tangent-Plane Diffusion Transformer for High-Quality 360° Panorama Generation
by: Çapuk, Hakan, et al.
Published: (2025)
by: Çapuk, Hakan, et al.
Published: (2025)
HyperGAN-CLIP: A Unified Framework for Domain Adaptation, Image Synthesis and Manipulation
by: Anees, Abdul Basit, et al.
Published: (2024)
by: Anees, Abdul Basit, et al.
Published: (2024)
Commonsense-T2I Challenge: Can Text-to-Image Generation Models Understand Commonsense?
by: Fu, Xingyu, et al.
Published: (2024)
by: Fu, Xingyu, et al.
Published: (2024)
How to Augment for Atmospheric Turbulence Effects on Thermal Adapted Object Detection Models?
by: Uzun, Engin, et al.
Published: (2024)
by: Uzun, Engin, et al.
Published: (2024)
Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation
by: Meng, Fanqing, et al.
Published: (2024)
by: Meng, Fanqing, et al.
Published: (2024)
Deep Learning-based Average Shear Wave Velocity Prediction using Accelerometer Records
by: Yılmaz, Barış, et al.
Published: (2024)
by: Yılmaz, Barış, et al.
Published: (2024)
PhyBench: A Physical Commonsense Benchmark for Evaluating Text-to-Image Models
by: Meng, Fanqing, et al.
Published: (2024)
by: Meng, Fanqing, et al.
Published: (2024)
Infrared Domain Adaptation with Zero-Shot Quantization
by: Sevsay, Burak, et al.
Published: (2024)
by: Sevsay, Burak, et al.
Published: (2024)
FuseFormer: A Transformer for Visual and Thermal Image Fusion
by: Erdogan, Aytekin, et al.
Published: (2024)
by: Erdogan, Aytekin, et al.
Published: (2024)
VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation
by: Bansal, Hritik, et al.
Published: (2025)
by: Bansal, Hritik, et al.
Published: (2025)
VideoPhy: Evaluating Physical Commonsense for Video Generation
by: Bansal, Hritik, et al.
Published: (2024)
by: Bansal, Hritik, et al.
Published: (2024)
FewMMBench: A Benchmark for Multimodal Few-Shot Learning
by: Dogan, Mustafa, et al.
Published: (2026)
by: Dogan, Mustafa, et al.
Published: (2026)
PhysGame: Uncovering Physical Commonsense Violations in Gameplay Videos
by: Cao, Meng, et al.
Published: (2024)
by: Cao, Meng, et al.
Published: (2024)
Enhancing Visual Question Answering through Question-Driven Image Captions as Prompts
by: Özdemir, Övgü, et al.
Published: (2024)
by: Özdemir, Övgü, et al.
Published: (2024)
ORIC: Benchmarking Object Recognition under Contextual Incongruity in Large Vision-Language Models
by: Li, Zhaoyang, et al.
Published: (2025)
by: Li, Zhaoyang, et al.
Published: (2025)
Morpheus: Benchmarking Physical Reasoning of Video Generative Models with Real Physical Experiments
by: Zhang, Chenyu, et al.
Published: (2025)
by: Zhang, Chenyu, et al.
Published: (2025)
Sequential Compositional Generalization in Multimodal Models
by: Yagcioglu, Semih, et al.
Published: (2024)
by: Yagcioglu, Semih, et al.
Published: (2024)
UVLM: Benchmarking Video Language Model for Underwater World Understanding
by: Xue, Xizhe, et al.
Published: (2025)
by: Xue, Xizhe, et al.
Published: (2025)
Exploring Challenges in Deep Learning of Single-Station Ground Motion Records
by: Çağlar, Ümit Mert, et al.
Published: (2024)
by: Çağlar, Ümit Mert, et al.
Published: (2024)
Implicit and Explicit Commonsense for Multi-sentence Video Captioning
by: Chou, Shih-Han, et al.
Published: (2023)
by: Chou, Shih-Han, et al.
Published: (2023)
Droplet3D: Commonsense Priors from Videos Facilitate 3D Generation
by: Li, Xiaochuan, et al.
Published: (2025)
by: Li, Xiaochuan, et al.
Published: (2025)
VideoGLUE: Video General Understanding Evaluation of Foundation Models
by: Yuan, Liangzhe, et al.
Published: (2023)
by: Yuan, Liangzhe, et al.
Published: (2023)
Can Generative Video Models Help Pose Estimation?
by: Cai, Ruojin, et al.
Published: (2024)
by: Cai, Ruojin, et al.
Published: (2024)
CRONOS: Benchmarking Counterfactual Physical Consistency in Video Models
by: Begiristain, León, et al.
Published: (2026)
by: Begiristain, León, et al.
Published: (2026)
Exploring Hallucination of Large Multimodal Models in Video Understanding: Benchmark, Analysis and Mitigation
by: Gao, Hongcheng, et al.
Published: (2025)
by: Gao, Hongcheng, et al.
Published: (2025)
VideoVerse: Does Your T2V Generator Have World Model Capability to Synthesize Videos?
by: Wang, Zeqing, et al.
Published: (2025)
by: Wang, Zeqing, et al.
Published: (2025)
Similar Items
-
LAMP: Language-Assisted Motion Planning for Controllable Video Generation
by: Kizil, Muhammed Burak, et al.
Published: (2025) -
Auteur: Language-Driven Cinematographic Framing for Human-Centric Video Generation
by: Kizil, Muhammed Burak, et al.
Published: (2026) -
A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features
by: Karanfil, Enes, et al.
Published: (2025) -
EVREAL: Towards a Comprehensive Benchmark and Analysis Suite for Event-based Video Reconstruction
by: Ercan, Burak, et al.
Published: (2023) -
HUE Dataset: High-Resolution Event and Frame Sequences for Low-Light Vision
by: Ercan, Burak, et al.
Published: (2024)