IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bordes, Florian, Garrido, Quentin, Kao, Justine T, Williams, Adina, Rabbat, Michael, Dupoux, Emmanuel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Intuitive physics understanding emerges from self-supervised pretraining on natural videos
von: Garrido, Quentin, et al.
Veröffentlicht: (2025)
von: Garrido, Quentin, et al.
Veröffentlicht: (2025)
What's in Common? Multimodal Models Hallucinate When Reasoning Across Scenes
von: Ross, Candace, et al.
Veröffentlicht: (2025)
von: Ross, Candace, et al.
Veröffentlicht: (2025)
LikePhys: Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood Preference
von: Yuan, Jianhao, et al.
Veröffentlicht: (2025)
von: Yuan, Jianhao, et al.
Veröffentlicht: (2025)
PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs
von: Zhang, Zixin, et al.
Veröffentlicht: (2025)
von: Zhang, Zixin, et al.
Veröffentlicht: (2025)
Learning Latent Action World Models In The Wild
von: Garrido, Quentin, et al.
Veröffentlicht: (2026)
von: Garrido, Quentin, et al.
Veröffentlicht: (2026)
Interpreting Physics in Video World Models
von: Joseph, Sonia, et al.
Veröffentlicht: (2026)
von: Joseph, Sonia, et al.
Veröffentlicht: (2026)
A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs
von: Krojer, Benno, et al.
Veröffentlicht: (2025)
von: Krojer, Benno, et al.
Veröffentlicht: (2025)
PhysDepth: Plug-and-Play Physical Refinement for Monocular Depth Estimation in Challenging Environments
von: Peng, Kebin, et al.
Veröffentlicht: (2024)
von: Peng, Kebin, et al.
Veröffentlicht: (2024)
PhysLab: A Benchmark Dataset for Multi-Granularity Visual Parsing of Physics Experiments
von: Zou, Minghao, et al.
Veröffentlicht: (2025)
von: Zou, Minghao, et al.
Veröffentlicht: (2025)
PhysVLM-AVR: Active Visual Reasoning for Multimodal Large Language Models in Physical Environments
von: Zhou, Weijie, et al.
Veröffentlicht: (2025)
von: Zhou, Weijie, et al.
Veröffentlicht: (2025)
PhysVLM: Enabling Visual Language Models to Understand Robotic Physical Reachability
von: Zhou, Weijie, et al.
Veröffentlicht: (2025)
von: Zhou, Weijie, et al.
Veröffentlicht: (2025)
Revisiting Feature Prediction for Learning Visual Representations from Video
von: Bardes, Adrien, et al.
Veröffentlicht: (2024)
von: Bardes, Adrien, et al.
Veröffentlicht: (2024)
PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding
von: Chow, Wei, et al.
Veröffentlicht: (2025)
von: Chow, Wei, et al.
Veröffentlicht: (2025)
PhysEditBench: A Protocol-Conditioned Benchmark for Dense Physical-Map Prediction with Image Editors
von: Yang, Jiaxin, et al.
Veröffentlicht: (2026)
von: Yang, Jiaxin, et al.
Veröffentlicht: (2026)
Diffusion-based 3D Hand Motion Recovery with Intuitive Physics
von: Zhang, Yufei, et al.
Veröffentlicht: (2025)
von: Zhang, Yufei, et al.
Veröffentlicht: (2025)
MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
von: Tong, Shengbang, et al.
Veröffentlicht: (2024)
von: Tong, Shengbang, et al.
Veröffentlicht: (2024)
ForestSim: A Synthetic Benchmark for Intelligent Vehicle Perception in Unstructured Forest Environments
von: Wagle, Pragat, et al.
Veröffentlicht: (2026)
von: Wagle, Pragat, et al.
Veröffentlicht: (2026)
Grounding Social Perception in Intuitive Physics
von: Ying, Lance, et al.
Veröffentlicht: (2026)
von: Ying, Lance, et al.
Veröffentlicht: (2026)
PhysCtrl: Generative Physics for Controllable and Physics-Grounded Video Generation
von: Wang, Chen, et al.
Veröffentlicht: (2025)
von: Wang, Chen, et al.
Veröffentlicht: (2025)
CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models
von: Foss, Aaron, et al.
Veröffentlicht: (2025)
von: Foss, Aaron, et al.
Veröffentlicht: (2025)
Beyond Static Vision: Scene Dynamic Field Unlocks Intuitive Physics Understanding in Multi-modal Large Language Models
von: Li, Nanxi, et al.
Veröffentlicht: (2026)
von: Li, Nanxi, et al.
Veröffentlicht: (2026)
A Picture is Worth More Than 77 Text Tokens: Evaluating CLIP-Style Models on Dense Captions
von: Urbanek, Jack, et al.
Veröffentlicht: (2023)
von: Urbanek, Jack, et al.
Veröffentlicht: (2023)
Eval Factsheets: A Structured Framework for Documenting AI Evaluations
von: Bordes, Florian, et al.
Veröffentlicht: (2025)
von: Bordes, Florian, et al.
Veröffentlicht: (2025)
Measuring Déjà vu Memorization Efficiently
von: Kokhlikyan, Narine, et al.
Veröffentlicht: (2025)
von: Kokhlikyan, Narine, et al.
Veröffentlicht: (2025)
PhysAvatar: Learning the Physics of Dressed 3D Avatars from Visual Observations
von: Zheng, Yang, et al.
Veröffentlicht: (2024)
von: Zheng, Yang, et al.
Veröffentlicht: (2024)
PhysGame: Uncovering Physical Commonsense Violations in Gameplay Videos
von: Cao, Meng, et al.
Veröffentlicht: (2024)
von: Cao, Meng, et al.
Veröffentlicht: (2024)
Evaluation of Conversational Agents: Understanding Culture, Context and Environment in Emotion Detection
von: Teye, Martha Teiko, et al.
Veröffentlicht: (2026)
von: Teye, Martha Teiko, et al.
Veröffentlicht: (2026)
PhysAnimator: Physics-Guided Generative Cartoon Animation
von: Xie, Tianyi, et al.
Veröffentlicht: (2025)
von: Xie, Tianyi, et al.
Veröffentlicht: (2025)
PhysMotion: Physics-Grounded Dynamics From a Single Image
von: Tan, Xiyang, et al.
Veröffentlicht: (2024)
von: Tan, Xiyang, et al.
Veröffentlicht: (2024)
PhysLayer: Language-Guided Layered Animation with Depth-Aware Physics
von: Xie, Tianyidan, et al.
Veröffentlicht: (2026)
von: Xie, Tianyidan, et al.
Veröffentlicht: (2026)
ReconPhys: Reconstruct Appearance and Physical Attributes from Single Video
von: Wang, Boyuan, et al.
Veröffentlicht: (2026)
von: Wang, Boyuan, et al.
Veröffentlicht: (2026)
Understanding Contrastive Representation Learning from Positive Unlabeled (PU) Data
von: Acharya, Anish, et al.
Veröffentlicht: (2024)
von: Acharya, Anish, et al.
Veröffentlicht: (2024)
Object-centric Binding in Contrastive Language-Image Pretraining
von: Assouel, Rim, et al.
Veröffentlicht: (2025)
von: Assouel, Rim, et al.
Veröffentlicht: (2025)
A Lightweight Library for Energy-Based Joint-Embedding Predictive Architectures
von: Terver, Basile, et al.
Veröffentlicht: (2026)
von: Terver, Basile, et al.
Veröffentlicht: (2026)
Benchmarking Vision-Based Object Tracking for USVs in Complex Maritime Environments
von: Din, Muhayy Ud, et al.
Veröffentlicht: (2024)
von: Din, Muhayy Ud, et al.
Veröffentlicht: (2024)
Learning to Play Video Games with Intuitive Physics Priors
von: Jaiswal, Abhishek, et al.
Veröffentlicht: (2024)
von: Jaiswal, Abhishek, et al.
Veröffentlicht: (2024)
PhysX-3D: Physical-Grounded 3D Asset Generation
von: Cao, Ziang, et al.
Veröffentlicht: (2025)
von: Cao, Ziang, et al.
Veröffentlicht: (2025)
PhysGen: Physically Grounded 3D Shape Generation for Industrial Design
von: You, Yingxuan, et al.
Veröffentlicht: (2025)
von: You, Yingxuan, et al.
Veröffentlicht: (2025)
MultiPhys: Multi-Person Physics-aware 3D Motion Estimation
von: Ugrinovic, Nicolas, et al.
Veröffentlicht: (2024)
von: Ugrinovic, Nicolas, et al.
Veröffentlicht: (2024)
PhysRVG: Physics-Aware Unified Reinforcement Learning for Video Generative Models
von: Zhang, Qiyuan, et al.
Veröffentlicht: (2026)
von: Zhang, Qiyuan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Intuitive physics understanding emerges from self-supervised pretraining on natural videos
von: Garrido, Quentin, et al.
Veröffentlicht: (2025) -
What's in Common? Multimodal Models Hallucinate When Reasoning Across Scenes
von: Ross, Candace, et al.
Veröffentlicht: (2025) -
LikePhys: Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood Preference
von: Yuan, Jianhao, et al.
Veröffentlicht: (2025) -
PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs
von: Zhang, Zixin, et al.
Veröffentlicht: (2025) -
Learning Latent Action World Models In The Wild
von: Garrido, Quentin, et al.
Veröffentlicht: (2026)