LLMPhy: Parameter-Identifiable Physical Reasoning Combining Large Language Models and Physics Engines
Fuente:
arXiv
Saved in:
| Main Authors: | Cherian, Anoop, Corcodel, Radu, Jain, Siddarth, Romeres, Diego |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SpinBench: Perspective and Rotation as a Lens on Spatial Reasoning in VLMs
by: Zhang, Yuyou, et al.
Published: (2025)
by: Zhang, Yuyou, et al.
Published: (2025)
Robot Confirmation Generation and Action Planning Using Long-context Q-Former Integrated with Multimodal LLM
by: Hori, Chiori, et al.
Published: (2025)
by: Hori, Chiori, et al.
Published: (2025)
I-PHYRE: Interactive Physical Reasoning
by: Li, Shiqian, et al.
Published: (2023)
by: Li, Shiqian, et al.
Published: (2023)
Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
PhysInOne: Visual Physics Learning and Reasoning in One Suite
by: Zhou, Siyuan, et al.
Published: (2026)
by: Zhou, Siyuan, et al.
Published: (2026)
DiffGen: Robot Demonstration Generation via Differentiable Physics Simulation, Differentiable Rendering, and Vision-Language Model
by: Jin, Yang, et al.
Published: (2024)
by: Jin, Yang, et al.
Published: (2024)
CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models
by: Zhao, Qingqing, et al.
Published: (2025)
by: Zhao, Qingqing, et al.
Published: (2025)
Evaluating Large Vision-and-Language Models on Children's Mathematical Olympiads
by: Cherian, Anoop, et al.
Published: (2024)
by: Cherian, Anoop, et al.
Published: (2024)
Cosmos World Foundation Model Platform for Physical AI
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
World Simulation with Video Foundation Models for Physical AI
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
Solving Physics Olympiad via Reinforcement Learning on Physics Simulators
by: Prabhudesai, Mihir, et al.
Published: (2026)
by: Prabhudesai, Mihir, et al.
Published: (2026)
Advancing Egocentric Video Question Answering with Multimodal Large Language Models
by: Patel, Alkesh, et al.
Published: (2025)
by: Patel, Alkesh, et al.
Published: (2025)
Physically Grounded Vision-Language Models for Robotic Manipulation
by: Gao, Jensen, et al.
Published: (2023)
by: Gao, Jensen, et al.
Published: (2023)
Flame3D: Zero-shot Compositional Reasoning of 3D Scenes with Agentic Language Models
by: Bharadwaj, Sagar, et al.
Published: (2026)
by: Bharadwaj, Sagar, et al.
Published: (2026)
Efficient Driving Behavior Narration and Reasoning on Edge Device Using Large Language Models
by: Huang, Yizhou, et al.
Published: (2024)
by: Huang, Yizhou, et al.
Published: (2024)
Visual Perception Engine: Fast and Flexible Multi-Head Inference for Robotic Vision Tasks
by: Łucki, Jakub, et al.
Published: (2025)
by: Łucki, Jakub, et al.
Published: (2025)
VTAM: Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAs
by: Yuan, Haoran, et al.
Published: (2026)
by: Yuan, Haoran, et al.
Published: (2026)
Tag Map: A Text-Based Map for Spatial Reasoning and Navigation with Large Language Models
by: Zhang, Mike, et al.
Published: (2024)
by: Zhang, Mike, et al.
Published: (2024)
Seeking Physics in Diffusion Noise
by: Tang, Chujun, et al.
Published: (2026)
by: Tang, Chujun, et al.
Published: (2026)
Robotic Manipulation by Imitating Generated Videos Without Physical Demonstrations
by: Patel, Shivansh, et al.
Published: (2025)
by: Patel, Shivansh, et al.
Published: (2025)
PH-Dreamer: A Physics-Driven World Model via Port-Hamiltonian Generative Dynamics
by: Luan, Xueyu, et al.
Published: (2026)
by: Luan, Xueyu, et al.
Published: (2026)
ICAT: Incident-Case-Grounded Adaptive Testing for Physical-Risk Prediction in Embodied World Models
by: Lai, Zhenglin, et al.
Published: (2026)
by: Lai, Zhenglin, et al.
Published: (2026)
Surfer: Progressive Reasoning with World Models for Robotic Manipulation
by: Ren, Pengzhen, et al.
Published: (2023)
by: Ren, Pengzhen, et al.
Published: (2023)
ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning
by: Huang, Chi-Pin, et al.
Published: (2025)
by: Huang, Chi-Pin, et al.
Published: (2025)
PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding
by: Chow, Wei, et al.
Published: (2025)
by: Chow, Wei, et al.
Published: (2025)
Fast-ThinkAct: Efficient Vision-Language-Action Reasoning via Verbalizable Latent Planning
by: Huang, Chi-Pin, et al.
Published: (2026)
by: Huang, Chi-Pin, et al.
Published: (2026)
Vision-Language Model Fine-Tuning via Simple Parameter-Efficient Modification
by: Li, Ming, et al.
Published: (2024)
by: Li, Ming, et al.
Published: (2024)
PRISM: A Multi-View Multi-Capability Retail Video Dataset for Embodied Vision-Language Models
by: Rouhi, Amirreza, et al.
Published: (2026)
by: Rouhi, Amirreza, et al.
Published: (2026)
PhysMoDPO: Physically-Plausible Humanoid Motion with Preference Optimization
by: Zhang, Yangsong, et al.
Published: (2026)
by: Zhang, Yangsong, et al.
Published: (2026)
SORT3D: Spatial Object-centric Reasoning Toolbox for Zero-Shot 3D Grounding Using Large Language Models
by: Zantout, Nader, et al.
Published: (2025)
by: Zantout, Nader, et al.
Published: (2025)
ODIN: A Single Model for 2D and 3D Segmentation
by: Jain, Ayush, et al.
Published: (2024)
by: Jain, Ayush, et al.
Published: (2024)
PhyScene: Physically Interactable 3D Scene Synthesis for Embodied AI
by: Yang, Yandan, et al.
Published: (2024)
by: Yang, Yandan, et al.
Published: (2024)
UAV-VLA: Vision-Language-Action System for Large Scale Aerial Mission Generation
by: Sautenkov, Oleg, et al.
Published: (2025)
by: Sautenkov, Oleg, et al.
Published: (2025)
Reason--Imagine--Act: Closed-Loop LLM Decision Making with World Models for Autonomous Driving
by: Sun, Zhengqi, et al.
Published: (2026)
by: Sun, Zhengqi, et al.
Published: (2026)
Hybrid Training for Vision-Language-Action Models
by: Mazzaglia, Pietro, et al.
Published: (2025)
by: Mazzaglia, Pietro, et al.
Published: (2025)
Zero-Shot Peg Insertion: Identifying Mating Holes and Estimating SE(2) Poses with Vision-Language Models
by: Yajima, Masaru, et al.
Published: (2025)
by: Yajima, Masaru, et al.
Published: (2025)
RoCoDA: Counterfactual Data Augmentation for Data-Efficient Robot Learning from Demonstrations
by: Ameperosa, Ezra, et al.
Published: (2024)
by: Ameperosa, Ezra, et al.
Published: (2024)
LanguageMPC: Large Language Models as Decision Makers for Autonomous Driving
by: Sha, Hao, et al.
Published: (2023)
by: Sha, Hao, et al.
Published: (2023)
Out of Sight, Still in Mind: Reasoning and Planning about Unobserved Objects with Video Tracking Enabled Memory Models
by: Huang, Yixuan, et al.
Published: (2023)
by: Huang, Yixuan, et al.
Published: (2023)
Where Bits Matter in World Model Planning: A Paired Mixed-Bit Study for Efficient Spatial Reasoning
by: Ranganath, Suraj, et al.
Published: (2026)
by: Ranganath, Suraj, et al.
Published: (2026)
Similar Items
-
SpinBench: Perspective and Rotation as a Lens on Spatial Reasoning in VLMs
by: Zhang, Yuyou, et al.
Published: (2025) -
Robot Confirmation Generation and Action Planning Using Long-context Q-Former Integrated with Multimodal LLM
by: Hori, Chiori, et al.
Published: (2025) -
I-PHYRE: Interactive Physical Reasoning
by: Li, Shiqian, et al.
Published: (2023) -
Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning
by: NVIDIA, et al.
Published: (2025) -
PhysInOne: Visual Physics Learning and Reasoning in One Suite
by: Zhou, Siyuan, et al.
Published: (2026)