Thinking in Dynamics: How Multimodal Large Language Models Perceive, Track, and Reason Dynamics in Physical 4D World
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Yuzhi, Wen, Kairun, Gao, Rongxin, Liu, Dongxuan, Lou, Yibin, Wu, Jie, Xu, Jing, Zhang, Jian, Yang, Zheng, Lin, Yunlong, Li, Chenxin, Pan, Panwang, Lu, Junbin, Jiang, Jingyan, Ding, Xinghao, Huang, Yue, Wang, Zhi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DynamicVerse: A Physically-Aware Multimodal Framework for 4D World Modeling
by: Wen, Kairun, et al.
Published: (2025)
by: Wen, Kairun, et al.
Published: (2025)
HumanCrafter: Synergizing Generalizable Human Reconstruction and Semantic 3D Segmentation
by: Pan, Panwang, et al.
Published: (2025)
by: Pan, Panwang, et al.
Published: (2025)
Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models
by: Pan, Panwang, et al.
Published: (2025)
by: Pan, Panwang, et al.
Published: (2025)
Pan-LUT: Efficient Pan-sharpening via Learnable Look-Up Tables
by: Cai, Zhongnan, et al.
Published: (2025)
by: Cai, Zhongnan, et al.
Published: (2025)
Track Any Anomalous Object: A Granular Video Anomaly Detection Pipeline
by: Huang, Yuzhi, et al.
Published: (2025)
by: Huang, Yuzhi, et al.
Published: (2025)
JarvisIR: Elevating Autonomous Driving Perception with Intelligent Image Restoration
by: Lin, Yunlong, et al.
Published: (2025)
by: Lin, Yunlong, et al.
Published: (2025)
PCMamba: Physics-Informed Cross-Modal State Space Model for Dual-Camera Compressive Hyperspectral Imaging
by: Meng, Ge, et al.
Published: (2025)
by: Meng, Ge, et al.
Published: (2025)
JarvisArt: Liberating Human Artistic Creativity via an Intelligent Photo Retouching Agent
by: Lin, Yunlong, et al.
Published: (2025)
by: Lin, Yunlong, et al.
Published: (2025)
Unsupervised Low-light Image Enhancement with Lookup Tables and Diffusion Priors
by: Lin, Yunlong, et al.
Published: (2024)
by: Lin, Yunlong, et al.
Published: (2024)
World2VLM: Distilling World Model Imagination into VLMs for Dynamic Spatial Reasoning
by: Zhang, Wanyue, et al.
Published: (2026)
by: Zhang, Wanyue, et al.
Published: (2026)
FRN: Fractal-Based Recursive Spectral Reconstruction Network
by: Meng, Ge, et al.
Published: (2025)
by: Meng, Ge, et al.
Published: (2025)
Dissecting Generalized Category Discovery: Multiplex Consensus under Self-Deconstruction
by: Tang, Luyao, et al.
Published: (2025)
by: Tang, Luyao, et al.
Published: (2025)
Think in Safety: Unveiling and Mitigating Safety Alignment Collapse in Multimodal Large Reasoning Model
by: Lou, Xinyue, et al.
Published: (2025)
by: Lou, Xinyue, et al.
Published: (2025)
RoboStream: Weaving Spatio-Temporal Reasoning with Memory in Vision-Language Models for Robotics
by: Huang, Yuzhi, et al.
Published: (2026)
by: Huang, Yuzhi, et al.
Published: (2026)
Beyond English-Centric Training: How Reinforcement Learning Improves Cross-Lingual Reasoning in LLMs
by: Huang, Shulin, et al.
Published: (2025)
by: Huang, Shulin, et al.
Published: (2025)
Boosting Accelerated Proximal Gradient Method with Adaptive Sampling for Stochastic Composite Optimization
by: Zhu, Dongxuan, et al.
Published: (2025)
by: Zhu, Dongxuan, et al.
Published: (2025)
CheXWorld: Exploring Image World Modeling for Radiograph Representation Learning
by: Yue, Yang, et al.
Published: (2025)
by: Yue, Yang, et al.
Published: (2025)
Discover Your Neighbors: Advanced Stable Test-Time Adaptation in Dynamic World
by: Jiang, Qinting, et al.
Published: (2024)
by: Jiang, Qinting, et al.
Published: (2024)
SAT: Balancing Reasoning Accuracy and Efficiency with Stepwise Adaptive Thinking
by: Huang, Weiyang, et al.
Published: (2026)
by: Huang, Weiyang, et al.
Published: (2026)
ThinkBench: Dynamic Out-of-Distribution Evaluation for Robust LLM Reasoning
by: Huang, Shulin, et al.
Published: (2025)
by: Huang, Shulin, et al.
Published: (2025)
AtomThink: Multimodal Slow Thinking with Atomic Step Reasoning
by: Xiang, Kun, et al.
Published: (2024)
by: Xiang, Kun, et al.
Published: (2024)
Detection of a Sparse Change in High-Dimensional Time Series
by: Huang, Jingyan
Published: (2025)
by: Huang, Jingyan
Published: (2025)
GaussianStego: A Generalizable Stenography Pipeline for Generative 3D Gaussians Splatting
by: Li, Chenxin, et al.
Published: (2024)
by: Li, Chenxin, et al.
Published: (2024)
DynOPETs: A Versatile Benchmark for Dynamic Object Pose Estimation and Tracking in Moving Camera Scenarios
by: Meng, Xiangting, et al.
Published: (2025)
by: Meng, Xiangting, et al.
Published: (2025)
Fully Dynamic Spectral and Cut Sparsifiers for Directed Graphs
by: Zhao, Yibin
Published: (2025)
by: Zhao, Yibin
Published: (2025)
Feature-Based Instance Neighbor Discovery: Advanced Stable Test-Time Adaptation in Dynamic World
by: Jiang, Qinting, et al.
Published: (2025)
by: Jiang, Qinting, et al.
Published: (2025)
Test-time Diverse Reasoning by Riemannian Activation Steering
by: Khanh, Ly Tran Ho, et al.
Published: (2025)
by: Khanh, Ly Tran Ho, et al.
Published: (2025)
TwiSTAR:Think Fast, Think Slow, Then Act,Generative Recommendation with Adaptive Reasoning
by: Cao, Shiteng, et al.
Published: (2026)
by: Cao, Shiteng, et al.
Published: (2026)
Think Dense, Not Long: Dynamic Decoupled Conditional Advantage for Efficient Reasoning
by: Peng, Keqin, et al.
Published: (2026)
by: Peng, Keqin, et al.
Published: (2026)
MM-UAVBench: How Well Do Multimodal Large Language Models See, Think, and Plan in Low-Altitude UAV Scenarios?
by: Dai, Shiqi, et al.
Published: (2025)
by: Dai, Shiqi, et al.
Published: (2025)
The Dynamics of Relational Innovation Leadership: A Case Study of Lip‐Bu Tan
by: Eric Li, et al.
Published: (2026)
by: Eric Li, et al.
Published: (2026)
WorldMM: Dynamic Multimodal Memory Agent for Long Video Reasoning
by: Yeo, Woongyeong, et al.
Published: (2025)
by: Yeo, Woongyeong, et al.
Published: (2025)
DASICS White Paper: Enhancing Memory Protection with Dynamic Compartmentalization
by: Jin, Yue, et al.
Published: (2023)
by: Jin, Yue, et al.
Published: (2023)
Think Natively: Unlocking Multilingual Reasoning with Consistency-Enhanced Reinforcement Learning
by: Zhang, Xue, et al.
Published: (2025)
by: Zhang, Xue, et al.
Published: (2025)
Don't Think Longer, Think Wisely: Optimizing Thinking Dynamics for Large Reasoning Models
by: An, Sohyun, et al.
Published: (2025)
by: An, Sohyun, et al.
Published: (2025)
Modeling LLM Agent Reviewer Dynamics in Elo-Ranked Review System
by: Huang, Hsiang-Wei, et al.
Published: (2026)
by: Huang, Hsiang-Wei, et al.
Published: (2026)
How Emotional Appeals and Deliberative Thinking Relate to Perceived Effectiveness
by: Shuichiro Kawaguchi, et al.
Published: (2026)
by: Shuichiro Kawaguchi, et al.
Published: (2026)
An optimization-based equilibrium measure describes non-equilibrium steady state dynamics: application to edge of chaos
by: Qiu, Junbin, et al.
Published: (2024)
by: Qiu, Junbin, et al.
Published: (2024)
ThinkQE: Query Expansion via an Evolving Thinking Process
by: Lei, Yibin, et al.
Published: (2025)
by: Lei, Yibin, et al.
Published: (2025)
ID-Crafter: VLM-Grounded Online RL for Compositional Multi-Subject Video Generation
by: Pan, Panwang, et al.
Published: (2025)
by: Pan, Panwang, et al.
Published: (2025)
Similar Items
-
DynamicVerse: A Physically-Aware Multimodal Framework for 4D World Modeling
by: Wen, Kairun, et al.
Published: (2025) -
HumanCrafter: Synergizing Generalizable Human Reconstruction and Semantic 3D Segmentation
by: Pan, Panwang, et al.
Published: (2025) -
Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models
by: Pan, Panwang, et al.
Published: (2025) -
Pan-LUT: Efficient Pan-sharpening via Learnable Look-Up Tables
by: Cai, Zhongnan, et al.
Published: (2025) -
Track Any Anomalous Object: A Granular Video Anomaly Detection Pipeline
by: Huang, Yuzhi, et al.
Published: (2025)