Thinking Ahead: Foresight Intelligence in MLLMs and World Models
Fuente:
arXiv
Saved in:
| Main Authors: | Gong, Zhantao, Fan, Liaoyuan, Guo, Qing, Xu, Xun, Yang, Xulei, Li, Shijie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Ego2World: Compiling Egocentric Cooking Videos into Executable Worlds for Belief-State Planning
by: Cheng, Qinchuan, et al.
Published: (2026)
by: Cheng, Qinchuan, et al.
Published: (2026)
GRIT: Teaching MLLMs to Think with Images
by: Fan, Yue, et al.
Published: (2025)
by: Fan, Yue, et al.
Published: (2025)
ETA: Efficiency through Thinking Ahead, A Dual Approach to Self-Driving with Large Models
by: Hamdan, Shadi, et al.
Published: (2025)
by: Hamdan, Shadi, et al.
Published: (2025)
CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving
by: Huang, Minqing, et al.
Published: (2026)
by: Huang, Minqing, et al.
Published: (2026)
Chain of World: World Model Thinking in Latent Motion
by: Yang, Fuxiang, et al.
Published: (2026)
by: Yang, Fuxiang, et al.
Published: (2026)
RynnEC: Bringing MLLMs into Embodied World
by: Dang, Ronghao, et al.
Published: (2025)
by: Dang, Ronghao, et al.
Published: (2025)
Towards Unified World Models for Visual Navigation via Memory-Augmented Planning and Foresight
by: Dong, Yifei, et al.
Published: (2025)
by: Dong, Yifei, et al.
Published: (2025)
Visual Foresight for Robotic Stow: A Diffusion-Based World Model from Sparse Snapshots
by: Zhang, Lijun, et al.
Published: (2026)
by: Zhang, Lijun, et al.
Published: (2026)
CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs
by: Zhang, Jiaming, et al.
Published: (2025)
by: Zhang, Jiaming, et al.
Published: (2025)
Mantis: A Versatile Vision-Language-Action Model with Disentangled Visual Foresight
by: Yang, Yi, et al.
Published: (2025)
by: Yang, Yi, et al.
Published: (2025)
SODA: Out-of-Distribution Detection in Domain-Shifted Point Clouds via Neighborhood Propagation
by: Goodge, Adam, et al.
Published: (2025)
by: Goodge, Adam, et al.
Published: (2025)
Think Before You Drive: World Model-Inspired Multimodal Grounding for Autonomous Vehicles
by: Liao, Haicheng, et al.
Published: (2025)
by: Liao, Haicheng, et al.
Published: (2025)
CrossView Suite: Harnessing Cross-view Spatial Intelligence of MLLMs with Dataset, Model and Benchmark
by: Wang, Wei, et al.
Published: (2026)
by: Wang, Wei, et al.
Published: (2026)
SUPERGLASSES: Benchmarking Vision Language Models as Intelligent Agents for AI Smart Glasses
by: Jiang, Zhuohang, et al.
Published: (2026)
by: Jiang, Zhuohang, et al.
Published: (2026)
Discrete Diffusion Models with MLLMs for Unified Medical Multimodal Generation
by: Mao, Jiawei, et al.
Published: (2025)
by: Mao, Jiawei, et al.
Published: (2025)
Dense Connector for MLLMs
by: Yao, Huanjin, et al.
Published: (2024)
by: Yao, Huanjin, et al.
Published: (2024)
Future-Aware Interaction Network For Motion Forecasting
by: Li, Shijie, et al.
Published: (2025)
by: Li, Shijie, et al.
Published: (2025)
VAP-Diffusion: Enriching Descriptions with MLLMs for Enhanced Medical Image Generation
by: Huang, Peng, et al.
Published: (2025)
by: Huang, Peng, et al.
Published: (2025)
Finding Lottery Tickets in Vision Models via Data-driven Spectral Foresight Pruning
by: Iurada, Leonardo, et al.
Published: (2024)
by: Iurada, Leonardo, et al.
Published: (2024)
V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators
by: Zhou, Jiazhou, et al.
Published: (2026)
by: Zhou, Jiazhou, et al.
Published: (2026)
CrystaL: Spontaneous Emergence of Visual Latents in MLLMs
by: Zhang, Yang, et al.
Published: (2026)
by: Zhang, Yang, et al.
Published: (2026)
Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors
by: Zheng, Duo, et al.
Published: (2025)
by: Zheng, Duo, et al.
Published: (2025)
CLIP-based Camera-Agnostic Feature Learning for Intra-camera Person Re-Identification
by: Tan, Xuan, et al.
Published: (2024)
by: Tan, Xuan, et al.
Published: (2024)
Enhancing Spatial Reasoning through Visual and Textual Thinking
by: Liang, Xun, et al.
Published: (2025)
by: Liang, Xun, et al.
Published: (2025)
A Benchmark for Ultra-High-Resolution Remote Sensing MLLMs
by: Dang, Yunkai, et al.
Published: (2025)
by: Dang, Yunkai, et al.
Published: (2025)
From Attributes to Natural Language: A Survey and Foresight on Text-based Person Re-identification
by: Jiang, Fanzhi, et al.
Published: (2024)
by: Jiang, Fanzhi, et al.
Published: (2024)
VidComposition: Can MLLMs Analyze Compositions in Compiled Videos?
by: Tang, Yolo Y., et al.
Published: (2024)
by: Tang, Yolo Y., et al.
Published: (2024)
Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events
by: Liu, Xiaolin, et al.
Published: (2026)
by: Liu, Xiaolin, et al.
Published: (2026)
Multi-View Industrial Anomaly Detection with Epipolar Constrained Cross-View Fusion
by: Liu, Yifan, et al.
Published: (2025)
by: Liu, Yifan, et al.
Published: (2025)
VRAG: Learning World Models for Interactive Video Generation
by: Chen, Taiye, et al.
Published: (2025)
by: Chen, Taiye, et al.
Published: (2025)
Non-Markov Multi-Round Conversational Image Generation with History-Conditioned MLLMs
by: Zhang, Haochen, et al.
Published: (2026)
by: Zhang, Haochen, et al.
Published: (2026)
PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension
by: Ouyang, Kun, et al.
Published: (2024)
by: Ouyang, Kun, et al.
Published: (2024)
SCBench: A Sports Commentary Benchmark for Video LLMs
by: Ge, Kuangzhi, et al.
Published: (2024)
by: Ge, Kuangzhi, et al.
Published: (2024)
VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos
by: Liu, Wenqi, et al.
Published: (2026)
by: Liu, Wenqi, et al.
Published: (2026)
STELAR-VISION: Self-Topology-Aware Efficient Learning for Aligned Reasoning in Vision
by: Li, Chen, et al.
Published: (2025)
by: Li, Chen, et al.
Published: (2025)
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models
by: Zhan, Yufei, et al.
Published: (2025)
by: Zhan, Yufei, et al.
Published: (2025)
GranViT: A Fine-Grained Vision Model With Autoregressive Perception For MLLMs
by: Zheng, Guanghao, et al.
Published: (2025)
by: Zheng, Guanghao, et al.
Published: (2025)
MedUnifier: Unifying Vision-and-Language Pre-training on Medical Data with Vision Generation Task using Discrete Visual Representations
by: Zhang, Ziyang, et al.
Published: (2025)
by: Zhang, Ziyang, et al.
Published: (2025)
ORL-LDM: Offline Reinforcement Learning Guided Latent Diffusion Model Super-Resolution Reconstruction
by: Lyu, Shijie
Published: (2025)
by: Lyu, Shijie
Published: (2025)
Towards Unified Surgical Scene Understanding:Bridging Reasoning and Grounding via MLLMs
by: Huang, Jincai, et al.
Published: (2026)
by: Huang, Jincai, et al.
Published: (2026)
Similar Items
-
Ego2World: Compiling Egocentric Cooking Videos into Executable Worlds for Belief-State Planning
by: Cheng, Qinchuan, et al.
Published: (2026) -
GRIT: Teaching MLLMs to Think with Images
by: Fan, Yue, et al.
Published: (2025) -
ETA: Efficiency through Thinking Ahead, A Dual Approach to Self-Driving with Large Models
by: Hamdan, Shadi, et al.
Published: (2025) -
CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving
by: Huang, Minqing, et al.
Published: (2026) -
Chain of World: World Model Thinking in Latent Motion
by: Yang, Fuxiang, et al.
Published: (2026)