Visual Planning: Let's Think Only with Images
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Yi, Li, Chengzu, Zhou, Han, Wan, Xingchen, Zhang, Caiqi, Korhonen, Anna, Vulić, Ivan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Thinking in Frames: How Visual Context and Test-Time Scaling Empower Video Reasoning
by: Li, Chengzu, et al.
Published: (2026)
by: Li, Chengzu, et al.
Published: (2026)
TopViewRS: Vision-Language Models as Top-View Spatial Reasoners
by: Li, Chengzu, et al.
Published: (2024)
by: Li, Chengzu, et al.
Published: (2024)
Translation-Enhanced Multilingual Text-to-Image Generation
by: Li, Yaoyiran, et al.
Published: (2023)
by: Li, Yaoyiran, et al.
Published: (2023)
Lost in Embeddings: Information Loss in Vision-Language Models
by: Li, Wenyan, et al.
Published: (2025)
by: Li, Wenyan, et al.
Published: (2025)
Agentic Policy Optimization via Instruction-Policy Co-Evolution
by: Zhou, Han, et al.
Published: (2025)
by: Zhou, Han, et al.
Published: (2025)
AutoPEFT: Automatic Configuration Search for Parameter-Efficient Fine-Tuning
by: Zhou, Han, et al.
Published: (2023)
by: Zhou, Han, et al.
Published: (2023)
Large Language Models are Miscalibrated In-Context Learners
by: Li, Chengzu, et al.
Published: (2023)
by: Li, Chengzu, et al.
Published: (2023)
Chatting with Images for Introspective Visual Thinking
by: Wu, Junfei, et al.
Published: (2026)
by: Wu, Junfei, et al.
Published: (2026)
Re-Thinking the Automatic Evaluation of Image-Text Alignment in Text-to-Image Models
by: Zhang, Huixuan, et al.
Published: (2025)
by: Zhang, Huixuan, et al.
Published: (2025)
Let's Think Outside the Box: Exploring Leap-of-Thought in Large Language Models with Creative Humor Generation
by: Zhong, Shanshan, et al.
Published: (2023)
by: Zhong, Shanshan, et al.
Published: (2023)
Bootstrapping Action-Grounded Visual Dynamics in Unified Vision-Language Models
by: Qiu, Yifu, et al.
Published: (2025)
by: Qiu, Yifu, et al.
Published: (2025)
Fairer Preferences Elicit Improved Human-Aligned Large Language Model Judgments
by: Zhou, Han, et al.
Published: (2024)
by: Zhou, Han, et al.
Published: (2024)
Semantic Map-based Generation of Navigation Instructions
by: Li, Chengzu, et al.
Published: (2024)
by: Li, Chengzu, et al.
Published: (2024)
Thinking with Generated Images
by: Chern, Ethan, et al.
Published: (2025)
by: Chern, Ethan, et al.
Published: (2025)
Can Rule-Based Insights Enhance LLMs for Radiology Report Classification? Introducing the RadPrompt Methodology
by: Fytas, Panagiotis, et al.
Published: (2024)
by: Fytas, Panagiotis, et al.
Published: (2024)
Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
Let's Think with Images Efficiently! An Interleaved-Modal Chain-of-Thought Reasoning Framework with Dynamic and Precise Visual Thoughts
by: Liu, Xu, et al.
Published: (2026)
by: Liu, Xu, et al.
Published: (2026)
GRIT: Teaching MLLMs to Think with Images
by: Fan, Yue, et al.
Published: (2025)
by: Fan, Yue, et al.
Published: (2025)
RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction
by: Wang, Yuchi, et al.
Published: (2025)
by: Wang, Yuchi, et al.
Published: (2025)
Thinking with Patterns: Breaking the Perceptual Bottleneck in Visual Planning via Pattern Induction
by: Jian, Yichang, et al.
Published: (2026)
by: Jian, Yichang, et al.
Published: (2026)
Think Visually, Reason Textually: Vision-Language Synergy in ARC
by: Zhang, Beichen, et al.
Published: (2025)
by: Zhang, Beichen, et al.
Published: (2025)
Imagine while Reasoning in Space: Multimodal Visualization-of-Thought
by: Li, Chengzu, et al.
Published: (2025)
by: Li, Chengzu, et al.
Published: (2025)
Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation
by: Yu, Seonghoon, et al.
Published: (2026)
by: Yu, Seonghoon, et al.
Published: (2026)
TOPA: Extending Large Language Models for Video Understanding via Text-Only Pre-Alignment
by: Li, Wei, et al.
Published: (2024)
by: Li, Wei, et al.
Published: (2024)
Seeing but Not Thinking: Routing Distraction in Multimodal Mixture-of-Experts
by: Xu, Haolei, et al.
Published: (2026)
by: Xu, Haolei, et al.
Published: (2026)
Seeing the Image: Prioritizing Visual Correlation by Contrastive Alignment
by: Xiao, Xin, et al.
Published: (2024)
by: Xiao, Xin, et al.
Published: (2024)
Image-of-Thought Prompting for Visual Reasoning Refinement in Multimodal Large Language Models
by: Zhou, Qiji, et al.
Published: (2024)
by: Zhou, Qiji, et al.
Published: (2024)
Think While You Generate: Discrete Diffusion with Planned Denoising
by: Liu, Sulin, et al.
Published: (2024)
by: Liu, Sulin, et al.
Published: (2024)
VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI
by: Cheng, Sijie, et al.
Published: (2024)
by: Cheng, Sijie, et al.
Published: (2024)
VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images
by: Zhou, Guanyu, et al.
Published: (2026)
by: Zhou, Guanyu, et al.
Published: (2026)
More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models
by: Liu, Chengzhi, et al.
Published: (2025)
by: Liu, Chengzhi, et al.
Published: (2025)
Decoder-Only LLMs are Better Controllers for Diffusion Models
by: Dong, Ziyi, et al.
Published: (2025)
by: Dong, Ziyi, et al.
Published: (2025)
Benchmarking and Enhancing Text-to-Image Models for Generating Visual Representations in Early Arithmetic Education
by: Wang, Junling, et al.
Published: (2026)
by: Wang, Junling, et al.
Published: (2026)
Automatic Layout Planning for Visually-Rich Documents with Instruction-Following Models
by: Zhu, Wanrong, et al.
Published: (2024)
by: Zhu, Wanrong, et al.
Published: (2024)
POP: Prefill-Only Pruning for Efficient Large Model Inference
by: He, Junhui, et al.
Published: (2026)
by: He, Junhui, et al.
Published: (2026)
Image Matters: A New Dataset and Empirical Study for Multimodal Hyperbole Detection
by: Zhang, Huixuan, et al.
Published: (2023)
by: Zhang, Huixuan, et al.
Published: (2023)
VideoExplorer: Think With Videos For Agentic Long-Video Understanding
by: Yuan, Huaying, et al.
Published: (2025)
by: Yuan, Huaying, et al.
Published: (2025)
MindGYM: What Matters in Question Synthesis for Thinking-Centric Fine-Tuning?
by: Xu, Zhe, et al.
Published: (2025)
by: Xu, Zhe, et al.
Published: (2025)
Why Only Text: Empowering Vision-and-Language Navigation with Multi-modal Prompts
by: Hong, Haodong, et al.
Published: (2024)
by: Hong, Haodong, et al.
Published: (2024)
Similar Items
-
Thinking in Frames: How Visual Context and Test-Time Scaling Empower Video Reasoning
by: Li, Chengzu, et al.
Published: (2026) -
TopViewRS: Vision-Language Models as Top-View Spatial Reasoners
by: Li, Chengzu, et al.
Published: (2024) -
Translation-Enhanced Multilingual Text-to-Image Generation
by: Li, Yaoyiran, et al.
Published: (2023) -
Lost in Embeddings: Information Loss in Vision-Language Models
by: Li, Wenyan, et al.
Published: (2025) -
Agentic Policy Optimization via Instruction-Policy Co-Evolution
by: Zhou, Han, et al.
Published: (2025)