Visualizing Thought: Conceptual Diagrams Enable Robust Planning in LMMs
Fuente:
arXiv
Saved in:
| Main Authors: | Borazjanizadeh, Nasim, Herzig, Roei, Oks, Eduard, Darrell, Trevor, Feris, Rogerio, Karlinsky, Leonid |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Navigating the Labyrinth: Evaluating LLMs' Ability to Reason About Search Problems
by: Borazjanizadeh, Nasim, et al.
Published: (2024)
by: Borazjanizadeh, Nasim, et al.
Published: (2024)
Latent Implicit Visual Reasoning
by: Li, Kelvin, et al.
Published: (2025)
by: Li, Kelvin, et al.
Published: (2025)
Multimodal Task Vectors Enable Many-Shot Multimodal In-Context Learning
by: Huang, Brandon, et al.
Published: (2024)
by: Huang, Brandon, et al.
Published: (2024)
Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features
by: Mitra, Chancharik, et al.
Published: (2024)
by: Mitra, Chancharik, et al.
Published: (2024)
Compositional Chain-of-Thought Prompting for Large Multimodal Models
by: Mitra, Chancharik, et al.
Published: (2023)
by: Mitra, Chancharik, et al.
Published: (2023)
Modeling Language as a Sequence of Thoughts
by: Borazjanizadeh, Nasim, et al.
Published: (2025)
by: Borazjanizadeh, Nasim, et al.
Published: (2025)
Reliable Reasoning Beyond Natural Language
by: Borazjanizadeh, Nasim, et al.
Published: (2024)
by: Borazjanizadeh, Nasim, et al.
Published: (2024)
TraveLER: A Modular Multi-LMM Agent Framework for Video Question-Answering
by: Shang, Chuyi, et al.
Published: (2024)
by: Shang, Chuyi, et al.
Published: (2024)
State-Space Large Audio Language Models
by: Bhati, Saurabhchand, et al.
Published: (2024)
by: Bhati, Saurabhchand, et al.
Published: (2024)
DAVE: A VLM Vision Encoder for Document Understanding and Web Agents
by: Huang, Brandon, et al.
Published: (2025)
by: Huang, Brandon, et al.
Published: (2025)
Do What? Teaching Vision-Language-Action Models to Reject the Impossible
by: Hsieh, Wen-Han, et al.
Published: (2025)
by: Hsieh, Wen-Han, et al.
Published: (2025)
$\textit{Trans-LoRA}$: towards data-free Transferable Parameter Efficient Finetuning
by: Wang, Runqian, et al.
Published: (2024)
by: Wang, Runqian, et al.
Published: (2024)
Pre-training Auto-regressive Robotic Models with 4D Representations
by: Niu, Dantong, et al.
Published: (2025)
by: Niu, Dantong, et al.
Published: (2025)
Large Scale Generative AI Text Applied to Sports and Music
by: Baughman, Aaron, et al.
Published: (2024)
by: Baughman, Aaron, et al.
Published: (2024)
Activation Reward Models for Few-Shot Model Alignment
by: Chai, Tianning, et al.
Published: (2025)
by: Chai, Tianning, et al.
Published: (2025)
Self-Specialization: Uncovering Latent Expertise within Large Language Models
by: Kang, Junmo, et al.
Published: (2023)
by: Kang, Junmo, et al.
Published: (2023)
Recursive Visual Programming
by: Ge, Jiaxin, et al.
Published: (2023)
by: Ge, Jiaxin, et al.
Published: (2023)
Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
by: Qin, Yiming, et al.
Published: (2025)
by: Qin, Yiming, et al.
Published: (2025)
TULIP: Towards Unified Language-Image Pretraining
by: Tang, Zineng, et al.
Published: (2025)
by: Tang, Zineng, et al.
Published: (2025)
On the Diagram of Thought
by: Zhang, Yifan, et al.
Published: (2024)
by: Zhang, Yifan, et al.
Published: (2024)
In-Context Learning Enables Robot Action Prediction in LLMs
by: Yin, Yida, et al.
Published: (2024)
by: Yin, Yida, et al.
Published: (2024)
UnSAMv2: Self-Supervised Learning Enables Segment Anything at Any Granularity
by: Yu, Junwei, et al.
Published: (2025)
by: Yu, Junwei, et al.
Published: (2025)
ConMe: Rethinking Evaluation of Compositional Reasoning for Modern VLMs
by: Huang, Irene, et al.
Published: (2024)
by: Huang, Irene, et al.
Published: (2024)
Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation
by: Rouditchenko, Andrew, et al.
Published: (2024)
by: Rouditchenko, Andrew, et al.
Published: (2024)
$\texttt{BATCLIP}$: Bimodal Online Test-Time Adaptation for CLIP
by: Maharana, Sarthak Kumar, et al.
Published: (2024)
by: Maharana, Sarthak Kumar, et al.
Published: (2024)
Lifting Embodied World Models for Planning and Control
by: Wang, Alex N., et al.
Published: (2026)
by: Wang, Alex N., et al.
Published: (2026)
Draw with Thought: Unleashing Multimodal Reasoning for Scientific Diagram Generation
by: Cui, Zhiqing, et al.
Published: (2025)
by: Cui, Zhiqing, et al.
Published: (2025)
MANAR: Memory-augmented Attention with Navigational Abstract Conceptual Representation
by: Jahshan, Zuher, et al.
Published: (2026)
by: Jahshan, Zuher, et al.
Published: (2026)
Comparison Visual Instruction Tuning
by: Lin, Wei, et al.
Published: (2024)
by: Lin, Wei, et al.
Published: (2024)
VITAL: Vision-Encoder-centered Pre-training for LMMs in Visual Quality Assessment
by: Jia, Ziheng, et al.
Published: (2025)
by: Jia, Ziheng, et al.
Published: (2025)
OmniDiagram: Advancing Unified Diagram Code Generation via Visual Interrogation Reward
by: Yang, Haoyue, et al.
Published: (2026)
by: Yang, Haoyue, et al.
Published: (2026)
Towards Audio Token Compression in Large Audio Language Models
by: Bhati, Saurabhchand, et al.
Published: (2025)
by: Bhati, Saurabhchand, et al.
Published: (2025)
End-to-End Breast Cancer Radiotherapy Planning via LMMs with Consistency Embedding
by: Kim, Kwanyoung, et al.
Published: (2023)
by: Kim, Kwanyoung, et al.
Published: (2023)
Visual Lexicon: Rich Image Features in Language Space
by: Wang, XuDong, et al.
Published: (2024)
by: Wang, XuDong, et al.
Published: (2024)
MIBench: Evaluating LMMs on Multimodal Interaction
by: Miao, Yu, et al.
Published: (2026)
by: Miao, Yu, et al.
Published: (2026)
Meta-Prompting for Automating Zero-shot Visual Recognition with LLMs
by: Mirza, M. Jehanzeb, et al.
Published: (2024)
by: Mirza, M. Jehanzeb, et al.
Published: (2024)
ChartGen: Scaling Chart Understanding Via Code-Guided Synthetic Chart Generation
by: Kondic, Jovana, et al.
Published: (2025)
by: Kondic, Jovana, et al.
Published: (2025)
CALM: Class-Conditional Sparse Attention Vectors for Large Audio-Language Models
by: Mehta, Videet, et al.
Published: (2026)
by: Mehta, Videet, et al.
Published: (2026)
DASS: Distilled Audio State Space Models Are Stronger and More Duration-Scalable Learners
by: Bhati, Saurabhchand, et al.
Published: (2024)
by: Bhati, Saurabhchand, et al.
Published: (2024)
Planning in the Dark: LLM-Symbolic Planning Pipeline without Experts
by: Huang, Sukai, et al.
Published: (2024)
by: Huang, Sukai, et al.
Published: (2024)
Similar Items
-
Navigating the Labyrinth: Evaluating LLMs' Ability to Reason About Search Problems
by: Borazjanizadeh, Nasim, et al.
Published: (2024) -
Latent Implicit Visual Reasoning
by: Li, Kelvin, et al.
Published: (2025) -
Multimodal Task Vectors Enable Many-Shot Multimodal In-Context Learning
by: Huang, Brandon, et al.
Published: (2024) -
Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features
by: Mitra, Chancharik, et al.
Published: (2024) -
Compositional Chain-of-Thought Prompting for Large Multimodal Models
by: Mitra, Chancharik, et al.
Published: (2023)