MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Jiarui, Khayatkhoei, Mahyar, Chhikara, Prateek, Ilievski, Filip |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Towards Perceiving Small Visual Details in Zero-shot Visual Question Answering with Multimodal LLMs
di: Zhang, Jiarui, et al.
Pubblicazione: (2023)
di: Zhang, Jiarui, et al.
Pubblicazione: (2023)
Exploring Perceptual Limitation of Multimodal Large Language Models
di: Zhang, Jiarui, et al.
Pubblicazione: (2024)
di: Zhang, Jiarui, et al.
Pubblicazione: (2024)
FIRE: Food Image to REcipe generation
di: Chhikara, Prateek, et al.
Pubblicazione: (2023)
di: Chhikara, Prateek, et al.
Pubblicazione: (2023)
Euclid: Supercharging Multimodal LLMs with Synthetic High-Fidelity Visual Descriptions
di: Zhang, Jiarui, et al.
Pubblicazione: (2024)
di: Zhang, Jiarui, et al.
Pubblicazione: (2024)
LookWise: Knowing When and Where to Look for Fine-Grained Visual Reasoning in Multimodal Large Language Models
di: Shen, Yuxiang, et al.
Pubblicazione: (2026)
di: Shen, Yuxiang, et al.
Pubblicazione: (2026)
Look, Learn and Leverage (L$^3$): Mitigating Visual-Domain Shift and Discovering Intrinsic Relations via Symbolic Alignment
di: Xie, Hanchen, et al.
Pubblicazione: (2024)
di: Xie, Hanchen, et al.
Pubblicazione: (2024)
Look Twice: Training-Free Evidence Highlighting in Multimodal Large Language Models
di: Morini, Marco, et al.
Pubblicazione: (2026)
di: Morini, Marco, et al.
Pubblicazione: (2026)
Cube Bench: A Benchmark for Spatial Visual Reasoning in MLLMs
di: Anand, Dhruv, et al.
Pubblicazione: (2025)
di: Anand, Dhruv, et al.
Pubblicazione: (2025)
AdaCodec: A Predictive Visual Code for Video MLLMs
di: Hou, Haowen, et al.
Pubblicazione: (2026)
di: Hou, Haowen, et al.
Pubblicazione: (2026)
Unsupervised Multimodal Deepfake Detection Using Intra- and Cross-Modal Inconsistencies
di: Tian, Mulin, et al.
Pubblicazione: (2023)
di: Tian, Mulin, et al.
Pubblicazione: (2023)
Beyond Words: Multimodal LLM Knows When to Speak
di: Liao, Zikai, et al.
Pubblicazione: (2025)
di: Liao, Zikai, et al.
Pubblicazione: (2025)
Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs
di: Zhang, Huanyu, et al.
Pubblicazione: (2025)
di: Zhang, Huanyu, et al.
Pubblicazione: (2025)
Do MLLMs Really See It: Reinforcing Visual Attention in Multimodal LLMs
di: Ou, Siqu, et al.
Pubblicazione: (2026)
di: Ou, Siqu, et al.
Pubblicazione: (2026)
Linking Perception, Confidence and Accuracy in MLLMs
di: Du, Yuetian, et al.
Pubblicazione: (2026)
di: Du, Yuetian, et al.
Pubblicazione: (2026)
Mitigating Hallucinations in Multimodal Spatial Relations through Constraint-Aware Prompting
di: Wu, Jiarui, et al.
Pubblicazione: (2025)
di: Wu, Jiarui, et al.
Pubblicazione: (2025)
Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation
di: Yuan, Qianhao, et al.
Pubblicazione: (2026)
di: Yuan, Qianhao, et al.
Pubblicazione: (2026)
Multimodal LLMs Struggle with Basic Visual Network Analysis: a VNA Benchmark
di: Williams, Evan M., et al.
Pubblicazione: (2024)
di: Williams, Evan M., et al.
Pubblicazione: (2024)
MLLMs-Augmented Visual-Language Representation Learning
di: Liu, Yanqing, et al.
Pubblicazione: (2023)
di: Liu, Yanqing, et al.
Pubblicazione: (2023)
Training-free Guidance in Text-to-Video Generation via Multimodal Planning and Structured Noise Initialization
di: Li, Jialu, et al.
Pubblicazione: (2025)
di: Li, Jialu, et al.
Pubblicazione: (2025)
A Study of Commonsense Reasoning over Visual Object Properties
di: Kolari, Abhishek, et al.
Pubblicazione: (2025)
di: Kolari, Abhishek, et al.
Pubblicazione: (2025)
GridPrune: From "Where to Look" to "What to Select" in Visual Token Pruning for MLLMs
di: Duan, Yuxiang, et al.
Pubblicazione: (2025)
di: Duan, Yuxiang, et al.
Pubblicazione: (2025)
Advancing Object Detection in Transportation with Multimodal Large Language Models (MLLMs): A Comprehensive Review and Empirical Testing
di: Ashqar, Huthaifa I., et al.
Pubblicazione: (2024)
di: Ashqar, Huthaifa I., et al.
Pubblicazione: (2024)
PoSh: Using Scene Graphs To Guide LLMs-as-a-Judge For Detailed Image Descriptions
di: Ananthram, Amith, et al.
Pubblicazione: (2025)
di: Ananthram, Amith, et al.
Pubblicazione: (2025)
Cross-Modal Projection in Multimodal LLMs Doesn't Really Project Visual Attributes to Textual Space
di: Verma, Gaurav, et al.
Pubblicazione: (2024)
di: Verma, Gaurav, et al.
Pubblicazione: (2024)
When Big Models Train Small Ones: Label-Free Model Parity Alignment for Efficient Visual Question Answering using Small VLMs
di: Penamakuri, Abhirama Subramanyam, et al.
Pubblicazione: (2025)
di: Penamakuri, Abhirama Subramanyam, et al.
Pubblicazione: (2025)
When Looking Is Not Enough: Visual Attention Structure Reveals Hallucination in MLLMs
di: Cao, Fanpu, et al.
Pubblicazione: (2026)
di: Cao, Fanpu, et al.
Pubblicazione: (2026)
Sound and Complete Neurosymbolic Reasoning with LLM-Grounded Interpretations
di: Allen, Bradley P., et al.
Pubblicazione: (2025)
di: Allen, Bradley P., et al.
Pubblicazione: (2025)
Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers
di: Wen, Yuxin, et al.
Pubblicazione: (2024)
di: Wen, Yuxin, et al.
Pubblicazione: (2024)
GRIT: Teaching MLLMs to Think with Images
di: Fan, Yue, et al.
Pubblicazione: (2025)
di: Fan, Yue, et al.
Pubblicazione: (2025)
Free Video-LLM: Prompt-guided Visual Perception for Efficient Training-free Video LLMs
di: Han, Kai, et al.
Pubblicazione: (2024)
di: Han, Kai, et al.
Pubblicazione: (2024)
Mind the Confidence Gap: Overconfidence, Calibration, and Distractor Effects in Large Language Models
di: Chhikara, Prateek
Pubblicazione: (2025)
di: Chhikara, Prateek
Pubblicazione: (2025)
Tables as Texts or Images: Evaluating the Table Reasoning Ability of LLMs and MLLMs
di: Deng, Naihao, et al.
Pubblicazione: (2024)
di: Deng, Naihao, et al.
Pubblicazione: (2024)
Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities
di: Yang, Enneng, et al.
Pubblicazione: (2024)
di: Yang, Enneng, et al.
Pubblicazione: (2024)
How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding
di: Yu, Zhuoran, et al.
Pubblicazione: (2025)
di: Yu, Zhuoran, et al.
Pubblicazione: (2025)
The Percept-V Challenge: Can Multimodal LLMs Crack Simple Perception Problems?
di: Ghosh, Samrajnee, et al.
Pubblicazione: (2025)
di: Ghosh, Samrajnee, et al.
Pubblicazione: (2025)
Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning
di: Liu, Huabin, et al.
Pubblicazione: (2025)
di: Liu, Huabin, et al.
Pubblicazione: (2025)
COLUMBUS: Evaluating COgnitive Lateral Understanding through Multiple-choice reBUSes
di: Kraaijveld, Koen, et al.
Pubblicazione: (2024)
di: Kraaijveld, Koen, et al.
Pubblicazione: (2024)
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering
di: Cocchi, Federico, et al.
Pubblicazione: (2024)
di: Cocchi, Federico, et al.
Pubblicazione: (2024)
Vision Enhancing LLMs: Empowering Multimodal Knowledge Storage and Sharing in LLMs
di: Li, Yunxin, et al.
Pubblicazione: (2023)
di: Li, Yunxin, et al.
Pubblicazione: (2023)
Perception-R1: Advancing Multimodal Reasoning Capabilities of MLLMs via Visual Perception Reward
di: Xiao, Tong, et al.
Pubblicazione: (2025)
di: Xiao, Tong, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Towards Perceiving Small Visual Details in Zero-shot Visual Question Answering with Multimodal LLMs
di: Zhang, Jiarui, et al.
Pubblicazione: (2023) -
Exploring Perceptual Limitation of Multimodal Large Language Models
di: Zhang, Jiarui, et al.
Pubblicazione: (2024) -
FIRE: Food Image to REcipe generation
di: Chhikara, Prateek, et al.
Pubblicazione: (2023) -
Euclid: Supercharging Multimodal LLMs with Synthetic High-Fidelity Visual Descriptions
di: Zhang, Jiarui, et al.
Pubblicazione: (2024) -
LookWise: Knowing When and Where to Look for Fine-Grained Visual Reasoning in Multimodal Large Language Models
di: Shen, Yuxiang, et al.
Pubblicazione: (2026)