See It, Say It, Sorted: An Iterative Training-Free Framework for Visually-Grounded Multimodal Reasoning in LVLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Yongchang, Ma, Oliver, Liu, Tianyi, Zhou, Guangquan, Chen, Yang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
See it. Say it. Sorted: Agentic System for Compositional Diagram Generation
von: Zhang, Hantao, et al.
Veröffentlicht: (2025)
von: Zhang, Hantao, et al.
Veröffentlicht: (2025)
See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMs
von: Dai, Ziyun, et al.
Veröffentlicht: (2025)
von: Dai, Ziyun, et al.
Veröffentlicht: (2025)
DeepScan: A Training-Free Framework for Visually Grounded Reasoning in Large Vision-Language Models
von: Li, Yangfu, et al.
Veröffentlicht: (2026)
von: Li, Yangfu, et al.
Veröffentlicht: (2026)
Think Twice to See More: Iterative Visual Reasoning in Medical VLMs
von: Chen, Kaitao, et al.
Veröffentlicht: (2025)
von: Chen, Kaitao, et al.
Veröffentlicht: (2025)
Connecting the Dots: Training-Free Visual Grounding via Agentic Reasoning
von: Luo, Liqin, et al.
Veröffentlicht: (2025)
von: Luo, Liqin, et al.
Veröffentlicht: (2025)
Visual Description Grounding Reduces Hallucinations and Boosts Reasoning in LVLMs
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination
von: Shi, Chufan, et al.
Veröffentlicht: (2026)
von: Shi, Chufan, et al.
Veröffentlicht: (2026)
Paying More Attention to Image: A Training-Free Method for Alleviating Hallucination in LVLMs
von: Liu, Shi, et al.
Veröffentlicht: (2024)
von: Liu, Shi, et al.
Veröffentlicht: (2024)
ClipTTT: CLIP-Guided Test-Time Training Helps LVLMs See Better
von: Nath, Mriganka, et al.
Veröffentlicht: (2026)
von: Nath, Mriganka, et al.
Veröffentlicht: (2026)
DIVER: Dynamic Iterative Visual Evidence Reasoning for Multimodal Fake News Detection
von: Zhou, Weilin, et al.
Veröffentlicht: (2026)
von: Zhou, Weilin, et al.
Veröffentlicht: (2026)
Elevating Visual Question Answering through Implicitly Learned Reasoning Pathways in LVLMs
von: Jing, Liu, et al.
Veröffentlicht: (2025)
von: Jing, Liu, et al.
Veröffentlicht: (2025)
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding
von: Li, Rong, et al.
Veröffentlicht: (2024)
von: Li, Rong, et al.
Veröffentlicht: (2024)
MIRROR: Multimodal Iterative Reasoning via Reflection on Visual Regions
von: Zhang, Haoyu, et al.
Veröffentlicht: (2026)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2026)
See Less, See Right: Bi-directional Perceptual Shaping For Multimodal Reasoning
von: Zhang, Shuoshuo, et al.
Veröffentlicht: (2025)
von: Zhang, Shuoshuo, et al.
Veröffentlicht: (2025)
Self-Improving Small Object Grounding in LVLMs
von: Yang, Tianze, et al.
Veröffentlicht: (2026)
von: Yang, Tianze, et al.
Veröffentlicht: (2026)
Boosting Visual Knowledge-Intensive Training for LVLMs Through Causality-Driven Visual Object Completion
von: Hu, Qingguo, et al.
Veröffentlicht: (2025)
von: Hu, Qingguo, et al.
Veröffentlicht: (2025)
Seeing Culture: A Benchmark for Visual Reasoning and Grounding
von: Satar, Burak, et al.
Veröffentlicht: (2025)
von: Satar, Burak, et al.
Veröffentlicht: (2025)
IVC-Prune: Revealing the Implicit Visual Coordinates in LVLMs for Vision Token Pruning
von: Sun, Zhichao, et al.
Veröffentlicht: (2026)
von: Sun, Zhichao, et al.
Veröffentlicht: (2026)
Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation
von: Ying, Kaining, et al.
Veröffentlicht: (2025)
von: Ying, Kaining, et al.
Veröffentlicht: (2025)
Language Models Can See Better: Visual Contrastive Decoding For LLM Multimodal Reasoning
von: Pang, Yuqi, et al.
Veröffentlicht: (2025)
von: Pang, Yuqi, et al.
Veröffentlicht: (2025)
Unlocking Multilingual Reasoning Capability of LLMs and LVLMs through Representation Engineering
von: Li, Qiming, et al.
Veröffentlicht: (2025)
von: Li, Qiming, et al.
Veröffentlicht: (2025)
Language-to-Space Programming for Training-Free 3D Visual Grounding
von: Mi, Boyu, et al.
Veröffentlicht: (2025)
von: Mi, Boyu, et al.
Veröffentlicht: (2025)
Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning
von: Ni, Minheng, et al.
Veröffentlicht: (2025)
von: Ni, Minheng, et al.
Veröffentlicht: (2025)
Self-Prophetic Decoding to Unlock Visual Search in LVLMs
von: He, Zhendong, et al.
Veröffentlicht: (2026)
von: He, Zhendong, et al.
Veröffentlicht: (2026)
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning
von: Wu, Chengfei, et al.
Veröffentlicht: (2025)
von: Wu, Chengfei, et al.
Veröffentlicht: (2025)
Seeing Clearly by Layer Two: Enhancing Attention Heads to Alleviate Hallucination in LVLMs
von: Zhang, Xiaofeng, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaofeng, et al.
Veröffentlicht: (2024)
REVEAL: Reference-Grounded Reasoning for Multimodal Manipulation Detection
von: Zhou, Jun, et al.
Veröffentlicht: (2026)
von: Zhou, Jun, et al.
Veröffentlicht: (2026)
Causally-Grounded Dual-Path Attention Intervention for Object Hallucination Mitigation in LVLMs
von: Yu, Liu, et al.
Veröffentlicht: (2025)
von: Yu, Liu, et al.
Veröffentlicht: (2025)
Knowing the Answer Isn't Enough: Fixing Reasoning Path Failures in LVLMs
von: Wang, Chaoyang, et al.
Veröffentlicht: (2025)
von: Wang, Chaoyang, et al.
Veröffentlicht: (2025)
See&Trek: Training-Free Spatial Prompting for Multimodal Large Language Model
von: Li, Pengteng, et al.
Veröffentlicht: (2025)
von: Li, Pengteng, et al.
Veröffentlicht: (2025)
Beyond Accuracy: Evaluating Visual Grounding In Multimodal Medical Reasoning
von: Zafar, Anas, et al.
Veröffentlicht: (2026)
von: Zafar, Anas, et al.
Veröffentlicht: (2026)
Seeing Together: Multi-Robot Cooperative Egocentric Spatial Reasoning with Multimodal Large Language Models
von: Peng, Kunyu, et al.
Veröffentlicht: (2026)
von: Peng, Kunyu, et al.
Veröffentlicht: (2026)
Mitigating Visual Context Degradation in Large Multimodal Models: A Training-Free Decoupled Agentic Framework
von: Jia, Hongrui, et al.
Veröffentlicht: (2025)
von: Jia, Hongrui, et al.
Veröffentlicht: (2025)
Seeing to Ground: Visual Attention for Hallucination-Resilient MDLLMs
von: Narnaware, Vishal, et al.
Veröffentlicht: (2026)
von: Narnaware, Vishal, et al.
Veröffentlicht: (2026)
Do You See What I Say? Generalizable Deepfake Detection based on Visual Speech Recognition
von: Bora, Maheswar, et al.
Veröffentlicht: (2025)
von: Bora, Maheswar, et al.
Veröffentlicht: (2025)
Context-Aware Multi-Turn Visual-Textual Reasoning in LVLMs via Dynamic Memory and Adaptive Visual Guidance
von: Shen, Weijie, et al.
Veröffentlicht: (2025)
von: Shen, Weijie, et al.
Veröffentlicht: (2025)
SkyNative: A Native Multimodal Framework for Remote Sensing Visual Evidence Reasoning
von: Yang, Xiao, et al.
Veröffentlicht: (2026)
von: Yang, Xiao, et al.
Veröffentlicht: (2026)
UniGround: Universal 3D Visual Grounding via Training-Free Scene Parsing
von: Zhang, Jiaxi, et al.
Veröffentlicht: (2026)
von: Zhang, Jiaxi, et al.
Veröffentlicht: (2026)
RASA: Replace Anyone, Say Anything -- A Training-Free Framework for Audio-Driven and Universal Portrait Video Editing
von: Pan, Tianrui, et al.
Veröffentlicht: (2025)
von: Pan, Tianrui, et al.
Veröffentlicht: (2025)
Seeing Right but Saying Wrong: Inter- and Intra-Layer Refinement in MLLMs without Training
von: Song, Shezheng, et al.
Veröffentlicht: (2026)
von: Song, Shezheng, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
See it. Say it. Sorted: Agentic System for Compositional Diagram Generation
von: Zhang, Hantao, et al.
Veröffentlicht: (2025) -
See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMs
von: Dai, Ziyun, et al.
Veröffentlicht: (2025) -
DeepScan: A Training-Free Framework for Visually Grounded Reasoning in Large Vision-Language Models
von: Li, Yangfu, et al.
Veröffentlicht: (2026) -
Think Twice to See More: Iterative Visual Reasoning in Medical VLMs
von: Chen, Kaitao, et al.
Veröffentlicht: (2025) -
Connecting the Dots: Training-Free Visual Grounding via Agentic Reasoning
von: Luo, Liqin, et al.
Veröffentlicht: (2025)