Vision Verification Enhanced Fusion of VLMs for Efficient Visual Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Tekin, Selim Furkan, Xu, Yichang, Liu, Gaowen, Kompella, Ramana Rao, Loper, Margaret L., Liu, Ling |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Neurosymbolic Agent System for Compositional Visual Reasoning
by: Xu, Yichang, et al.
Published: (2025)
by: Xu, Yichang, et al.
Published: (2025)
Attention-aware Inference Optimizations for Large Vision-Language Models with Memory-efficient Decoding
by: Ilhan, Fatih, et al.
Published: (2026)
by: Ilhan, Fatih, et al.
Published: (2026)
A Multi-Agent Perception-Action Alliance for Efficient Long Video Reasoning
by: Xu, Yichang, et al.
Published: (2026)
by: Xu, Yichang, et al.
Published: (2026)
Adversarial Attention Perturbations for Large Object Detection Transformers
by: Yahn, Zachary, et al.
Published: (2025)
by: Yahn, Zachary, et al.
Published: (2025)
Efficient Multitask Dense Predictor via Binarization
by: Shang, Yuzhang, et al.
Published: (2024)
by: Shang, Yuzhang, et al.
Published: (2024)
Personalized Face Privacy Protection From a Single Image
by: Yahn, Zachary, et al.
Published: (2026)
by: Yahn, Zachary, et al.
Published: (2026)
Targeted Forgetting of Image Subgroups in CLIP Models
by: Zhang, Zeliang, et al.
Published: (2025)
by: Zhang, Zeliang, et al.
Published: (2025)
Robust Few-Shot Ensemble Learning with Focal Diversity-Based Pruning
by: Tekin, Selim Furkan, et al.
Published: (2024)
by: Tekin, Selim Furkan, et al.
Published: (2024)
Orientation-anchored Hyper-Gaussian for 4D Reconstruction from Casual Videos
by: Wu, Junyi, et al.
Published: (2025)
by: Wu, Junyi, et al.
Published: (2025)
Urban Scene Diffusion through Semantic Occupancy Map
by: Zhang, Junge, et al.
Published: (2024)
by: Zhang, Junge, et al.
Published: (2024)
Motion Marionette: Rethinking Rigid Motion Transfer via Prior Guidance
by: Wang, Haoxuan, et al.
Published: (2025)
by: Wang, Haoxuan, et al.
Published: (2025)
Dynamic Optimizations of LLM Ensembles with Two-Stage Reinforcement Learning Agents
by: Tekin, Selim Furkan, et al.
Published: (2025)
by: Tekin, Selim Furkan, et al.
Published: (2025)
SwiftNDC: Fast Neural Depth Correction for High-Fidelity 3D Reconstruction
by: Han, Kang, et al.
Published: (2026)
by: Han, Kang, et al.
Published: (2026)
GIFSplat: Generative Prior-Guided Iterative Feed-Forward 3D Gaussian Splatting from Sparse Views
by: Chen, Tianyu, et al.
Published: (2026)
by: Chen, Tianyu, et al.
Published: (2026)
Training-Free Semantic Segmentation via LLM-Supervision
by: Sun, Wenfang, et al.
Published: (2024)
by: Sun, Wenfang, et al.
Published: (2024)
UnlearnCanvas: Stylized Image Dataset for Enhanced Machine Unlearning Evaluation in Diffusion Models
by: Zhang, Yihua, et al.
Published: (2024)
by: Zhang, Yihua, et al.
Published: (2024)
Prompt Diffusion Robustifies Any-Modality Prompt Learning
by: Du, Yingjun, et al.
Published: (2024)
by: Du, Yingjun, et al.
Published: (2024)
Spectrally Distilled Representations Aligned with Instruction-Augmented LLMs for Satellite Imagery
by: Do, Minh Kha, et al.
Published: (2026)
by: Do, Minh Kha, et al.
Published: (2026)
A Survey on Large Language Model-Based Game Agents
by: Hu, Sihao, et al.
Published: (2024)
by: Hu, Sihao, et al.
Published: (2024)
Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning
by: Ge, Yuyao, et al.
Published: (2025)
by: Ge, Yuyao, et al.
Published: (2025)
VLMs have Tunnel Vision: Evaluating Nonlocal Visual Reasoning in Leading VLMs
by: Berman, Shmuel, et al.
Published: (2025)
by: Berman, Shmuel, et al.
Published: (2025)
Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs
by: Zhang, Xintong, et al.
Published: (2025)
by: Zhang, Xintong, et al.
Published: (2025)
ViFP: A Framework for Visual False Positive Detection to Enhance Reasoning Reliability in VLMs
by: Zhang, Ben, et al.
Published: (2025)
by: Zhang, Ben, et al.
Published: (2025)
SegVG: Transferring Object Bounding Box to Segmentation for Visual Grounding
by: Kang, Weitai, et al.
Published: (2024)
by: Kang, Weitai, et al.
Published: (2024)
Decomposing Visual Classification: Assessing Tree-Based Reasoning in VLMs
by: Elmansoury, Sary, et al.
Published: (2025)
by: Elmansoury, Sary, et al.
Published: (2025)
VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMs
by: Törtei, Brigitta Malagurski, et al.
Published: (2025)
by: Törtei, Brigitta Malagurski, et al.
Published: (2025)
VisionPulse: Dynamic Visual Sparsity for Efficient Multimodal Reasoning
by: Xu, Hengbo, et al.
Published: (2026)
by: Xu, Hengbo, et al.
Published: (2026)
How Auxiliary Reasoning Unleashes GUI Grounding in VLMs
by: Li, Weiming, et al.
Published: (2025)
by: Li, Weiming, et al.
Published: (2025)
Autonomous Imagination: Closed-Loop Decomposition of Visual-to-Textual Conversion in Visual Reasoning for Multimodal Large Language Models
by: Liu, Jingming, et al.
Published: (2024)
by: Liu, Jingming, et al.
Published: (2024)
Thinking with Gaze: Sequential Eye-Tracking as Visual Reasoning Supervision for Medical VLMs
by: Li, Yiwei, et al.
Published: (2026)
by: Li, Yiwei, et al.
Published: (2026)
Better Reasoning with Less Data: Enhancing VLMs Through Unified Modality Scoring
by: Xu, Mingjie, et al.
Published: (2025)
by: Xu, Mingjie, et al.
Published: (2025)
LightPure: Realtime Adversarial Image Purification for Mobile Devices Using Diffusion Models
by: Khalili, Hossein, et al.
Published: (2024)
by: Khalili, Hossein, et al.
Published: (2024)
Think Twice to See More: Iterative Visual Reasoning in Medical VLMs
by: Chen, Kaitao, et al.
Published: (2025)
by: Chen, Kaitao, et al.
Published: (2025)
Self-Adapting Large Visual-Language Models to Edge Devices across Visual Modalities
by: Cai, Kaiwen, et al.
Published: (2024)
by: Cai, Kaiwen, et al.
Published: (2024)
CLGRPO: Reasoning Ability Enhancement for Small VLMs
by: Wang, Fanyi, et al.
Published: (2025)
by: Wang, Fanyi, et al.
Published: (2025)
VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images
by: Zhou, Guanyu, et al.
Published: (2026)
by: Zhou, Guanyu, et al.
Published: (2026)
VRIQ: Benchmarking and Analyzing Visual-Reasoning IQ of VLMs
by: Khezresmaeilzadeh, Tina, et al.
Published: (2026)
by: Khezresmaeilzadeh, Tina, et al.
Published: (2026)
Enhancing Infrared Vision: Progressive Prompt Fusion Network and Benchmark
by: Liu, Jinyuan, et al.
Published: (2025)
by: Liu, Jinyuan, et al.
Published: (2025)
Beyond Shortcuts: Mitigating Visual Illusions in Frozen VLMs via Qualitative Reasoning
by: Guo, Hao, et al.
Published: (2026)
by: Guo, Hao, et al.
Published: (2026)
Can VLMs be used on videos for action recognition? LLMs are Visual Reasoning Coordinators
by: Lunia, Harsh
Published: (2024)
by: Lunia, Harsh
Published: (2024)
Similar Items
-
A Neurosymbolic Agent System for Compositional Visual Reasoning
by: Xu, Yichang, et al.
Published: (2025) -
Attention-aware Inference Optimizations for Large Vision-Language Models with Memory-efficient Decoding
by: Ilhan, Fatih, et al.
Published: (2026) -
A Multi-Agent Perception-Action Alliance for Efficient Long Video Reasoning
by: Xu, Yichang, et al.
Published: (2026) -
Adversarial Attention Perturbations for Large Object Detection Transformers
by: Yahn, Zachary, et al.
Published: (2025) -
Efficient Multitask Dense Predictor via Binarization
by: Shang, Yuzhang, et al.
Published: (2024)