Decomposing Visual Classification: Assessing Tree-Based Reasoning in VLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Elmansoury, Sary, Mesabah, Islam, Großmann, Gerrit, Neigel, Peter, Bhalwankar, Raj, Kondermann, Daniel, Vollmer, Sebastian J. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning
von: Galougah, Siminfar Samakoush, et al.
Veröffentlicht: (2025)
von: Galougah, Siminfar Samakoush, et al.
Veröffentlicht: (2025)
Ambiguous Annotations: When is a Pedestrian not a Pedestrian?
von: Schwirten, Luisa, et al.
Veröffentlicht: (2024)
von: Schwirten, Luisa, et al.
Veröffentlicht: (2024)
VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMs
von: Törtei, Brigitta Malagurski, et al.
Veröffentlicht: (2025)
von: Törtei, Brigitta Malagurski, et al.
Veröffentlicht: (2025)
VLMs have Tunnel Vision: Evaluating Nonlocal Visual Reasoning in Leading VLMs
von: Berman, Shmuel, et al.
Veröffentlicht: (2025)
von: Berman, Shmuel, et al.
Veröffentlicht: (2025)
Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning
von: Ge, Yuyao, et al.
Veröffentlicht: (2025)
von: Ge, Yuyao, et al.
Veröffentlicht: (2025)
VRIQ: Benchmarking and Analyzing Visual-Reasoning IQ of VLMs
von: Khezresmaeilzadeh, Tina, et al.
Veröffentlicht: (2026)
von: Khezresmaeilzadeh, Tina, et al.
Veröffentlicht: (2026)
Beyond Shortcuts: Mitigating Visual Illusions in Frozen VLMs via Qualitative Reasoning
von: Guo, Hao, et al.
Veröffentlicht: (2026)
von: Guo, Hao, et al.
Veröffentlicht: (2026)
Can VLMs be used on videos for action recognition? LLMs are Visual Reasoning Coordinators
von: Lunia, Harsh
Veröffentlicht: (2024)
von: Lunia, Harsh
Veröffentlicht: (2024)
Vision Verification Enhanced Fusion of VLMs for Efficient Visual Reasoning
von: Tekin, Selim Furkan, et al.
Veröffentlicht: (2026)
von: Tekin, Selim Furkan, et al.
Veröffentlicht: (2026)
Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs
von: Zhang, Xintong, et al.
Veröffentlicht: (2025)
von: Zhang, Xintong, et al.
Veröffentlicht: (2025)
No Need to Sacrifice Data Quality for Quantity: Crowd-Informed Machine Annotation for Cost-Effective Understanding of Visual Data
von: Klugmann, Christopher, et al.
Veröffentlicht: (2024)
von: Klugmann, Christopher, et al.
Veröffentlicht: (2024)
Activating Visual Context and Commonsense Reasoning through Masked Prediction in VLMs
von: Yu, Jiaao, et al.
Veröffentlicht: (2025)
von: Yu, Jiaao, et al.
Veröffentlicht: (2025)
Think Twice to See More: Iterative Visual Reasoning in Medical VLMs
von: Chen, Kaitao, et al.
Veröffentlicht: (2025)
von: Chen, Kaitao, et al.
Veröffentlicht: (2025)
CLGRPO: Reasoning Ability Enhancement for Small VLMs
von: Wang, Fanyi, et al.
Veröffentlicht: (2025)
von: Wang, Fanyi, et al.
Veröffentlicht: (2025)
Assessing Situational and Spatial Awareness of VLMs with Synthetically Generated Video
von: Benschop, Pascal, et al.
Veröffentlicht: (2026)
von: Benschop, Pascal, et al.
Veröffentlicht: (2026)
Visual Structures Helps Visual Reasoning: Addressing the Binding Problem in VLMs
von: Izadi, Amirmohammad, et al.
Veröffentlicht: (2025)
von: Izadi, Amirmohammad, et al.
Veröffentlicht: (2025)
Thinking with Gaze: Sequential Eye-Tracking as Visual Reasoning Supervision for Medical VLMs
von: Li, Yiwei, et al.
Veröffentlicht: (2026)
von: Li, Yiwei, et al.
Veröffentlicht: (2026)
Test-Time Reasoning Through Visual Human Preferences with VLMs and Soft Rewards
von: Gambashidze, Alexander, et al.
Veröffentlicht: (2025)
von: Gambashidze, Alexander, et al.
Veröffentlicht: (2025)
Decompose and Compare Consistency: Measuring VLMs' Answer Reliability via Task-Decomposition Consistency Comparison
von: Yang, Qian, et al.
Veröffentlicht: (2024)
von: Yang, Qian, et al.
Veröffentlicht: (2024)
Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding
von: Yoon, Hee Suk, et al.
Veröffentlicht: (2026)
von: Yoon, Hee Suk, et al.
Veröffentlicht: (2026)
How Auxiliary Reasoning Unleashes GUI Grounding in VLMs
von: Li, Weiming, et al.
Veröffentlicht: (2025)
von: Li, Weiming, et al.
Veröffentlicht: (2025)
iVISPAR -- An Interactive Visual-Spatial Reasoning Benchmark for VLMs
von: Mayer, Julius, et al.
Veröffentlicht: (2025)
von: Mayer, Julius, et al.
Veröffentlicht: (2025)
Improved Zero-Shot Classification by Adapting VLMs with Text Descriptions
von: Saha, Oindrila, et al.
Veröffentlicht: (2024)
von: Saha, Oindrila, et al.
Veröffentlicht: (2024)
Caption This, Reason That: VLMs Caught in the Middle
von: Weng, Zihan, et al.
Veröffentlicht: (2025)
von: Weng, Zihan, et al.
Veröffentlicht: (2025)
ConMe: Rethinking Evaluation of Compositional Reasoning for Modern VLMs
von: Huang, Irene, et al.
Veröffentlicht: (2024)
von: Huang, Irene, et al.
Veröffentlicht: (2024)
TIR-Flow: Active Video Search and Reasoning with Frozen VLMs
von: Jin, Hongbo, et al.
Veröffentlicht: (2026)
von: Jin, Hongbo, et al.
Veröffentlicht: (2026)
Toward Inherently Robust VLMs Against Visual Perception Attacks
von: MohajerAnsari, Pedram, et al.
Veröffentlicht: (2025)
von: MohajerAnsari, Pedram, et al.
Veröffentlicht: (2025)
Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs
von: Qiao, Yuxuan, et al.
Veröffentlicht: (2024)
von: Qiao, Yuxuan, et al.
Veröffentlicht: (2024)
Probing Visual Language Priors in VLMs
von: Luo, Tiange, et al.
Veröffentlicht: (2024)
von: Luo, Tiange, et al.
Veröffentlicht: (2024)
De-fine: Decomposing and Refining Visual Programs with Auto-Feedback
von: Gao, Minghe, et al.
Veröffentlicht: (2023)
von: Gao, Minghe, et al.
Veröffentlicht: (2023)
ViFP: A Framework for Visual False Positive Detection to Enhance Reasoning Reliability in VLMs
von: Zhang, Ben, et al.
Veröffentlicht: (2025)
von: Zhang, Ben, et al.
Veröffentlicht: (2025)
Can you SPLICE it together? A Human Curated Benchmark for Probing Visual Reasoning in VLMs
von: Ballout, Mohamad, et al.
Veröffentlicht: (2025)
von: Ballout, Mohamad, et al.
Veröffentlicht: (2025)
Benchmarking VLMs' Reasoning About Persuasive Atypical Images
von: Malakouti, Sina, et al.
Veröffentlicht: (2024)
von: Malakouti, Sina, et al.
Veröffentlicht: (2024)
NICE FACT: Diagnosing and Calibrating VLMs in Quantitative Reasoning for Kinematic Physics
von: Lan, Jian, et al.
Veröffentlicht: (2026)
von: Lan, Jian, et al.
Veröffentlicht: (2026)
From Label Error Detection to Correction: A Modular Framework and Benchmark for Object Detection Datasets
von: Penquitt, Sarina, et al.
Veröffentlicht: (2025)
von: Penquitt, Sarina, et al.
Veröffentlicht: (2025)
Minority Reports: Balancing Cost and Quality in Ground Truth Data Annotation
von: Liao, Hsuan Wei, et al.
Veröffentlicht: (2025)
von: Liao, Hsuan Wei, et al.
Veröffentlicht: (2025)
[De|Re]constructing VLMs' Reasoning in Counting
von: Alghisi, Simone, et al.
Veröffentlicht: (2025)
von: Alghisi, Simone, et al.
Veröffentlicht: (2025)
Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth
von: Wu, Yuhuan, et al.
Veröffentlicht: (2026)
von: Wu, Yuhuan, et al.
Veröffentlicht: (2026)
HIVTP: A Training-Free Method to Improve VLMs Efficiency via Hierarchical Visual Token Pruning Using Middle-Layer-Based Importance Score
von: Xu, Jingqi, et al.
Veröffentlicht: (2025)
von: Xu, Jingqi, et al.
Veröffentlicht: (2025)
Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation
von: Han, Su Ho, et al.
Veröffentlicht: (2025)
von: Han, Su Ho, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning
von: Galougah, Siminfar Samakoush, et al.
Veröffentlicht: (2025) -
Ambiguous Annotations: When is a Pedestrian not a Pedestrian?
von: Schwirten, Luisa, et al.
Veröffentlicht: (2024) -
VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMs
von: Törtei, Brigitta Malagurski, et al.
Veröffentlicht: (2025) -
VLMs have Tunnel Vision: Evaluating Nonlocal Visual Reasoning in Leading VLMs
von: Berman, Shmuel, et al.
Veröffentlicht: (2025) -
Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning
von: Ge, Yuyao, et al.
Veröffentlicht: (2025)