VISCO: Benchmarking Fine-Grained Critique and Correction Towards Self-Improvement in Visual Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Xueqing, Ding, Yuheng, Li, Bingxuan, Lu, Pan, Yin, Da, Chang, Kai-Wei, Peng, Nanyun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On Asymmetric Optimization of Reasoning and Perception in Vision-Language Model Post-Training
von: Wu, Xueqing, et al.
Veröffentlicht: (2026)
von: Wu, Xueqing, et al.
Veröffentlicht: (2026)
FronTalk: Benchmarking Front-End Development as Conversational Code Generation with Multi-Modal Feedback
von: Wu, Xueqing, et al.
Veröffentlicht: (2025)
von: Wu, Xueqing, et al.
Veröffentlicht: (2025)
KPEval: Towards Fine-Grained Semantic-Based Keyphrase Evaluation
von: Wu, Di, et al.
Veröffentlicht: (2023)
von: Wu, Di, et al.
Veröffentlicht: (2023)
VDebugger: Harnessing Execution Feedback for Debugging Visual Programs
von: Wu, Xueqing, et al.
Veröffentlicht: (2024)
von: Wu, Xueqing, et al.
Veröffentlicht: (2024)
METAL: A Multi-Agent Framework for Chart Generation with Test-Time Scaling
von: Li, Bingxuan, et al.
Veröffentlicht: (2025)
von: Li, Bingxuan, et al.
Veröffentlicht: (2025)
Control Large Language Models via Divide and Conquer
von: Li, Bingxuan, et al.
Veröffentlicht: (2024)
von: Li, Bingxuan, et al.
Veröffentlicht: (2024)
Re-ReST: Reflection-Reinforced Self-Training for Language Agents
von: Dou, Zi-Yi, et al.
Veröffentlicht: (2024)
von: Dou, Zi-Yi, et al.
Veröffentlicht: (2024)
Guiding Through Complexity: What Makes Good Supervision for Hard Math Reasoning Tasks?
von: He, Xuan, et al.
Veröffentlicht: (2024)
von: He, Xuan, et al.
Veröffentlicht: (2024)
Are LLMs Capable of Data-based Statistical and Causal Reasoning? Benchmarking Advanced Quantitative Reasoning with Data
von: Liu, Xiao, et al.
Veröffentlicht: (2024)
von: Liu, Xiao, et al.
Veröffentlicht: (2024)
DialectGen: Benchmarking and Improving Dialect Robustness in Multimodal Generation
von: Zhou, Yu, et al.
Veröffentlicht: (2025)
von: Zhou, Yu, et al.
Veröffentlicht: (2025)
ConTextual: Evaluating Context-Sensitive Text-Rich Visual Reasoning in Large Multimodal Models
von: Wadhawan, Rohan, et al.
Veröffentlicht: (2024)
von: Wadhawan, Rohan, et al.
Veröffentlicht: (2024)
OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles
von: Deng, Yihe, et al.
Veröffentlicht: (2025)
von: Deng, Yihe, et al.
Veröffentlicht: (2025)
Contrastive Visual Data Augmentation
von: Zhou, Yu, et al.
Veröffentlicht: (2025)
von: Zhou, Yu, et al.
Veröffentlicht: (2025)
OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks
von: Hu, Wenbo, et al.
Veröffentlicht: (2026)
von: Hu, Wenbo, et al.
Veröffentlicht: (2026)
SafeWorld: Geo-Diverse Safety Alignment
von: Yin, Da, et al.
Veröffentlicht: (2024)
von: Yin, Da, et al.
Veröffentlicht: (2024)
Self-Routing RAG: Binding Selective Retrieval with Knowledge Verbalization
von: Wu, Di, et al.
Veröffentlicht: (2025)
von: Wu, Di, et al.
Veröffentlicht: (2025)
CoKe: Customizable Fine-Grained Story Evaluation via Chain-of-Keyword Rationalization
von: Joshi, Brihi, et al.
Veröffentlicht: (2025)
von: Joshi, Brihi, et al.
Veröffentlicht: (2025)
CompAlign: Improving Compositional Text-to-Image Generation with a Complex Benchmark and Fine-Grained Feedback
von: Wan, Yixin, et al.
Veröffentlicht: (2025)
von: Wan, Yixin, et al.
Veröffentlicht: (2025)
REFFLY: Melody-Constrained Lyrics Editing Model
von: Zhao, Songyan, et al.
Veröffentlicht: (2024)
von: Zhao, Songyan, et al.
Veröffentlicht: (2024)
LLM Self-Correction with DeCRIM: Decompose, Critique, and Refine for Enhanced Following of Instructions with Multiple Constraints
von: Ferraz, Thomas Palmeira, et al.
Veröffentlicht: (2024)
von: Ferraz, Thomas Palmeira, et al.
Veröffentlicht: (2024)
CriticBench: Benchmarking LLMs for Critique-Correct Reasoning
von: Lin, Zicheng, et al.
Veröffentlicht: (2024)
von: Lin, Zicheng, et al.
Veröffentlicht: (2024)
ReasonMap: Towards Fine-Grained Visual Reasoning from Transit Maps
von: Feng, Sicheng, et al.
Veröffentlicht: (2025)
von: Feng, Sicheng, et al.
Veröffentlicht: (2025)
Catching the Details: Self-Distilled RoI Predictors for Fine-Grained MLLM Perception
von: Shi, Yuheng, et al.
Veröffentlicht: (2025)
von: Shi, Yuheng, et al.
Veröffentlicht: (2025)
BRIEF: Bridging Retrieval and Inference for Multi-hop Reasoning via Compression
von: Li, Yuankai, et al.
Veröffentlicht: (2024)
von: Li, Yuankai, et al.
Veröffentlicht: (2024)
Visual Reasoning Tracer: Object-Level Grounded Reasoning Benchmark
von: Yuan, Haobo, et al.
Veröffentlicht: (2025)
von: Yuan, Haobo, et al.
Veröffentlicht: (2025)
Verbalized Representation Learning for Interpretable Few-Shot Generalization
von: Yang, Cheng-Fu, et al.
Veröffentlicht: (2024)
von: Yang, Cheng-Fu, et al.
Veröffentlicht: (2024)
DACO: Towards Application-Driven and Comprehensive Data Analysis via Code Generation
von: Wu, Xueqing, et al.
Veröffentlicht: (2024)
von: Wu, Xueqing, et al.
Veröffentlicht: (2024)
Synchronous Faithfulness Monitoring for Trustworthy Retrieval-Augmented Generation
von: Wu, Di, et al.
Veröffentlicht: (2024)
von: Wu, Di, et al.
Veröffentlicht: (2024)
DiCoRe: Enhancing Zero-shot Event Detection via Divergent-Convergent LLM Reasoning
von: Parekh, Tanmay, et al.
Veröffentlicht: (2025)
von: Parekh, Tanmay, et al.
Veröffentlicht: (2025)
Towards Fine-Grained Recognition with Large Visual Language Models: Benchmark and Optimization Strategies
von: Pang, Cong, et al.
Veröffentlicht: (2025)
von: Pang, Cong, et al.
Veröffentlicht: (2025)
AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning
von: Galougah, Siminfar Samakoush, et al.
Veröffentlicht: (2025)
von: Galougah, Siminfar Samakoush, et al.
Veröffentlicht: (2025)
Joint Optimization of Reasoning and Dual-Memory for Self-Learning Diagnostic Agent
von: Li, Bingxuan, et al.
Veröffentlicht: (2026)
von: Li, Bingxuan, et al.
Veröffentlicht: (2026)
BRIEF-Pro: Universal Context Compression with Short-to-Long Synthesis for Fast and Accurate Multi-Hop Reasoning
von: Gu, Jia-Chen, et al.
Veröffentlicht: (2025)
von: Gu, Jia-Chen, et al.
Veröffentlicht: (2025)
Fine-Grained Self-Endorsement Improves Factuality and Reasoning
von: Wang, Ante, et al.
Veröffentlicht: (2024)
von: Wang, Ante, et al.
Veröffentlicht: (2024)
DeepEdit: Knowledge Editing as Decoding with Constraints
von: Wang, Yiwei, et al.
Veröffentlicht: (2024)
von: Wang, Yiwei, et al.
Veröffentlicht: (2024)
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning
von: Wu, Chengfei, et al.
Veröffentlicht: (2025)
von: Wu, Chengfei, et al.
Veröffentlicht: (2025)
TikArt: Stabilizing Aperture-Guided Fine-Grained Visual Reasoning with Reinforcement Learning
von: Ding, Hao, et al.
Veröffentlicht: (2026)
von: Ding, Hao, et al.
Veröffentlicht: (2026)
Med-CMR: A Fine-Grained Benchmark Integrating Visual Evidence and Clinical Logic for Medical Complex Multimodal Reasoning
von: Gong, Haozhen, et al.
Veröffentlicht: (2025)
von: Gong, Haozhen, et al.
Veröffentlicht: (2025)
Dancing with Critiques: Enhancing LLM Reasoning with Stepwise Natural Language Self-Critique
von: Li, Yansi, et al.
Veröffentlicht: (2025)
von: Li, Yansi, et al.
Veröffentlicht: (2025)
Beyond Query-Level Comparison: Fine-Grained Reinforcement Learning for Text-to-SQL with Automated Interpretable Critiques
von: Wang, Guifeng, et al.
Veröffentlicht: (2025)
von: Wang, Guifeng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
On Asymmetric Optimization of Reasoning and Perception in Vision-Language Model Post-Training
von: Wu, Xueqing, et al.
Veröffentlicht: (2026) -
FronTalk: Benchmarking Front-End Development as Conversational Code Generation with Multi-Modal Feedback
von: Wu, Xueqing, et al.
Veröffentlicht: (2025) -
KPEval: Towards Fine-Grained Semantic-Based Keyphrase Evaluation
von: Wu, Di, et al.
Veröffentlicht: (2023) -
VDebugger: Harnessing Execution Feedback for Debugging Visual Programs
von: Wu, Xueqing, et al.
Veröffentlicht: (2024) -
METAL: A Multi-Agent Framework for Chart Generation with Test-Time Scaling
von: Li, Bingxuan, et al.
Veröffentlicht: (2025)