No Labels, No Problem: Training Visual Reasoners with Multimodal Verifiers
Fuente:
arXiv
Saved in:
| Main Authors: | Marsili, Damiano, Gkioxari, Georgia |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Visual Agentic AI for Spatial Reasoning with a Dynamic API
by: Marsili, Damiano, et al.
Published: (2025)
by: Marsili, Damiano, et al.
Published: (2025)
Same or Not? Enhancing Visual Perception in Vision-Language Models
by: Marsili, Damiano, et al.
Published: (2025)
by: Marsili, Damiano, et al.
Published: (2025)
Feedforward 3D Editing via Text-Steerable Image-to-3D
by: Ma, Ziqi, et al.
Published: (2025)
by: Ma, Ziqi, et al.
Published: (2025)
Generative Universal Verifier as Multimodal Meta-Reasoner
by: Zhang, Xinchen, et al.
Published: (2025)
by: Zhang, Xinchen, et al.
Published: (2025)
Multimodal Fact-Level Attribution for Verifiable Reasoning
by: Wan, David, et al.
Published: (2026)
by: Wan, David, et al.
Published: (2026)
Unlocking Complex Visual Generation via Closed-Loop Verified Reasoning
by: Cheng, Hanbo, et al.
Published: (2026)
by: Cheng, Hanbo, et al.
Published: (2026)
CoRGI: Verified Chain-of-Thought Reasoning with Post-hoc Visual Grounding
by: Yi, Shixin, et al.
Published: (2025)
by: Yi, Shixin, et al.
Published: (2025)
Conversational Image Segmentation: Grounding Abstract Concepts with Scalable Supervision
by: Sahoo, Aadarsh, et al.
Published: (2026)
by: Sahoo, Aadarsh, et al.
Published: (2026)
The Role of Visual Modality in Multimodal Mathematical Reasoning: Challenges and Insights
by: Liu, Yufang, et al.
Published: (2025)
by: Liu, Yufang, et al.
Published: (2025)
GenEARL: A Training-Free Generative Framework for Multimodal Event Argument Role Labeling
by: Bansal, Hritik, et al.
Published: (2024)
by: Bansal, Hritik, et al.
Published: (2024)
Linear Mechanisms for Spatiotemporal Reasoning in Vision Language Models
by: Kang, Raphi, et al.
Published: (2026)
by: Kang, Raphi, et al.
Published: (2026)
Progressive Multimodal Search and Reasoning for Knowledge-Intensive Visual Question Answering
by: Choi, Changin, et al.
Published: (2025)
by: Choi, Changin, et al.
Published: (2025)
Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens
by: Yang, Zeyuan, et al.
Published: (2025)
by: Yang, Zeyuan, et al.
Published: (2025)
Self-Rewarded Multimodal Coherent Reasoning Across Diverse Visual Domains
by: Zhang, Jesen, et al.
Published: (2025)
by: Zhang, Jesen, et al.
Published: (2025)
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models
by: Zhan, Yufei, et al.
Published: (2025)
by: Zhan, Yufei, et al.
Published: (2025)
Chain-of-Thought Degrades Visual Spatial Reasoning Capabilities of Multimodal LLMs
by: Kancheti, Sai Srinivas, et al.
Published: (2026)
by: Kancheti, Sai Srinivas, et al.
Published: (2026)
PerturboLLaVA: Reducing Multimodal Hallucinations with Perturbative Visual Training
by: Chen, Cong, et al.
Published: (2025)
by: Chen, Cong, et al.
Published: (2025)
Training-Free Mitigation of Language Reasoning Degradation After Multimodal Instruction Tuning
by: Ratzlaff, Neale, et al.
Published: (2024)
by: Ratzlaff, Neale, et al.
Published: (2024)
Multi-Label Contrastive Learning for Abstract Visual Reasoning
by: Małkiński, Mikołaj, et al.
Published: (2020)
by: Małkiński, Mikołaj, et al.
Published: (2020)
SATORI-R1: Incentivizing Multimodal Reasoning through Explicit Visual Anchoring
by: Shen, Chuming, et al.
Published: (2025)
by: Shen, Chuming, et al.
Published: (2025)
Revisiting Visual Understanding in Multimodal Reasoning through a Lens of Image Perturbation
by: Li, Yuting, et al.
Published: (2025)
by: Li, Yuting, et al.
Published: (2025)
DIVER: Dynamic Iterative Visual Evidence Reasoning for Multimodal Fake News Detection
by: Zhou, Weilin, et al.
Published: (2026)
by: Zhou, Weilin, et al.
Published: (2026)
PaLMR: Towards Faithful Visual Reasoning via Multimodal Process Alignment
by: Li, Yantao, et al.
Published: (2026)
by: Li, Yantao, et al.
Published: (2026)
Improved Visual-Spatial Reasoning via R1-Zero-Like Training
by: Liao, Zhenyi, et al.
Published: (2025)
by: Liao, Zhenyi, et al.
Published: (2025)
Seg-Agent: Test-Time Multimodal Reasoning for Training-Free Language-Guided Segmentation
by: Hao, Chao, et al.
Published: (2026)
by: Hao, Chao, et al.
Published: (2026)
MM-CoT:A Benchmark for Probing Visual Chain-of-Thought Reasoning in Multimodal Models
by: Zhang, Jusheng, et al.
Published: (2025)
by: Zhang, Jusheng, et al.
Published: (2025)
Language Models Can See Better: Visual Contrastive Decoding For LLM Multimodal Reasoning
by: Pang, Yuqi, et al.
Published: (2025)
by: Pang, Yuqi, et al.
Published: (2025)
Thinking Diffusion: Penalize and Guide Visual-Grounded Reasoning in Diffusion Multimodal Language Models
by: Kim, Keuntae, et al.
Published: (2026)
by: Kim, Keuntae, et al.
Published: (2026)
Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination
by: Zheng, Haojie, et al.
Published: (2024)
by: Zheng, Haojie, et al.
Published: (2024)
EVE: Verifiable Self-Evolution of MLLMs via Executable Visual Transformations
by: Heng, Yongrui, et al.
Published: (2026)
by: Heng, Yongrui, et al.
Published: (2026)
SARE: Sample-wise Adaptive Reasoning for Training-free Fine-grained Visual Recognition
by: Yang, Jingxiao, et al.
Published: (2026)
by: Yang, Jingxiao, et al.
Published: (2026)
ClueTracer: Question-to-Vision Clue Tracing for Training-Free Hallucination Suppression in Multimodal Reasoning
by: Xi, Gongli, et al.
Published: (2026)
by: Xi, Gongli, et al.
Published: (2026)
CogFlow: Bridging Perception and Reasoning through Knowledge Internalization for Visual Mathematical Problem Solving
by: Chen, Shuhang, et al.
Published: (2026)
by: Chen, Shuhang, et al.
Published: (2026)
GoViG: Goal-Conditioned Visual Navigation Instruction Generation via Multimodal Reasoning
by: Wu, Fengyi, et al.
Published: (2025)
by: Wu, Fengyi, et al.
Published: (2025)
VTPerception-R1: Enhancing Multimodal Reasoning via Explicit Visual and Textual Perceptual Grounding
by: Ding, Yizhuo, et al.
Published: (2025)
by: Ding, Yizhuo, et al.
Published: (2025)
Graph-of-Mark: Promote Spatial Reasoning in Multimodal Language Models with Graph-Based Visual Prompting
by: Frisoni, Giacomo, et al.
Published: (2026)
by: Frisoni, Giacomo, et al.
Published: (2026)
ScalSelect: Scalable Training-Free Multimodal Data Selection for Efficient Visual Instruction Tuning
by: Wu, Changti, et al.
Published: (2026)
by: Wu, Changti, et al.
Published: (2026)
Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning
by: Ou, Siqu, et al.
Published: (2025)
by: Ou, Siqu, et al.
Published: (2025)
Long Grounded Thoughts: Synthesizing Visual Problems and Reasoning Chains at Scale
by: Acuna, David, et al.
Published: (2025)
by: Acuna, David, et al.
Published: (2025)
Visual Structures Helps Visual Reasoning: Addressing the Binding Problem in VLMs
by: Izadi, Amirmohammad, et al.
Published: (2025)
by: Izadi, Amirmohammad, et al.
Published: (2025)
Similar Items
-
Visual Agentic AI for Spatial Reasoning with a Dynamic API
by: Marsili, Damiano, et al.
Published: (2025) -
Same or Not? Enhancing Visual Perception in Vision-Language Models
by: Marsili, Damiano, et al.
Published: (2025) -
Feedforward 3D Editing via Text-Steerable Image-to-3D
by: Ma, Ziqi, et al.
Published: (2025) -
Generative Universal Verifier as Multimodal Meta-Reasoner
by: Zhang, Xinchen, et al.
Published: (2025) -
Multimodal Fact-Level Attribution for Verifiable Reasoning
by: Wan, David, et al.
Published: (2026)