Journey Before Destination: On the importance of Visual Faithfulness in Slow Thinking
Fuente:
arXiv
Saved in:
| Main Authors: | Uppaal, Rheeya, Htut, Phu Mon, Bai, Min, Pappas, Nikolaos, Qi, Zheng, Swamy, Sandesh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When Safety Fails Before the Answer: Benchmarking Harmful Behavior Detection in Reasoning Chains
by: Kakkar, Ishita, et al.
Published: (2026)
by: Kakkar, Ishita, et al.
Published: (2026)
How Useful is Continued Pre-Training for Generative Unsupervised Domain Adaptation?
by: Uppaal, Rheeya, et al.
Published: (2024)
by: Uppaal, Rheeya, et al.
Published: (2024)
Rethinking LLM Uncertainty: A Multi-Agent Approach to Estimating Black-Box Model Uncertainty
by: Feng, Yu, et al.
Published: (2024)
by: Feng, Yu, et al.
Published: (2024)
Towards Long Context Hallucination Detection
by: Liu, Siyi, et al.
Published: (2025)
by: Liu, Siyi, et al.
Published: (2025)
Faithful-MR1: Faithful Multimodal Reasoning via Anchoring and Reinforcing Visual Attention
by: Tian, Changyuan, et al.
Published: (2026)
by: Tian, Changyuan, et al.
Published: (2026)
Model Editing as a Robust and Denoised variant of DPO: A Case Study on Toxicity
by: Uppaal, Rheeya, et al.
Published: (2024)
by: Uppaal, Rheeya, et al.
Published: (2024)
On the Faithfulness of Visual Thinking: Measurement and Enhancement
by: Liu, Zujing, et al.
Published: (2025)
by: Liu, Zujing, et al.
Published: (2025)
Forest Before Trees: Latent Superposition for Efficient Visual Reasoning
by: Wang, Yubo, et al.
Published: (2026)
by: Wang, Yubo, et al.
Published: (2026)
Fast-Slow Thinking GRPO for Large Vision-Language Model Reasoning
by: Xiao, Wenyi, et al.
Published: (2025)
by: Xiao, Wenyi, et al.
Published: (2025)
Contextual Knowledge Pursuit for Faithful Visual Synthesis
by: Luo, Jinqi, et al.
Published: (2023)
by: Luo, Jinqi, et al.
Published: (2023)
VinaBench: Benchmark for Faithful and Consistent Visual Narratives
by: Gao, Silin, et al.
Published: (2025)
by: Gao, Silin, et al.
Published: (2025)
Think Twice, Click Once: Enhancing GUI Grounding via Fast and Slow Systems
by: Tang, Fei, et al.
Published: (2025)
by: Tang, Fei, et al.
Published: (2025)
Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation
by: Qi, Zheng, et al.
Published: (2025)
by: Qi, Zheng, et al.
Published: (2025)
Chatting with Images for Introspective Visual Thinking
by: Wu, Junfei, et al.
Published: (2026)
by: Wu, Junfei, et al.
Published: (2026)
Look Again, Think Slowly: Enhancing Visual Reflection in Vision-Language Models
by: Jian, Pu, et al.
Published: (2025)
by: Jian, Pu, et al.
Published: (2025)
Learning to Think Fast and Slow for Visual Language Models
by: Lin, Chenyu, et al.
Published: (2025)
by: Lin, Chenyu, et al.
Published: (2025)
Reference-Based Post-OCR Processing with LLM for Precise Diacritic Text in Historical Document Recognition
by: Do, Thao, et al.
Published: (2024)
by: Do, Thao, et al.
Published: (2024)
Benchmarking and Improving Large Vision-Language Models for Fundamental Visual Graph Understanding and Reasoning
by: Zhu, Yingjie, et al.
Published: (2024)
by: Zhu, Yingjie, et al.
Published: (2024)
Watch Before You Answer: Learning from Visually Grounded Post-Training
by: Zhang, Yuxuan, et al.
Published: (2026)
by: Zhang, Yuxuan, et al.
Published: (2026)
SemVink: Advancing VLMs' Semantic Understanding of Optical Illusions via Visual Global Thinking
by: Li, Sifan, et al.
Published: (2025)
by: Li, Sifan, et al.
Published: (2025)
Reading, Not Thinking: Understanding and Bridging the Modality Gap When Text Becomes Pixels in Multimodal LLMs
by: Sun, Kaiser, et al.
Published: (2026)
by: Sun, Kaiser, et al.
Published: (2026)
Thinking with Programming Vision: Towards a Unified View for Thinking with Images
by: Guo, Zirun, et al.
Published: (2025)
by: Guo, Zirun, et al.
Published: (2025)
SPD-Faith Bench: Diagnosing and Improving Faithfulness in Chain-of-Thought for Multimodal Large Language Models
by: Lv, Weijiang, et al.
Published: (2026)
by: Lv, Weijiang, et al.
Published: (2026)
Think Visually, Reason Textually: Vision-Language Synergy in ARC
by: Zhang, Beichen, et al.
Published: (2025)
by: Zhang, Beichen, et al.
Published: (2025)
Evaluating Reasoning Faithfulness in Medical Vision-Language Models using Multimodal Perturbations
by: Moll, Johannes, et al.
Published: (2025)
by: Moll, Johannes, et al.
Published: (2025)
VALOR-EVAL: Holistic Coverage and Faithfulness Evaluation of Large Vision-Language Models
by: Qiu, Haoyi, et al.
Published: (2024)
by: Qiu, Haoyi, et al.
Published: (2024)
Multimodal Representation Alignment for Image Generation: Text-Image Interleaved Control Is Easier Than You Think
by: Chen, Liang, et al.
Published: (2025)
by: Chen, Liang, et al.
Published: (2025)
SafeLens: Deliberate and Efficient Video Guardrails with Fast-and-Slow Screening
by: Nahin, Shahriar Kabir, et al.
Published: (2026)
by: Nahin, Shahriar Kabir, et al.
Published: (2026)
Weak-eval-Strong: Evaluating and Eliciting Lateral Thinking of LLMs with Situation Puzzles
by: Chen, Qi, et al.
Published: (2024)
by: Chen, Qi, et al.
Published: (2024)
Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models
by: Li, Yunxin, et al.
Published: (2025)
by: Li, Yunxin, et al.
Published: (2025)
EgoThink: Evaluating First-Person Perspective Thinking Capability of Vision-Language Models
by: Cheng, Sijie, et al.
Published: (2023)
by: Cheng, Sijie, et al.
Published: (2023)
Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
Cognitive Visual-Language Mapper: Advancing Multimodal Comprehension with Enhanced Visual Knowledge Alignment
by: Li, Yunxin, et al.
Published: (2024)
by: Li, Yunxin, et al.
Published: (2024)
CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning
by: Qi, Ji, et al.
Published: (2024)
by: Qi, Ji, et al.
Published: (2024)
Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification
by: Bai, Tianyi, et al.
Published: (2025)
by: Bai, Tianyi, et al.
Published: (2025)
GRIT: Teaching MLLMs to Think with Images
by: Fan, Yue, et al.
Published: (2025)
by: Fan, Yue, et al.
Published: (2025)
Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm
by: Tong, Jingqi, et al.
Published: (2025)
by: Tong, Jingqi, et al.
Published: (2025)
Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation
by: Yu, Seonghoon, et al.
Published: (2026)
by: Yu, Seonghoon, et al.
Published: (2026)
VLMInferSlow: Evaluating the Efficiency Robustness of Large Vision-Language Models as a Service
by: Wang, Xiasi, et al.
Published: (2025)
by: Wang, Xiasi, et al.
Published: (2025)
Expressive and Generalizable Low-rank Adaptation for Large Models via Slow Cascaded Learning
by: Li, Siwei, et al.
Published: (2024)
by: Li, Siwei, et al.
Published: (2024)
Similar Items
-
When Safety Fails Before the Answer: Benchmarking Harmful Behavior Detection in Reasoning Chains
by: Kakkar, Ishita, et al.
Published: (2026) -
How Useful is Continued Pre-Training for Generative Unsupervised Domain Adaptation?
by: Uppaal, Rheeya, et al.
Published: (2024) -
Rethinking LLM Uncertainty: A Multi-Agent Approach to Estimating Black-Box Model Uncertainty
by: Feng, Yu, et al.
Published: (2024) -
Towards Long Context Hallucination Detection
by: Liu, Siyi, et al.
Published: (2025) -
Faithful-MR1: Faithful Multimodal Reasoning via Anchoring and Reinforcing Visual Attention
by: Tian, Changyuan, et al.
Published: (2026)