Chatting with Images for Introspective Visual Thinking
Fuente:
arXiv
Salvato in:
| Autori principali: | Wu, Junfei, Guan, Jian, Liu, Qiang, Wu, Shu, Wang, Liang, Wu, Wei, Tan, Tieniu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing
di: Wu, Junfei, et al.
Pubblicazione: (2025)
di: Wu, Junfei, et al.
Pubblicazione: (2025)
Logical Closed Loop: Uncovering Object Hallucinations in Large Vision-Language Models
di: Wu, Junfei, et al.
Pubblicazione: (2024)
di: Wu, Junfei, et al.
Pubblicazione: (2024)
Mixture of Decoding: An Attention-Inspired Adaptive Decoding Strategy to Mitigate Hallucinations in Large Vision-Language Models
di: Chen, Xinlong, et al.
Pubblicazione: (2025)
di: Chen, Xinlong, et al.
Pubblicazione: (2025)
VLKEB: A Large Vision-Language Model Knowledge Editing Benchmark
di: Huang, Han, et al.
Pubblicazione: (2024)
di: Huang, Han, et al.
Pubblicazione: (2024)
GRIT: Teaching MLLMs to Think with Images
di: Fan, Yue, et al.
Pubblicazione: (2025)
di: Fan, Yue, et al.
Pubblicazione: (2025)
More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models
di: Liu, Chengzhi, et al.
Pubblicazione: (2025)
di: Liu, Chengzhi, et al.
Pubblicazione: (2025)
Seeing the Image: Prioritizing Visual Correlation by Contrastive Alignment
di: Xiao, Xin, et al.
Pubblicazione: (2024)
di: Xiao, Xin, et al.
Pubblicazione: (2024)
Hallucination Benchmark in Medical Visual Question Answering
di: Wu, Jinge, et al.
Pubblicazione: (2024)
di: Wu, Jinge, et al.
Pubblicazione: (2024)
Thinking with Generated Images
di: Chern, Ethan, et al.
Pubblicazione: (2025)
di: Chern, Ethan, et al.
Pubblicazione: (2025)
Weak-eval-Strong: Evaluating and Eliciting Lateral Thinking of LLMs with Situation Puzzles
di: Chen, Qi, et al.
Pubblicazione: (2024)
di: Chen, Qi, et al.
Pubblicazione: (2024)
A Comprehensive Survey on Test-Time Adaptation under Distribution Shifts
di: Liang, Jian, et al.
Pubblicazione: (2023)
di: Liang, Jian, et al.
Pubblicazione: (2023)
Weaving Context Across Images: Improving Vision-Language Models through Focus-Centric Visual Chains
di: Zhang, Juntian, et al.
Pubblicazione: (2025)
di: Zhang, Juntian, et al.
Pubblicazione: (2025)
Visual Planning: Let's Think Only with Images
di: Xu, Yi, et al.
Pubblicazione: (2025)
di: Xu, Yi, et al.
Pubblicazione: (2025)
Think Visually, Reason Textually: Vision-Language Synergy in ARC
di: Zhang, Beichen, et al.
Pubblicazione: (2025)
di: Zhang, Beichen, et al.
Pubblicazione: (2025)
Vero: An Open RL Recipe for General Visual Reasoning
di: Sarch, Gabriel, et al.
Pubblicazione: (2026)
di: Sarch, Gabriel, et al.
Pubblicazione: (2026)
Thinking with Map: Reinforced Parallel Map-Augmented Agent for Geolocalization
di: Ji, Yuxiang, et al.
Pubblicazione: (2026)
di: Ji, Yuxiang, et al.
Pubblicazione: (2026)
VGR: Visual Grounded Reasoning
di: Wang, Jiacong, et al.
Pubblicazione: (2025)
di: Wang, Jiacong, et al.
Pubblicazione: (2025)
VisJudge-Bench: Aesthetics and Quality Assessment of Visualizations
di: Xie, Yupeng, et al.
Pubblicazione: (2025)
di: Xie, Yupeng, et al.
Pubblicazione: (2025)
Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation
di: Guo, Ziyu, et al.
Pubblicazione: (2025)
di: Guo, Ziyu, et al.
Pubblicazione: (2025)
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
di: Wu, Chengyue, et al.
Pubblicazione: (2024)
di: Wu, Chengyue, et al.
Pubblicazione: (2024)
Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities
di: Zhou, Ziwei, et al.
Pubblicazione: (2025)
di: Zhou, Ziwei, et al.
Pubblicazione: (2025)
CaughtCheating: Is Your MLLM a Good Cheating Detective? Exploring the Boundary of Visual Perception and Reasoning
di: Li, Ming, et al.
Pubblicazione: (2025)
di: Li, Ming, et al.
Pubblicazione: (2025)
VideoExplorer: Think With Videos For Agentic Long-Video Understanding
di: Yuan, Huaying, et al.
Pubblicazione: (2025)
di: Yuan, Huaying, et al.
Pubblicazione: (2025)
From Introspection to Best Practices: Principled Analysis of Demonstrations in Multimodal In-Context Learning
di: Xu, Nan, et al.
Pubblicazione: (2024)
di: Xu, Nan, et al.
Pubblicazione: (2024)
Re-Thinking the Automatic Evaluation of Image-Text Alignment in Text-to-Image Models
di: Zhang, Huixuan, et al.
Pubblicazione: (2025)
di: Zhang, Huixuan, et al.
Pubblicazione: (2025)
RIV: Recursive Introspection Mask Diffusion Vision Language Model
di: Li, YuQian, et al.
Pubblicazione: (2025)
di: Li, YuQian, et al.
Pubblicazione: (2025)
Visually Descriptive Language Model for Vector Graphics Reasoning
di: Wang, Zhenhailong, et al.
Pubblicazione: (2024)
di: Wang, Zhenhailong, et al.
Pubblicazione: (2024)
Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation
di: Yu, Seonghoon, et al.
Pubblicazione: (2026)
di: Yu, Seonghoon, et al.
Pubblicazione: (2026)
Connecting the Dots: Collaborative Fine-tuning for Black-Box Vision-Language Models
di: Wang, Zhengbo, et al.
Pubblicazione: (2024)
di: Wang, Zhengbo, et al.
Pubblicazione: (2024)
The Factuality Tax of Diversity-Intervened Text-to-Image Generation: Benchmark and Fact-Augmented Intervention
di: Wan, Yixin, et al.
Pubblicazione: (2024)
di: Wan, Yixin, et al.
Pubblicazione: (2024)
Adversarial Attacks on VQA-NLE: Exposing and Alleviating Inconsistencies in Visual Question Answering Explanations
di: Yeh, Yahsin, et al.
Pubblicazione: (2025)
di: Yeh, Yahsin, et al.
Pubblicazione: (2025)
Out-of-distribution Evidence-aware Fake News Detection via Dual Adversarial Debiasing
di: Liu, Qiang, et al.
Pubblicazione: (2023)
di: Liu, Qiang, et al.
Pubblicazione: (2023)
VISCO: Benchmarking Fine-Grained Critique and Correction Towards Self-Improvement in Visual Reasoning
di: Wu, Xueqing, et al.
Pubblicazione: (2024)
di: Wu, Xueqing, et al.
Pubblicazione: (2024)
CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning
di: Xing, Long, et al.
Pubblicazione: (2025)
di: Xing, Long, et al.
Pubblicazione: (2025)
Thinking with Patterns: Breaking the Perceptual Bottleneck in Visual Planning via Pattern Induction
di: Jian, Yichang, et al.
Pubblicazione: (2026)
di: Jian, Yichang, et al.
Pubblicazione: (2026)
GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents
di: Wu, Qianhui, et al.
Pubblicazione: (2025)
di: Wu, Qianhui, et al.
Pubblicazione: (2025)
CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching
di: Jiang, Dongzhi, et al.
Pubblicazione: (2024)
di: Jiang, Dongzhi, et al.
Pubblicazione: (2024)
SOK-Bench: A Situated Video Reasoning Benchmark with Aligned Open-World Knowledge
di: Wang, Andong, et al.
Pubblicazione: (2024)
di: Wang, Andong, et al.
Pubblicazione: (2024)
EditReward: A Human-Aligned Reward Model for Instruction-Guided Image Editing
di: Wu, Keming, et al.
Pubblicazione: (2025)
di: Wu, Keming, et al.
Pubblicazione: (2025)
MULTITEXTEDIT: Benchmarking Cross-Lingual Degradation in Text-in-Image Editing
di: Cheng, Liwei, et al.
Pubblicazione: (2026)
di: Cheng, Liwei, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing
di: Wu, Junfei, et al.
Pubblicazione: (2025) -
Logical Closed Loop: Uncovering Object Hallucinations in Large Vision-Language Models
di: Wu, Junfei, et al.
Pubblicazione: (2024) -
Mixture of Decoding: An Attention-Inspired Adaptive Decoding Strategy to Mitigate Hallucinations in Large Vision-Language Models
di: Chen, Xinlong, et al.
Pubblicazione: (2025) -
VLKEB: A Large Vision-Language Model Knowledge Editing Benchmark
di: Huang, Han, et al.
Pubblicazione: (2024) -
GRIT: Teaching MLLMs to Think with Images
di: Fan, Yue, et al.
Pubblicazione: (2025)