See, Think, Learn: A Self-Taught Multimodal Reasoner
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sharma, Sourabh, Gupta, Sonam, Sadbhawna |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models
von: Liu, Chengzhi, et al.
Veröffentlicht: (2025)
von: Liu, Chengzhi, et al.
Veröffentlicht: (2025)
Seeing but Not Thinking: Routing Distraction in Multimodal Mixture-of-Experts
von: Xu, Haolei, et al.
Veröffentlicht: (2026)
von: Xu, Haolei, et al.
Veröffentlicht: (2026)
Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models
von: Li, Yunxin, et al.
Veröffentlicht: (2025)
von: Li, Yunxin, et al.
Veröffentlicht: (2025)
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models
von: Wu, Juncheng, et al.
Veröffentlicht: (2026)
von: Wu, Juncheng, et al.
Veröffentlicht: (2026)
ReLoop: "Seeing Twice and Thinking Backwards" via Closed-loop Training to Mitigate Hallucinations in Multimodal understanding
von: Yang, Jianjiang, et al.
Veröffentlicht: (2025)
von: Yang, Jianjiang, et al.
Veröffentlicht: (2025)
Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm
von: Tong, Jingqi, et al.
Veröffentlicht: (2025)
von: Tong, Jingqi, et al.
Veröffentlicht: (2025)
Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation
von: Yu, Seonghoon, et al.
Veröffentlicht: (2026)
von: Yu, Seonghoon, et al.
Veröffentlicht: (2026)
From Perception to Reasoning: Deep Thinking Empowers Multimodal Large Language Models
von: Zhu, Wenxin, et al.
Veröffentlicht: (2025)
von: Zhu, Wenxin, et al.
Veröffentlicht: (2025)
Multimodal Reasoning for Science: Technical Report and 1st Place Solution to the ICML 2025 SeePhys Challenge
von: Liang, Hao, et al.
Veröffentlicht: (2025)
von: Liang, Hao, et al.
Veröffentlicht: (2025)
Seeing Beyond Words: Self-Supervised Visual Learning for Multimodal Large Language Models
von: Caffagni, Davide, et al.
Veröffentlicht: (2025)
von: Caffagni, Davide, et al.
Veröffentlicht: (2025)
Knowledge-Aware Reasoning over Multimodal Semi-structured Tables
von: Mathur, Suyash Vardhan, et al.
Veröffentlicht: (2024)
von: Mathur, Suyash Vardhan, et al.
Veröffentlicht: (2024)
See, Explain, and Intervene: A Few-Shot Multimodal Agent Framework for Hateful Meme Moderation
von: Rizwan, Naquee, et al.
Veröffentlicht: (2026)
von: Rizwan, Naquee, et al.
Veröffentlicht: (2026)
Seeing What Tastes Good: Revisiting Multimodal Distributional Semantics in the Billion Parameter Era
von: Oneata, Dan, et al.
Veröffentlicht: (2025)
von: Oneata, Dan, et al.
Veröffentlicht: (2025)
v1: Learning to Point Visual Tokens for Multimodal Grounded Reasoning
von: Chung, Jiwan, et al.
Veröffentlicht: (2025)
von: Chung, Jiwan, et al.
Veröffentlicht: (2025)
ViewFusion: Structured Spatial Thinking Chains for Multi-View Reasoning
von: Tao, Xingjian, et al.
Veröffentlicht: (2026)
von: Tao, Xingjian, et al.
Veröffentlicht: (2026)
Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
von: Yang, Jihan, et al.
Veröffentlicht: (2024)
von: Yang, Jihan, et al.
Veröffentlicht: (2024)
Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation
von: Yuan, Qianhao, et al.
Veröffentlicht: (2026)
von: Yuan, Qianhao, et al.
Veröffentlicht: (2026)
Draw2Think: Harnessing Geometry Reasoning through Constraint Engine Interaction
von: Hu, Juncheng, et al.
Veröffentlicht: (2026)
von: Hu, Juncheng, et al.
Veröffentlicht: (2026)
Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture
von: Zhang, Longxiang, et al.
Veröffentlicht: (2026)
von: Zhang, Longxiang, et al.
Veröffentlicht: (2026)
Reading, Not Thinking: Understanding and Bridging the Modality Gap When Text Becomes Pixels in Multimodal LLMs
von: Sun, Kaiser, et al.
Veröffentlicht: (2026)
von: Sun, Kaiser, et al.
Veröffentlicht: (2026)
LaRe: Latent Refocusing for Multimodal Reasoning
von: Ma, Jizheng, et al.
Veröffentlicht: (2025)
von: Ma, Jizheng, et al.
Veröffentlicht: (2025)
Reinforcing Multimodal Reasoning Against Visual Degradation
von: Liu, Rui, et al.
Veröffentlicht: (2026)
von: Liu, Rui, et al.
Veröffentlicht: (2026)
BLINK: Multimodal Large Language Models Can See but Not Perceive
von: Fu, Xingyu, et al.
Veröffentlicht: (2024)
von: Fu, Xingyu, et al.
Veröffentlicht: (2024)
From Sight to Insight: Improving Visual Reasoning Capabilities of Multimodal Models via Reinforcement Learning
von: Sharif, Omar, et al.
Veröffentlicht: (2026)
von: Sharif, Omar, et al.
Veröffentlicht: (2026)
Seeing Culture: A Benchmark for Visual Reasoning and Grounding
von: Satar, Burak, et al.
Veröffentlicht: (2025)
von: Satar, Burak, et al.
Veröffentlicht: (2025)
One RL to See Them All: Visual Triple Unified Reinforcement Learning
von: Ma, Yan, et al.
Veröffentlicht: (2025)
von: Ma, Yan, et al.
Veröffentlicht: (2025)
Learn to Think: Improving Multimodal Reasoning through Vision-Aware Self-Improvement Training
von: Zhong, Qihuang, et al.
Veröffentlicht: (2026)
von: Zhong, Qihuang, et al.
Veröffentlicht: (2026)
Multimodal Representation Alignment for Image Generation: Text-Image Interleaved Control Is Easier Than You Think
von: Chen, Liang, et al.
Veröffentlicht: (2025)
von: Chen, Liang, et al.
Veröffentlicht: (2025)
Seeing Through Deception: Uncovering Misleading Creator Intent in Multimodal News with Vision-Language Models
von: Wu, Jiaying, et al.
Veröffentlicht: (2025)
von: Wu, Jiaying, et al.
Veröffentlicht: (2025)
Unleashing Perception-Time Scaling to Multimodal Reasoning Models
von: Li, Yifan, et al.
Veröffentlicht: (2025)
von: Li, Yifan, et al.
Veröffentlicht: (2025)
Probabilistic Concept Graph Reasoning for Multimodal Misinformation Detection
von: Yang, Ruichao, et al.
Veröffentlicht: (2026)
von: Yang, Ruichao, et al.
Veröffentlicht: (2026)
Can Multimodal LLMs See Materials Clearly? A Multimodal Benchmark on Materials Characterization
von: Lai, Zhengzhao, et al.
Veröffentlicht: (2025)
von: Lai, Zhengzhao, et al.
Veröffentlicht: (2025)
Seeing Through the MiRAGE: Evaluating Multimodal Retrieval Augmented Generation
von: Martin, Alexander, et al.
Veröffentlicht: (2025)
von: Martin, Alexander, et al.
Veröffentlicht: (2025)
Unleashing Spatial Reasoning in Multimodal Large Language Models via Textual Representation Guided Reasoning
von: Hua, Jiacheng, et al.
Veröffentlicht: (2026)
von: Hua, Jiacheng, et al.
Veröffentlicht: (2026)
Shared and Private Information Learning in Multimodal Sentiment Analysis with Deep Modal Alignment and Self-supervised Multi-Task Learning
von: Lai, Songning, et al.
Veröffentlicht: (2023)
von: Lai, Songning, et al.
Veröffentlicht: (2023)
MME-Finance: A Multimodal Finance Benchmark for Expert-level Understanding and Reasoning
von: Gan, Ziliang, et al.
Veröffentlicht: (2024)
von: Gan, Ziliang, et al.
Veröffentlicht: (2024)
Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space
von: Liu, Chengzhi, et al.
Veröffentlicht: (2025)
von: Liu, Chengzhi, et al.
Veröffentlicht: (2025)
VisualPRM: An Effective Process Reward Model for Multimodal Reasoning
von: Wang, Weiyun, et al.
Veröffentlicht: (2025)
von: Wang, Weiyun, et al.
Veröffentlicht: (2025)
Plug-and-Play Grounding of Reasoning in Multimodal Large Language Models
von: Chen, Jiaxing, et al.
Veröffentlicht: (2024)
von: Chen, Jiaxing, et al.
Veröffentlicht: (2024)
Knowing When Not to Answer: Evaluating Abstention in Multimodal Reasoning Systems
von: Madhusudhan, Nishanth, et al.
Veröffentlicht: (2026)
von: Madhusudhan, Nishanth, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models
von: Liu, Chengzhi, et al.
Veröffentlicht: (2025) -
Seeing but Not Thinking: Routing Distraction in Multimodal Mixture-of-Experts
von: Xu, Haolei, et al.
Veröffentlicht: (2026) -
Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models
von: Li, Yunxin, et al.
Veröffentlicht: (2025) -
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models
von: Wu, Juncheng, et al.
Veröffentlicht: (2026) -
ReLoop: "Seeing Twice and Thinking Backwards" via Closed-loop Training to Mitigate Hallucinations in Multimodal understanding
von: Yang, Jianjiang, et al.
Veröffentlicht: (2025)