GRIT: Teaching MLLMs to Think with Images
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fan, Yue, He, Xuehai, Yang, Diji, Zheng, Kaizhi, Kuo, Ching-Chen, Zheng, Yuting, Narayanaraju, Sravana Jyothi, Guan, Xinze, Wang, Xin Eric |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
JARVIS: A Neuro-Symbolic Commonsense Reasoning Framework for Conversational Embodied Agents
von: Zheng, Kaizhi, et al.
Veröffentlicht: (2022)
von: Zheng, Kaizhi, et al.
Veröffentlicht: (2022)
MiniGPT-5: Interleaved Vision-and-Language Generation via Generative Vokens
von: Zheng, Kaizhi, et al.
Veröffentlicht: (2023)
von: Zheng, Kaizhi, et al.
Veröffentlicht: (2023)
MorphoSim: An Interactive, Controllable, and Editable Language-guided 4D World Simulator
von: He, Xuehai, et al.
Veröffentlicht: (2025)
von: He, Xuehai, et al.
Veröffentlicht: (2025)
Multimodal Inconsistency Reasoning (MMIR): A New Benchmark for Multimodal Reasoning Models
von: Yan, Qianqi, et al.
Veröffentlicht: (2025)
von: Yan, Qianqi, et al.
Veröffentlicht: (2025)
Read Anywhere Pointed: Layout-aware GUI Screen Reading with Tree-of-Lens Grounding
von: Fan, Yue, et al.
Veröffentlicht: (2024)
von: Fan, Yue, et al.
Veröffentlicht: (2024)
Self-Evolving 3D Scene Generation from a Single Image
von: Zheng, Kaizhi, et al.
Veröffentlicht: (2025)
von: Zheng, Kaizhi, et al.
Veröffentlicht: (2025)
Muffin or Chihuahua? Challenging Multimodal Large Language Models with Multipanel VQA
von: Fan, Yue, et al.
Veröffentlicht: (2024)
von: Fan, Yue, et al.
Veröffentlicht: (2024)
Beyond Introspection: Reinforcing Thinking via Externalist Behavioral Feedback
von: Yang, Diji, et al.
Veröffentlicht: (2024)
von: Yang, Diji, et al.
Veröffentlicht: (2024)
ComCLIP: Training-Free Compositional Image and Text Matching
von: Jiang, Kenan, et al.
Veröffentlicht: (2022)
von: Jiang, Kenan, et al.
Veröffentlicht: (2022)
Image captioning for Brazilian Portuguese using GRIT model
von: de Alencar, Rafael Silva, et al.
Veröffentlicht: (2024)
von: de Alencar, Rafael Silva, et al.
Veröffentlicht: (2024)
Many Minds from One Model: Bayesian-Inspired Transformers for Population Diversity
von: Yang, Diji, et al.
Veröffentlicht: (2025)
von: Yang, Diji, et al.
Veröffentlicht: (2025)
MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos
von: He, Xuehai, et al.
Veröffentlicht: (2024)
von: He, Xuehai, et al.
Veröffentlicht: (2024)
Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Space
von: Zhang, Zhen, et al.
Veröffentlicht: (2025)
von: Zhang, Zhen, et al.
Veröffentlicht: (2025)
MM-THEBench: Do Reasoning MLLMs Think Reasonably?
von: Huang, Zhidian, et al.
Veröffentlicht: (2026)
von: Huang, Zhidian, et al.
Veröffentlicht: (2026)
VULCA-Bench: A Multicultural Vision-Language Benchmark for Evaluating Cultural Understanding
von: Yu, Haorui, et al.
Veröffentlicht: (2026)
von: Yu, Haorui, et al.
Veröffentlicht: (2026)
Bridging the Gap Between Multimodal Foundation Models and World Models
von: He, Xuehai
Veröffentlicht: (2025)
von: He, Xuehai
Veröffentlicht: (2025)
ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning
von: Hou, Bairu, et al.
Veröffentlicht: (2025)
von: Hou, Bairu, et al.
Veröffentlicht: (2025)
Children's Intelligence Tests Pose Challenges for MLLMs? KidGym: A 2D Grid-Based Reasoning Benchmark for MLLMs
von: Ye, Hengwei, et al.
Veröffentlicht: (2026)
von: Ye, Hengwei, et al.
Veröffentlicht: (2026)
Fine-tuning MLLMs Without Forgetting Is Easier Than You Think
von: Li, He, et al.
Veröffentlicht: (2026)
von: Li, He, et al.
Veröffentlicht: (2026)
SAIL-RL: Guiding MLLMs in When and How to Think via Dual-Reward RL Tuning
von: Shu, Fangxun, et al.
Veröffentlicht: (2025)
von: Shu, Fangxun, et al.
Veröffentlicht: (2025)
SafePro: Evaluating the Safety of Professional-Level AI Agents
von: Zhou, Kaiwen, et al.
Veröffentlicht: (2026)
von: Zhou, Kaiwen, et al.
Veröffentlicht: (2026)
Mojito: Motion Trajectory and Intensity Control for Video Generation
von: He, Xuehai, et al.
Veröffentlicht: (2024)
von: He, Xuehai, et al.
Veröffentlicht: (2024)
Teaching Language Models to Think in Code
von: Hwang, Hyeon, et al.
Veröffentlicht: (2026)
von: Hwang, Hyeon, et al.
Veröffentlicht: (2026)
Tables as Texts or Images: Evaluating the Table Reasoning Ability of LLMs and MLLMs
von: Deng, Naihao, et al.
Veröffentlicht: (2024)
von: Deng, Naihao, et al.
Veröffentlicht: (2024)
Knowing You Don't Know: Learning When to Continue Search in Multi-round RAG through Self-Practicing
von: Yang, Diji, et al.
Veröffentlicht: (2025)
von: Yang, Diji, et al.
Veröffentlicht: (2025)
Distill Visual Chart Reasoning Ability from LLMs to MLLMs
von: He, Wei, et al.
Veröffentlicht: (2024)
von: He, Wei, et al.
Veröffentlicht: (2024)
Do MLLMs Really Understand the Charts?
von: Zhang, Xiao, et al.
Veröffentlicht: (2025)
von: Zhang, Xiao, et al.
Veröffentlicht: (2025)
XITE: Cross-lingual Interpolation for Transfer using Embeddings
von: Fazili, Barah, et al.
Veröffentlicht: (2026)
von: Fazili, Barah, et al.
Veröffentlicht: (2026)
See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles
von: Wu, Zongru, et al.
Veröffentlicht: (2025)
von: Wu, Zongru, et al.
Veröffentlicht: (2025)
Mitigating Object Hallucinations in MLLMs via Multi-Frequency Perturbations
von: Li, Shuo, et al.
Veröffentlicht: (2025)
von: Li, Shuo, et al.
Veröffentlicht: (2025)
Unhackable Temporal Rewarding for Scalable Video MLLMs
von: Yu, En, et al.
Veröffentlicht: (2025)
von: Yu, En, et al.
Veröffentlicht: (2025)
Chatting with Images for Introspective Visual Thinking
von: Wu, Junfei, et al.
Veröffentlicht: (2026)
von: Wu, Junfei, et al.
Veröffentlicht: (2026)
COPO: Causal-Oriented Policy Optimization for Hallucinations of MLLMs
von: Guo, Peizheng, et al.
Veröffentlicht: (2025)
von: Guo, Peizheng, et al.
Veröffentlicht: (2025)
O1 Embedder: Let Retrievers Think Before Action
von: Yan, Ruiran, et al.
Veröffentlicht: (2025)
von: Yan, Ruiran, et al.
Veröffentlicht: (2025)
Fast Thinking for Large Language Models
von: Zheng, Haoyu, et al.
Veröffentlicht: (2025)
von: Zheng, Haoyu, et al.
Veröffentlicht: (2025)
rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking
von: Guan, Xinyu, et al.
Veröffentlicht: (2025)
von: Guan, Xinyu, et al.
Veröffentlicht: (2025)
Teaching Thinking Models to Reason with Tools: A Full-Pipeline Recipe for Tool-Integrated Reasoning
von: Cheng, Qianjia, et al.
Veröffentlicht: (2026)
von: Cheng, Qianjia, et al.
Veröffentlicht: (2026)
MLLMs-Augmented Visual-Language Representation Learning
von: Liu, Yanqing, et al.
Veröffentlicht: (2023)
von: Liu, Yanqing, et al.
Veröffentlicht: (2023)
Interpretable and Reliable Detection of AI-Generated Images via Grounded Reasoning in MLLMs
von: Ji, Yikun, et al.
Veröffentlicht: (2025)
von: Ji, Yikun, et al.
Veröffentlicht: (2025)
Graft: Integrating the Domain Knowledge via Efficient Parameter Synergy for MLLMs
von: Dai, Yang, et al.
Veröffentlicht: (2025)
von: Dai, Yang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
JARVIS: A Neuro-Symbolic Commonsense Reasoning Framework for Conversational Embodied Agents
von: Zheng, Kaizhi, et al.
Veröffentlicht: (2022) -
MiniGPT-5: Interleaved Vision-and-Language Generation via Generative Vokens
von: Zheng, Kaizhi, et al.
Veröffentlicht: (2023) -
MorphoSim: An Interactive, Controllable, and Editable Language-guided 4D World Simulator
von: He, Xuehai, et al.
Veröffentlicht: (2025) -
Multimodal Inconsistency Reasoning (MMIR): A New Benchmark for Multimodal Reasoning Models
von: Yan, Qianqi, et al.
Veröffentlicht: (2025) -
Read Anywhere Pointed: Layout-aware GUI Screen Reading with Tree-of-Lens Grounding
von: Fan, Yue, et al.
Veröffentlicht: (2024)