Thinking with Programming Vision: Towards a Unified View for Thinking with Images
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Guo, Zirun, Hong, Minjie, Zhang, Feng, Jia, Kai, Jin, Tao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multimodal Prompt Learning with Missing Modalities for Sentiment Analysis and Emotion Recognition
von: Guo, Zirun, et al.
Veröffentlicht: (2024)
von: Guo, Zirun, et al.
Veröffentlicht: (2024)
Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning
von: Guo, Zirun, et al.
Veröffentlicht: (2025)
von: Guo, Zirun, et al.
Veröffentlicht: (2025)
ViewFusion: Structured Spatial Thinking Chains for Multi-View Reasoning
von: Tao, Xingjian, et al.
Veröffentlicht: (2026)
von: Tao, Xingjian, et al.
Veröffentlicht: (2026)
EgoThink: Evaluating First-Person Perspective Thinking Capability of Vision-Language Models
von: Cheng, Sijie, et al.
Veröffentlicht: (2023)
von: Cheng, Sijie, et al.
Veröffentlicht: (2023)
LLM-I: LLMs are Naturally Interleaved Multimodal Creators
von: Guo, Zirun, et al.
Veröffentlicht: (2025)
von: Guo, Zirun, et al.
Veröffentlicht: (2025)
Classifier-guided Gradient Modulation for Enhanced Multimodal Learning
von: Guo, Zirun, et al.
Veröffentlicht: (2024)
von: Guo, Zirun, et al.
Veröffentlicht: (2024)
Smoothing the Shift: Towards Stable Test-Time Adaptation under Complex Multimodal Noises
von: Guo, Zirun, et al.
Veröffentlicht: (2025)
von: Guo, Zirun, et al.
Veröffentlicht: (2025)
Thinking with Generated Images
von: Chern, Ethan, et al.
Veröffentlicht: (2025)
von: Chern, Ethan, et al.
Veröffentlicht: (2025)
ConceptGuard: Continual Personalized Text-to-Image Generation with Forgetting and Confusion Mitigation
von: Guo, Zirun, et al.
Veröffentlicht: (2025)
von: Guo, Zirun, et al.
Veröffentlicht: (2025)
Look Again, Think Slowly: Enhancing Visual Reflection in Vision-Language Models
von: Jian, Pu, et al.
Veröffentlicht: (2025)
von: Jian, Pu, et al.
Veröffentlicht: (2025)
Dolphin: Moving Towards Closed-loop Auto-research through Thinking, Practice, and Feedback
von: Yuan, Jiakang, et al.
Veröffentlicht: (2025)
von: Yuan, Jiakang, et al.
Veröffentlicht: (2025)
AgentThink: A Unified Framework for Tool-Augmented Chain-of-Thought Reasoning in Vision-Language Models for Autonomous Driving
von: Qian, Kangan, et al.
Veröffentlicht: (2025)
von: Qian, Kangan, et al.
Veröffentlicht: (2025)
Re-Thinking Inverse Graphics With Large Language Models
von: Kulits, Peter, et al.
Veröffentlicht: (2024)
von: Kulits, Peter, et al.
Veröffentlicht: (2024)
GRIT: Teaching MLLMs to Think with Images
von: Fan, Yue, et al.
Veröffentlicht: (2025)
von: Fan, Yue, et al.
Veröffentlicht: (2025)
Chatting with Images for Introspective Visual Thinking
von: Wu, Junfei, et al.
Veröffentlicht: (2026)
von: Wu, Junfei, et al.
Veröffentlicht: (2026)
When to Think and When to Look: Uncertainty-Guided Lookback
von: Bi, Jing, et al.
Veröffentlicht: (2025)
von: Bi, Jing, et al.
Veröffentlicht: (2025)
Think Visually, Reason Textually: Vision-Language Synergy in ARC
von: Zhang, Beichen, et al.
Veröffentlicht: (2025)
von: Zhang, Beichen, et al.
Veröffentlicht: (2025)
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models
von: Wu, Juncheng, et al.
Veröffentlicht: (2026)
von: Wu, Juncheng, et al.
Veröffentlicht: (2026)
Re-Thinking the Automatic Evaluation of Image-Text Alignment in Text-to-Image Models
von: Zhang, Huixuan, et al.
Veröffentlicht: (2025)
von: Zhang, Huixuan, et al.
Veröffentlicht: (2025)
VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning
von: Yang, Senqiao, et al.
Veröffentlicht: (2025)
von: Yang, Senqiao, et al.
Veröffentlicht: (2025)
Multimodal Representation Alignment for Image Generation: Text-Image Interleaved Control Is Easier Than You Think
von: Chen, Liang, et al.
Veröffentlicht: (2025)
von: Chen, Liang, et al.
Veröffentlicht: (2025)
Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation
von: Guo, Ziyu, et al.
Veröffentlicht: (2025)
von: Guo, Ziyu, et al.
Veröffentlicht: (2025)
Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models
von: Li, Yunxin, et al.
Veröffentlicht: (2025)
von: Li, Yunxin, et al.
Veröffentlicht: (2025)
Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm
von: Tong, Jingqi, et al.
Veröffentlicht: (2025)
von: Tong, Jingqi, et al.
Veröffentlicht: (2025)
Uni-cot: Towards Unified Chain-of-Thought Reasoning Across Text and Vision
von: Qin, Luozheng, et al.
Veröffentlicht: (2025)
von: Qin, Luozheng, et al.
Veröffentlicht: (2025)
Think-Reflect-Revise: A Policy-Guided Reflective Framework for Safety Alignment in Large Vision Language Models
von: Weng, Fenghua, et al.
Veröffentlicht: (2025)
von: Weng, Fenghua, et al.
Veröffentlicht: (2025)
Fast-Slow Thinking GRPO for Large Vision-Language Model Reasoning
von: Xiao, Wenyi, et al.
Veröffentlicht: (2025)
von: Xiao, Wenyi, et al.
Veröffentlicht: (2025)
Seeing but Not Thinking: Routing Distraction in Multimodal Mixture-of-Experts
von: Xu, Haolei, et al.
Veröffentlicht: (2026)
von: Xu, Haolei, et al.
Veröffentlicht: (2026)
Visual Planning: Let's Think Only with Images
von: Xu, Yi, et al.
Veröffentlicht: (2025)
von: Xu, Yi, et al.
Veröffentlicht: (2025)
Draw2Think: Harnessing Geometry Reasoning through Constraint Engine Interaction
von: Hu, Juncheng, et al.
Veröffentlicht: (2026)
von: Hu, Juncheng, et al.
Veröffentlicht: (2026)
See, Think, Learn: A Self-Taught Multimodal Reasoner
von: Sharma, Sourabh, et al.
Veröffentlicht: (2025)
von: Sharma, Sourabh, et al.
Veröffentlicht: (2025)
RoboUniView: Visual-Language Model with Unified View Representation for Robotic Manipulation
von: Liu, Fanfan, et al.
Veröffentlicht: (2024)
von: Liu, Fanfan, et al.
Veröffentlicht: (2024)
SAIL-RL: Guiding MLLMs in When and How to Think via Dual-Reward RL Tuning
von: Shu, Fangxun, et al.
Veröffentlicht: (2025)
von: Shu, Fangxun, et al.
Veröffentlicht: (2025)
Reading, Not Thinking: Understanding and Bridging the Modality Gap When Text Becomes Pixels in Multimodal LLMs
von: Sun, Kaiser, et al.
Veröffentlicht: (2026)
von: Sun, Kaiser, et al.
Veröffentlicht: (2026)
From Perception to Reasoning: Deep Thinking Empowers Multimodal Large Language Models
von: Zhu, Wenxin, et al.
Veröffentlicht: (2025)
von: Zhu, Wenxin, et al.
Veröffentlicht: (2025)
Low-rank Prompt Interaction for Continual Vision-Language Retrieval
von: Yan, Weicai, et al.
Veröffentlicht: (2025)
von: Yan, Weicai, et al.
Veröffentlicht: (2025)
VideoScore2: Think before You Score in Generative Video Evaluation
von: He, Xuan, et al.
Veröffentlicht: (2025)
von: He, Xuan, et al.
Veröffentlicht: (2025)
3ViewSense: Spatial and Mental Perspective Reasoning from Orthographic Views in Vision-Language Models
von: Zhan, Shaoxiong, et al.
Veröffentlicht: (2026)
von: Zhan, Shaoxiong, et al.
Veröffentlicht: (2026)
Unifying Vision-Language Latents for Zero-label Image Caption Enhancement
von: Byun, Sanghyun, et al.
Veröffentlicht: (2025)
von: Byun, Sanghyun, et al.
Veröffentlicht: (2025)
ThinkingViT: Matryoshka Thinking Vision Transformer for Elastic Inference
von: Hojjat, Ali, et al.
Veröffentlicht: (2025)
von: Hojjat, Ali, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Multimodal Prompt Learning with Missing Modalities for Sentiment Analysis and Emotion Recognition
von: Guo, Zirun, et al.
Veröffentlicht: (2024) -
Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning
von: Guo, Zirun, et al.
Veröffentlicht: (2025) -
ViewFusion: Structured Spatial Thinking Chains for Multi-View Reasoning
von: Tao, Xingjian, et al.
Veröffentlicht: (2026) -
EgoThink: Evaluating First-Person Perspective Thinking Capability of Vision-Language Models
von: Cheng, Sijie, et al.
Veröffentlicht: (2023) -
LLM-I: LLMs are Naturally Interleaved Multimodal Creators
von: Guo, Zirun, et al.
Veröffentlicht: (2025)