Thinking with Programming Vision: Towards a Unified View for Thinking with Images
Fuente:
arXiv
Salvato in:
| Autori principali: | Guo, Zirun, Hong, Minjie, Zhang, Feng, Jia, Kai, Jin, Tao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Multimodal Prompt Learning with Missing Modalities for Sentiment Analysis and Emotion Recognition
di: Guo, Zirun, et al.
Pubblicazione: (2024)
di: Guo, Zirun, et al.
Pubblicazione: (2024)
Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning
di: Guo, Zirun, et al.
Pubblicazione: (2025)
di: Guo, Zirun, et al.
Pubblicazione: (2025)
ViewFusion: Structured Spatial Thinking Chains for Multi-View Reasoning
di: Tao, Xingjian, et al.
Pubblicazione: (2026)
di: Tao, Xingjian, et al.
Pubblicazione: (2026)
EgoThink: Evaluating First-Person Perspective Thinking Capability of Vision-Language Models
di: Cheng, Sijie, et al.
Pubblicazione: (2023)
di: Cheng, Sijie, et al.
Pubblicazione: (2023)
LLM-I: LLMs are Naturally Interleaved Multimodal Creators
di: Guo, Zirun, et al.
Pubblicazione: (2025)
di: Guo, Zirun, et al.
Pubblicazione: (2025)
Classifier-guided Gradient Modulation for Enhanced Multimodal Learning
di: Guo, Zirun, et al.
Pubblicazione: (2024)
di: Guo, Zirun, et al.
Pubblicazione: (2024)
Smoothing the Shift: Towards Stable Test-Time Adaptation under Complex Multimodal Noises
di: Guo, Zirun, et al.
Pubblicazione: (2025)
di: Guo, Zirun, et al.
Pubblicazione: (2025)
Thinking with Generated Images
di: Chern, Ethan, et al.
Pubblicazione: (2025)
di: Chern, Ethan, et al.
Pubblicazione: (2025)
ConceptGuard: Continual Personalized Text-to-Image Generation with Forgetting and Confusion Mitigation
di: Guo, Zirun, et al.
Pubblicazione: (2025)
di: Guo, Zirun, et al.
Pubblicazione: (2025)
Look Again, Think Slowly: Enhancing Visual Reflection in Vision-Language Models
di: Jian, Pu, et al.
Pubblicazione: (2025)
di: Jian, Pu, et al.
Pubblicazione: (2025)
Dolphin: Moving Towards Closed-loop Auto-research through Thinking, Practice, and Feedback
di: Yuan, Jiakang, et al.
Pubblicazione: (2025)
di: Yuan, Jiakang, et al.
Pubblicazione: (2025)
AgentThink: A Unified Framework for Tool-Augmented Chain-of-Thought Reasoning in Vision-Language Models for Autonomous Driving
di: Qian, Kangan, et al.
Pubblicazione: (2025)
di: Qian, Kangan, et al.
Pubblicazione: (2025)
Re-Thinking Inverse Graphics With Large Language Models
di: Kulits, Peter, et al.
Pubblicazione: (2024)
di: Kulits, Peter, et al.
Pubblicazione: (2024)
GRIT: Teaching MLLMs to Think with Images
di: Fan, Yue, et al.
Pubblicazione: (2025)
di: Fan, Yue, et al.
Pubblicazione: (2025)
Chatting with Images for Introspective Visual Thinking
di: Wu, Junfei, et al.
Pubblicazione: (2026)
di: Wu, Junfei, et al.
Pubblicazione: (2026)
When to Think and When to Look: Uncertainty-Guided Lookback
di: Bi, Jing, et al.
Pubblicazione: (2025)
di: Bi, Jing, et al.
Pubblicazione: (2025)
Think Visually, Reason Textually: Vision-Language Synergy in ARC
di: Zhang, Beichen, et al.
Pubblicazione: (2025)
di: Zhang, Beichen, et al.
Pubblicazione: (2025)
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models
di: Wu, Juncheng, et al.
Pubblicazione: (2026)
di: Wu, Juncheng, et al.
Pubblicazione: (2026)
Re-Thinking the Automatic Evaluation of Image-Text Alignment in Text-to-Image Models
di: Zhang, Huixuan, et al.
Pubblicazione: (2025)
di: Zhang, Huixuan, et al.
Pubblicazione: (2025)
VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning
di: Yang, Senqiao, et al.
Pubblicazione: (2025)
di: Yang, Senqiao, et al.
Pubblicazione: (2025)
Multimodal Representation Alignment for Image Generation: Text-Image Interleaved Control Is Easier Than You Think
di: Chen, Liang, et al.
Pubblicazione: (2025)
di: Chen, Liang, et al.
Pubblicazione: (2025)
Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation
di: Guo, Ziyu, et al.
Pubblicazione: (2025)
di: Guo, Ziyu, et al.
Pubblicazione: (2025)
Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models
di: Li, Yunxin, et al.
Pubblicazione: (2025)
di: Li, Yunxin, et al.
Pubblicazione: (2025)
Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm
di: Tong, Jingqi, et al.
Pubblicazione: (2025)
di: Tong, Jingqi, et al.
Pubblicazione: (2025)
Uni-cot: Towards Unified Chain-of-Thought Reasoning Across Text and Vision
di: Qin, Luozheng, et al.
Pubblicazione: (2025)
di: Qin, Luozheng, et al.
Pubblicazione: (2025)
Think-Reflect-Revise: A Policy-Guided Reflective Framework for Safety Alignment in Large Vision Language Models
di: Weng, Fenghua, et al.
Pubblicazione: (2025)
di: Weng, Fenghua, et al.
Pubblicazione: (2025)
Fast-Slow Thinking GRPO for Large Vision-Language Model Reasoning
di: Xiao, Wenyi, et al.
Pubblicazione: (2025)
di: Xiao, Wenyi, et al.
Pubblicazione: (2025)
Seeing but Not Thinking: Routing Distraction in Multimodal Mixture-of-Experts
di: Xu, Haolei, et al.
Pubblicazione: (2026)
di: Xu, Haolei, et al.
Pubblicazione: (2026)
Visual Planning: Let's Think Only with Images
di: Xu, Yi, et al.
Pubblicazione: (2025)
di: Xu, Yi, et al.
Pubblicazione: (2025)
Draw2Think: Harnessing Geometry Reasoning through Constraint Engine Interaction
di: Hu, Juncheng, et al.
Pubblicazione: (2026)
di: Hu, Juncheng, et al.
Pubblicazione: (2026)
See, Think, Learn: A Self-Taught Multimodal Reasoner
di: Sharma, Sourabh, et al.
Pubblicazione: (2025)
di: Sharma, Sourabh, et al.
Pubblicazione: (2025)
RoboUniView: Visual-Language Model with Unified View Representation for Robotic Manipulation
di: Liu, Fanfan, et al.
Pubblicazione: (2024)
di: Liu, Fanfan, et al.
Pubblicazione: (2024)
SAIL-RL: Guiding MLLMs in When and How to Think via Dual-Reward RL Tuning
di: Shu, Fangxun, et al.
Pubblicazione: (2025)
di: Shu, Fangxun, et al.
Pubblicazione: (2025)
Reading, Not Thinking: Understanding and Bridging the Modality Gap When Text Becomes Pixels in Multimodal LLMs
di: Sun, Kaiser, et al.
Pubblicazione: (2026)
di: Sun, Kaiser, et al.
Pubblicazione: (2026)
From Perception to Reasoning: Deep Thinking Empowers Multimodal Large Language Models
di: Zhu, Wenxin, et al.
Pubblicazione: (2025)
di: Zhu, Wenxin, et al.
Pubblicazione: (2025)
Low-rank Prompt Interaction for Continual Vision-Language Retrieval
di: Yan, Weicai, et al.
Pubblicazione: (2025)
di: Yan, Weicai, et al.
Pubblicazione: (2025)
VideoScore2: Think before You Score in Generative Video Evaluation
di: He, Xuan, et al.
Pubblicazione: (2025)
di: He, Xuan, et al.
Pubblicazione: (2025)
3ViewSense: Spatial and Mental Perspective Reasoning from Orthographic Views in Vision-Language Models
di: Zhan, Shaoxiong, et al.
Pubblicazione: (2026)
di: Zhan, Shaoxiong, et al.
Pubblicazione: (2026)
Unifying Vision-Language Latents for Zero-label Image Caption Enhancement
di: Byun, Sanghyun, et al.
Pubblicazione: (2025)
di: Byun, Sanghyun, et al.
Pubblicazione: (2025)
ThinkingViT: Matryoshka Thinking Vision Transformer for Elastic Inference
di: Hojjat, Ali, et al.
Pubblicazione: (2025)
di: Hojjat, Ali, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Multimodal Prompt Learning with Missing Modalities for Sentiment Analysis and Emotion Recognition
di: Guo, Zirun, et al.
Pubblicazione: (2024) -
Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning
di: Guo, Zirun, et al.
Pubblicazione: (2025) -
ViewFusion: Structured Spatial Thinking Chains for Multi-View Reasoning
di: Tao, Xingjian, et al.
Pubblicazione: (2026) -
EgoThink: Evaluating First-Person Perspective Thinking Capability of Vision-Language Models
di: Cheng, Sijie, et al.
Pubblicazione: (2023) -
LLM-I: LLMs are Naturally Interleaved Multimodal Creators
di: Guo, Zirun, et al.
Pubblicazione: (2025)