Thinking with Generated Images
Fuente:
arXiv
Saved in:
| Main Authors: | Chern, Ethan, Hu, Zhulin, Chern, Steffi, Kou, Siqi, Su, Jiadi, Ma, Yan, Deng, Zhijie, Liu, Pengfei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation
by: Chern, Ethan, et al.
Published: (2024)
by: Chern, Ethan, et al.
Published: (2024)
LiveTalk: Real-Time Multimodal Interactive Video Diffusion via Improved On-Policy Distillation
by: Chern, Ethan, et al.
Published: (2025)
by: Chern, Ethan, et al.
Published: (2025)
Halu-J: Critique-Based Hallucination Judge
by: Wang, Binjie, et al.
Published: (2024)
by: Wang, Binjie, et al.
Published: (2024)
Can Large Language Models be Trusted for Evaluation? Scalable Meta-Evaluation of LLMs as Evaluators via Agent Debate
by: Chern, Steffi, et al.
Published: (2024)
by: Chern, Steffi, et al.
Published: (2024)
BeHonest: Benchmarking Honesty in Large Language Models
by: Chern, Steffi, et al.
Published: (2024)
by: Chern, Steffi, et al.
Published: (2024)
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads
by: Kou, Siqi, et al.
Published: (2024)
by: Kou, Siqi, et al.
Published: (2024)
Rethinking RL Scaling for Vision Language Models: A Transparent, From-Scratch Framework and Comprehensive Evaluation Scheme
by: Ma, Yan, et al.
Published: (2025)
by: Ma, Yan, et al.
Published: (2025)
Visual Programmability: A Guide for Code-as-Thought in Chart Understanding
by: Tang, Bohao, et al.
Published: (2025)
by: Tang, Bohao, et al.
Published: (2025)
Combating Adversarial Attacks with Multi-Agent Debate
by: Chern, Steffi, et al.
Published: (2024)
by: Chern, Steffi, et al.
Published: (2024)
ProductWebGen: Benchmarking Multimodal Product Webpage Generation
by: Liu, Zhihong, et al.
Published: (2026)
by: Liu, Zhihong, et al.
Published: (2026)
When VLMs Meet Image Classification: Test Sets Renovation via Missing Label Identification
by: Pang, Zirui, et al.
Published: (2025)
by: Pang, Zirui, et al.
Published: (2025)
Chatting with Images for Introspective Visual Thinking
by: Wu, Junfei, et al.
Published: (2026)
by: Wu, Junfei, et al.
Published: (2026)
Mantis: A Versatile Vision-Language-Action Model with Disentangled Visual Foresight
by: Yang, Yi, et al.
Published: (2025)
by: Yang, Yi, et al.
Published: (2025)
Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
GRIT: Teaching MLLMs to Think with Images
by: Fan, Yue, et al.
Published: (2025)
by: Fan, Yue, et al.
Published: (2025)
Thinking with Map: Reinforced Parallel Map-Augmented Agent for Geolocalization
by: Ji, Yuxiang, et al.
Published: (2026)
by: Ji, Yuxiang, et al.
Published: (2026)
VideoScore2: Think before You Score in Generative Video Evaluation
by: He, Xuan, et al.
Published: (2025)
by: He, Xuan, et al.
Published: (2025)
Re-Thinking the Automatic Evaluation of Image-Text Alignment in Text-to-Image Models
by: Zhang, Huixuan, et al.
Published: (2025)
by: Zhang, Huixuan, et al.
Published: (2025)
Think Twice, Click Once: Enhancing GUI Grounding via Fast and Slow Systems
by: Tang, Fei, et al.
Published: (2025)
by: Tang, Fei, et al.
Published: (2025)
VideoExplorer: Think With Videos For Agentic Long-Video Understanding
by: Yuan, Huaying, et al.
Published: (2025)
by: Yuan, Huaying, et al.
Published: (2025)
Instruct-Imagen: Image Generation with Multi-modal Instruction
by: Hu, Hexiang, et al.
Published: (2024)
by: Hu, Hexiang, et al.
Published: (2024)
Revisiting the Data Sampling in Multimodal Post-training from a Difficulty-Distinguish View
by: Qi, Jianyu, et al.
Published: (2025)
by: Qi, Jianyu, et al.
Published: (2025)
Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation
by: Ye, Junyan, et al.
Published: (2025)
by: Ye, Junyan, et al.
Published: (2025)
Unveiling Uncertainty: A Deep Dive into Calibration and Performance of Multimodal Large Language Models
by: Chen, Zijun, et al.
Published: (2024)
by: Chen, Zijun, et al.
Published: (2024)
Enhancing Autonomous Vehicle Perception in Adverse Weather through Image Augmentation during Semantic Segmentation Training
by: Kou, Ethan, et al.
Published: (2024)
by: Kou, Ethan, et al.
Published: (2024)
More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models
by: Liu, Chengzhi, et al.
Published: (2025)
by: Liu, Chengzhi, et al.
Published: (2025)
OmniGen2: Towards Instruction-Aligned Multimodal Generation
by: Wu, Chenyuan, et al.
Published: (2025)
by: Wu, Chenyuan, et al.
Published: (2025)
Let's Think Outside the Box: Exploring Leap-of-Thought in Large Language Models with Creative Humor Generation
by: Zhong, Shanshan, et al.
Published: (2023)
by: Zhong, Shanshan, et al.
Published: (2023)
Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction
by: Wang, Yuchi, et al.
Published: (2025)
by: Wang, Yuchi, et al.
Published: (2025)
Dolphin: Moving Towards Closed-loop Auto-research through Thinking, Practice, and Feedback
by: Yuan, Jiakang, et al.
Published: (2025)
by: Yuan, Jiakang, et al.
Published: (2025)
Holistic Evaluation for Interleaved Text-and-Image Generation
by: Liu, Minqian, et al.
Published: (2024)
by: Liu, Minqian, et al.
Published: (2024)
Visual Planning: Let's Think Only with Images
by: Xu, Yi, et al.
Published: (2025)
by: Xu, Yi, et al.
Published: (2025)
MINOS: A Multimodal Evaluation Model for Bidirectional Generation Between Image and Text
by: Zhang, Junzhe, et al.
Published: (2025)
by: Zhang, Junzhe, et al.
Published: (2025)
VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI
by: Cheng, Sijie, et al.
Published: (2024)
by: Cheng, Sijie, et al.
Published: (2024)
Interpretable and Reliable Detection of AI-Generated Images via Grounded Reasoning in MLLMs
by: Ji, Yikun, et al.
Published: (2025)
by: Ji, Yikun, et al.
Published: (2025)
Seeing but Not Thinking: Routing Distraction in Multimodal Mixture-of-Experts
by: Xu, Haolei, et al.
Published: (2026)
by: Xu, Haolei, et al.
Published: (2026)
Whiteboard-of-Thought: Thinking Step-by-Step Across Modalities
by: Menon, Sachit, et al.
Published: (2024)
by: Menon, Sachit, et al.
Published: (2024)
Alignment for Honesty
by: Yang, Yuqing, et al.
Published: (2023)
by: Yang, Yuqing, et al.
Published: (2023)
Think While You Generate: Discrete Diffusion with Planned Denoising
by: Liu, Sulin, et al.
Published: (2024)
by: Liu, Sulin, et al.
Published: (2024)
Similar Items
-
ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation
by: Chern, Ethan, et al.
Published: (2024) -
LiveTalk: Real-Time Multimodal Interactive Video Diffusion via Improved On-Policy Distillation
by: Chern, Ethan, et al.
Published: (2025) -
Halu-J: Critique-Based Hallucination Judge
by: Wang, Binjie, et al.
Published: (2024) -
Can Large Language Models be Trusted for Evaluation? Scalable Meta-Evaluation of LLMs as Evaluators via Agent Debate
by: Chern, Steffi, et al.
Published: (2024) -
BeHonest: Benchmarking Honesty in Large Language Models
by: Chern, Steffi, et al.
Published: (2024)