ThinkGen: Generalized Thinking for Visual Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Jiao, Siyu, Lin, Yiheng, Zhong, Yujie, She, Qi, Zhou, Wei, Lan, Xiaohan, Huang, Zilong, Yu, Fei, Yu, Yingchen, Zhao, Yunqing, Zhao, Yao, Wei, Yunchao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Let ViT Speak: Generative Language-Image Pre-training
by: Fang, Yan, et al.
Published: (2026)
by: Fang, Yan, et al.
Published: (2026)
TextSculptor: Training and Benchmarking Scene Text Editing
by: Lin, Yiheng, et al.
Published: (2026)
by: Lin, Yiheng, et al.
Published: (2026)
CodeDance: A Dynamic Tool-integrated MLLM for Executable Visual Reasoning
by: Song, Qi, et al.
Published: (2025)
by: Song, Qi, et al.
Published: (2025)
AlignGen: Boosting Personalized Image Generation with Cross-Modality Prior Alignment
by: Lin, Yiheng, et al.
Published: (2025)
by: Lin, Yiheng, et al.
Published: (2025)
Collaborative Feature-Logits Contrastive Learning for Open-Set Semi-Supervised Object Detection
by: Zhong, Xinhao, et al.
Published: (2024)
by: Zhong, Xinhao, et al.
Published: (2024)
OmniAD: Detect and Understand Industrial Anomaly via Multimodal Reasoning
by: Zhao, Shifang, et al.
Published: (2025)
by: Zhao, Shifang, et al.
Published: (2025)
On the Faithfulness of Visual Thinking: Measurement and Enhancement
by: Liu, Zujing, et al.
Published: (2025)
by: Liu, Zujing, et al.
Published: (2025)
DreamLCM: Towards High-Quality Text-to-3D Generation via Latent Consistency Model
by: Zhong, Yiming, et al.
Published: (2024)
by: Zhong, Yiming, et al.
Published: (2024)
Thinking LLMs: General Instruction Following with Thought Generation
by: Wu, Tianhao, et al.
Published: (2024)
by: Wu, Tianhao, et al.
Published: (2024)
Diffusion for Natural Image Matting
by: Hu, Yihan, et al.
Published: (2023)
by: Hu, Yihan, et al.
Published: (2023)
Think or Not Think: A Study of Explicit Thinking in Rule-Based Visual Reinforcement Fine-Tuning
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
Text4Seg++: Advancing Image Segmentation via Generative Language Modeling
by: Lan, Mengcheng, et al.
Published: (2025)
by: Lan, Mengcheng, et al.
Published: (2025)
TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding
by: Yang, Zuhao, et al.
Published: (2025)
by: Yang, Zuhao, et al.
Published: (2025)
ThinkRec: Thinking-based recommendation via LLM
by: Yu, Qihang, et al.
Published: (2025)
by: Yu, Qihang, et al.
Published: (2025)
Think-J: Learning to Think for Generative LLM-as-a-Judge
by: Huang, Hui, et al.
Published: (2025)
by: Huang, Hui, et al.
Published: (2025)
Thinking, Faithful and Stable: Mitigating Hallucinations in LLMs
by: Zou, Chelsea, et al.
Published: (2025)
by: Zou, Chelsea, et al.
Published: (2025)
Collaborative Vision-Text Representation Optimizing for Open-Vocabulary Segmentation
by: Jiao, Siyu, et al.
Published: (2024)
by: Jiao, Siyu, et al.
Published: (2024)
FlexVAR: Flexible Visual Autoregressive Modeling without Residual Prediction
by: Jiao, Siyu, et al.
Published: (2025)
by: Jiao, Siyu, et al.
Published: (2025)
PreFM: Online Audio-Visual Event Parsing via Predictive Future Modeling
by: Yu, Xiao, et al.
Published: (2025)
by: Yu, Xiao, et al.
Published: (2025)
ThinkPilot: Steering Reasoning Models via Automated Think-prefixes Optimization
by: Li, Sunzhu, et al.
Published: (2025)
by: Li, Sunzhu, et al.
Published: (2025)
Think Deep, Not Just Long: Measuring LLM Reasoning Effort via Deep-Thinking Tokens
by: Chen, Wei-Lin, et al.
Published: (2026)
by: Chen, Wei-Lin, et al.
Published: (2026)
CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder
by: Ma, Lichen, et al.
Published: (2024)
by: Ma, Lichen, et al.
Published: (2024)
KAG-Thinker: Interactive Thinking and Deep Reasoning in LLMs via Knowledge-Augmented Generation
by: Zhang, Dalong, et al.
Published: (2025)
by: Zhang, Dalong, et al.
Published: (2025)
IPSeg: Image Posterior Mitigates Semantic Drift in Class-Incremental Segmentation
by: Yu, Xiao, et al.
Published: (2025)
by: Yu, Xiao, et al.
Published: (2025)
Manga Generation via Layout-controllable Diffusion
by: Chen, Siyu, et al.
Published: (2024)
by: Chen, Siyu, et al.
Published: (2024)
Think Anywhere in Code Generation
by: Jiang, Xue, et al.
Published: (2026)
by: Jiang, Xue, et al.
Published: (2026)
ConFoThinking: Consolidated Focused Attention Driven Thinking for Visual Question Answering
by: Wu, Zhaodong, et al.
Published: (2026)
by: Wu, Zhaodong, et al.
Published: (2026)
SAGE: Exploring the Boundaries of Unsafe Concept Domain with Semantic-Augment Erasing
by: Zhu, Hongguang, et al.
Published: (2025)
by: Zhu, Hongguang, et al.
Published: (2025)
ThinkDrive: Chain-of-Thought Guided Progressive Reinforcement Learning Fine-Tuning for Autonomous Driving
by: Zhao, Chang, et al.
Published: (2026)
by: Zhao, Chang, et al.
Published: (2026)
Synergizing Understanding and Generation with Interleaved Analyzing-Drafting Thinking
by: Wu, Shengqiong, et al.
Published: (2026)
by: Wu, Shengqiong, et al.
Published: (2026)
Semantic Visual Anomaly Detection and Reasoning in AI-Generated Images
by: Tan, Chuangchuang, et al.
Published: (2025)
by: Tan, Chuangchuang, et al.
Published: (2025)
M4V: Multi-Modal Mamba for Text-to-Video Generation
by: Huang, Jiancheng, et al.
Published: (2025)
by: Huang, Jiancheng, et al.
Published: (2025)
4DGen: Grounded 4D Content Generation with Spatial-temporal Consistency
by: Yin, Yuyang, et al.
Published: (2023)
by: Yin, Yuyang, et al.
Published: (2023)
Extrapolating and Decoupling Image-to-Video Generation Models: Motion Modeling is Easier Than You Think
by: Tian, Jie, et al.
Published: (2025)
by: Tian, Jie, et al.
Published: (2025)
VisionCreator: A Native Visual-Generation Agentic Model with Understanding, Thinking, Planning and Creation
by: Lai, Jinxiang, et al.
Published: (2026)
by: Lai, Jinxiang, et al.
Published: (2026)
ThinkSound: Chain-of-Thought Reasoning in Multimodal Large Language Models for Audio Generation and Editing
by: Liu, Huadai, et al.
Published: (2025)
by: Liu, Huadai, et al.
Published: (2025)
Textual-to-Visual Iterative Self-Verification for Slide Generation
by: Xu, Yunqing, et al.
Published: (2025)
by: Xu, Yunqing, et al.
Published: (2025)
Bridging Performance Gaps for ECG Foundation Models: A Post-Training Strategy
by: Zhou, Ya, et al.
Published: (2025)
by: Zhou, Ya, et al.
Published: (2025)
Hume: Introducing System-2 Thinking in Visual-Language-Action Model
by: Song, Haoming, et al.
Published: (2025)
by: Song, Haoming, et al.
Published: (2025)
Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
Similar Items
-
Let ViT Speak: Generative Language-Image Pre-training
by: Fang, Yan, et al.
Published: (2026) -
TextSculptor: Training and Benchmarking Scene Text Editing
by: Lin, Yiheng, et al.
Published: (2026) -
CodeDance: A Dynamic Tool-integrated MLLM for Executable Visual Reasoning
by: Song, Qi, et al.
Published: (2025) -
AlignGen: Boosting Personalized Image Generation with Cross-Modality Prior Alignment
by: Lin, Yiheng, et al.
Published: (2025) -
Collaborative Feature-Logits Contrastive Learning for Open-Set Semi-Supervised Object Detection
by: Zhong, Xinhao, et al.
Published: (2024)