Improving Chain-of-Thought Efficiency for Autoregressive Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Gu, Zeqi, Georgopoulos, Markos, Dai, Xiaoliang, Ghazvininejad, Marjan, Wang, Chu, Juefei-Xu, Felix, Li, Kunpeng, Shi, Yujun, He, Zecheng, He, Zijian, Zhou, Jiawei, Davis, Abe, Wang, Jialiang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UniT: Unified Multimodal Chain-of-Thought Test-time Scaling
by: Chen, Leon Liangyu, et al.
Published: (2026)
by: Chen, Leon Liangyu, et al.
Published: (2026)
Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning
by: Zhang, Lei, et al.
Published: (2026)
by: Zhang, Lei, et al.
Published: (2026)
Token-Shuffle: Towards High-Resolution Image Generation with Autoregressive Models
by: Ma, Xu, et al.
Published: (2025)
by: Ma, Xu, et al.
Published: (2025)
Llama Learns to Direct: DirectorLLM for Human-Centric Video Generation
by: Song, Kunpeng, et al.
Published: (2024)
by: Song, Kunpeng, et al.
Published: (2024)
Non-Markov Multi-Round Conversational Image Generation with History-Conditioned MLLMs
by: Zhang, Haochen, et al.
Published: (2026)
by: Zhang, Haochen, et al.
Published: (2026)
Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models
by: Yasunaga, Michihiro, et al.
Published: (2025)
by: Yasunaga, Michihiro, et al.
Published: (2025)
MoCha: Towards Movie-Grade Talking Character Synthesis
by: Wei, Cong, et al.
Published: (2025)
by: Wei, Cong, et al.
Published: (2025)
David helps Goliath: Inference-Time Collaboration Between Small Specialized and Large General Diffusion LMs
by: Han, Xiaochuang, et al.
Published: (2023)
by: Han, Xiaochuang, et al.
Published: (2023)
PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation
by: Cai, Yuanhao, et al.
Published: (2025)
by: Cai, Yuanhao, et al.
Published: (2025)
Movie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored Prompts
by: Liang, Feng, et al.
Published: (2025)
by: Liang, Feng, et al.
Published: (2025)
JPEG-LM: LLMs as Image Generators with Canonical Codec Representations
by: Han, Xiaochuang, et al.
Published: (2024)
by: Han, Xiaochuang, et al.
Published: (2024)
Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge
by: Saha, Swarnadeep, et al.
Published: (2025)
by: Saha, Swarnadeep, et al.
Published: (2025)
Filter-Guided Diffusion for Controllable Image Generation
by: Gu, Zeqi, et al.
Published: (2023)
by: Gu, Zeqi, et al.
Published: (2023)
Evolving Demonstration Optimization for Chain-of-Thought Feature Transformation
by: Wang, Xinyuan, et al.
Published: (2026)
by: Wang, Xinyuan, et al.
Published: (2026)
StreamDiT: Real-Time Streaming Text-to-Video Generation
by: Kodaira, Akio, et al.
Published: (2025)
by: Kodaira, Akio, et al.
Published: (2025)
Unleashing In-context Learning of Autoregressive Models for Few-shot Image Manipulation
by: Lai, Bolin, et al.
Published: (2024)
by: Lai, Bolin, et al.
Published: (2024)
Eliciting Chain-of-Thought in Base LLMs via Gradient-Based Representation Optimization
by: Wang, Zijian, et al.
Published: (2025)
by: Wang, Zijian, et al.
Published: (2025)
Pixel-Space Post-Training of Latent Diffusion Models
by: Zhang, Christina, et al.
Published: (2024)
by: Zhang, Christina, et al.
Published: (2024)
Exploring MLLM-Diffusion Information Transfer with MetaCanvas
by: Lin, Han, et al.
Published: (2025)
by: Lin, Han, et al.
Published: (2025)
Representation Deficiency in Masked Language Modeling
by: Meng, Yu, et al.
Published: (2023)
by: Meng, Yu, et al.
Published: (2023)
Autoregressive Distillation of Diffusion Transformers
by: Kim, Yeongmin, et al.
Published: (2025)
by: Kim, Yeongmin, et al.
Published: (2025)
Populate-A-Scene: Affordance-Aware Human Video Generation
by: Shan, Mengyi, et al.
Published: (2025)
by: Shan, Mengyi, et al.
Published: (2025)
A Theory of Learning with Autoregressive Chain of Thought
by: Joshi, Nirmit, et al.
Published: (2025)
by: Joshi, Nirmit, et al.
Published: (2025)
ACIL: Auto Chain of Thoughts for In-Context Learning
by: Chu, Rui
Published: (2026)
by: Chu, Rui
Published: (2026)
Multimodal RewardBench 2: Evaluating Omni Reward Models for Interleaved Text and Image
by: Hu, Yushi, et al.
Published: (2025)
by: Hu, Yushi, et al.
Published: (2025)
reWordBench: Benchmarking and Improving the Robustness of Reward Models with Transformed Inputs
by: Wu, Zhaofeng, et al.
Published: (2025)
by: Wu, Zhaofeng, et al.
Published: (2025)
An Analysis on Quantizing Diffusion Transformers
by: Yang, Yuewei, et al.
Published: (2024)
by: Yang, Yuewei, et al.
Published: (2024)
ThoughtProbe: Classifier-Guided Thought Space Exploration Leveraging LLM Intrinsic Reasoning
by: Wang, Zijian, et al.
Published: (2025)
by: Wang, Zijian, et al.
Published: (2025)
ThoughtProbe: Classifier-Guided LLM Thought Space Exploration via Probing Representations
by: Wang, Zijian, et al.
Published: (2025)
by: Wang, Zijian, et al.
Published: (2025)
Pretraining with Token-Level Adaptive Latent Chain-of-Thought
by: Zeng, Boyi, et al.
Published: (2026)
by: Zeng, Boyi, et al.
Published: (2026)
RecycleGPT: An Autoregressive Language Model with Recyclable Module
by: Jiang, Yufan, et al.
Published: (2023)
by: Jiang, Yufan, et al.
Published: (2023)
Thinking with Sound: Audio Chain-of-Thought Enables Multimodal Reasoning in Large Audio-Language Models
by: Xiong, Zhen, et al.
Published: (2025)
by: Xiong, Zhen, et al.
Published: (2025)
How to Train Your Dragon: Automatic Diffusion‐Based Rigging for Characters with Diverse Topologies
by: Zeqi Gu, et al.
Published: (2025)
by: Zeqi Gu, et al.
Published: (2025)
Chain-of-Thought Prompting Obscures Hallucination Cues in Large Language Models: An Empirical Evaluation
by: Cheng, Jiahao, et al.
Published: (2025)
by: Cheng, Jiahao, et al.
Published: (2025)
Graph-Based Chain-of-Thought Pruning for Reducing Redundant Reflections in Reasoning LLMs
by: Yuan, Hongyuan, et al.
Published: (2026)
by: Yuan, Hongyuan, et al.
Published: (2026)
Visual Thoughts: A Unified Perspective of Understanding Multimodal Chain-of-Thought
by: Cheng, Zihui, et al.
Published: (2025)
by: Cheng, Zihui, et al.
Published: (2025)
MedCoT: Medical Chain of Thought via Hierarchical Expert
by: Liu, Jiaxiang, et al.
Published: (2024)
by: Liu, Jiaxiang, et al.
Published: (2024)
How to Train Your Dragon: Automatic Diffusion-Based Rigging for Characters with Diverse Topologies
by: Gu, Zeqi, et al.
Published: (2025)
by: Gu, Zeqi, et al.
Published: (2025)
Visual Chain of Thought: Bridging Logical Gaps with Multimodal Infillings
by: Rose, Daniel, et al.
Published: (2023)
by: Rose, Daniel, et al.
Published: (2023)
OPV: Outcome-based Process Verifier for Efficient Long Chain-of-Thought Verification
by: Wu, Zijian, et al.
Published: (2025)
by: Wu, Zijian, et al.
Published: (2025)
Similar Items
-
UniT: Unified Multimodal Chain-of-Thought Test-time Scaling
by: Chen, Leon Liangyu, et al.
Published: (2026) -
Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning
by: Zhang, Lei, et al.
Published: (2026) -
Token-Shuffle: Towards High-Resolution Image Generation with Autoregressive Models
by: Ma, Xu, et al.
Published: (2025) -
Llama Learns to Direct: DirectorLLM for Human-Centric Video Generation
by: Song, Kunpeng, et al.
Published: (2024) -
Non-Markov Multi-Round Conversational Image Generation with History-Conditioned MLLMs
by: Zhang, Haochen, et al.
Published: (2026)