Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Lei, Tian, Junjiao, Fan, Zhipeng, Li, Kunpeng, Wang, Jialiang, Chen, Weifeng, Georgopoulos, Markos, Juefei-Xu, Felix, Bao, Yuxiang, McAuley, Julian, Li, Manling, He, Zecheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On Generalization in Agentic Tool Calling: CoreThink Agentic Reasoner and MAVEN Dataset
von: Bhat, Vishvesh, et al.
Veröffentlicht: (2025)
von: Bhat, Vishvesh, et al.
Veröffentlicht: (2025)
Improving Chain-of-Thought Efficiency for Autoregressive Image Generation
von: Gu, Zeqi, et al.
Veröffentlicht: (2025)
von: Gu, Zeqi, et al.
Veröffentlicht: (2025)
ThinkRouter: Efficient Reasoning via Routing Thinking between Latent and Discrete Spaces
von: Xu, Xin, et al.
Veröffentlicht: (2026)
von: Xu, Xin, et al.
Veröffentlicht: (2026)
CoreThink: A Symbolic Reasoning Layer to reason over Long Horizon Tasks with LLMs
von: Vaghasiya, Jay, et al.
Veröffentlicht: (2025)
von: Vaghasiya, Jay, et al.
Veröffentlicht: (2025)
Code to Think, Think to Code: A Survey on Code-Enhanced Reasoning and Reasoning-Driven Code Intelligence in LLMs
von: Yang, Dayu, et al.
Veröffentlicht: (2025)
von: Yang, Dayu, et al.
Veröffentlicht: (2025)
Reddit2Deezer: A Scalable Dataset for Real-World Grounded Conversational Music Recommendation
von: Kim, Haven, et al.
Veröffentlicht: (2026)
von: Kim, Haven, et al.
Veröffentlicht: (2026)
Three central limit theorems for the unbounded excursion component of a Gaussian field
von: McAuley, Michael
Veröffentlicht: (2024)
von: McAuley, Michael
Veröffentlicht: (2024)
Children in Custody
von: McAuley, Mary
Veröffentlicht: (2022)
von: McAuley, Mary
Veröffentlicht: (2022)
Politics and the Soviet Union / Mary McAuley
von: McAuley, Mary
von: McAuley, Mary
Limit theorems for non-local functionals of smooth Gaussian fields via quasi-association
von: McAuley, Michael
Veröffentlicht: (2026)
von: McAuley, Michael
Veröffentlicht: (2026)
GSPRec: Temporal-Aware Graph Spectral Filtering for Recommendation
von: Rabiah, Ahmad Bin, et al.
Veröffentlicht: (2025)
von: Rabiah, Ahmad Bin, et al.
Veröffentlicht: (2025)
StreamDiT: Real-Time Streaming Text-to-Video Generation
von: Kodaira, Akio, et al.
Veröffentlicht: (2025)
von: Kodaira, Akio, et al.
Veröffentlicht: (2025)
PixelThink: Towards Efficient Chain-of-Pixel Reasoning
von: Wang, Song, et al.
Veröffentlicht: (2025)
von: Wang, Song, et al.
Veröffentlicht: (2025)
Extending Input Contexts of Language Models through Training on Segmented Sequences
von: Karypis, Petros, et al.
Veröffentlicht: (2023)
von: Karypis, Petros, et al.
Veröffentlicht: (2023)
FusID: Modality-Fused Semantic IDs for Generative Music Recommendation
von: Kim, Haven, et al.
Veröffentlicht: (2026)
von: Kim, Haven, et al.
Veröffentlicht: (2026)
Multi-Behavior Generative Recommendation
von: Liu, Zihan, et al.
Veröffentlicht: (2024)
von: Liu, Zihan, et al.
Veröffentlicht: (2024)
Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions
von: Hu, Yuanzhe, et al.
Veröffentlicht: (2025)
von: Hu, Yuanzhe, et al.
Veröffentlicht: (2025)
Inductive Generative Recommendation via Retrieval-based Speculation
von: Ding, Yijie, et al.
Veröffentlicht: (2024)
von: Ding, Yijie, et al.
Veröffentlicht: (2024)
Purely Semantic Indexing for LLM-based Generative Recommendation and Retrieval
von: Zhang, Ruohan, et al.
Veröffentlicht: (2025)
von: Zhang, Ruohan, et al.
Veröffentlicht: (2025)
PDMX: A Large-Scale Public Domain MusicXML Dataset for Symbolic Music Processing
von: Long, Phillip, et al.
Veröffentlicht: (2024)
von: Long, Phillip, et al.
Veröffentlicht: (2024)
Non-Markov Multi-Round Conversational Image Generation with History-Conditioned MLLMs
von: Zhang, Haochen, et al.
Veröffentlicht: (2026)
von: Zhang, Haochen, et al.
Veröffentlicht: (2026)
Exploring MLLM-Diffusion Information Transfer with MetaCanvas
von: Lin, Han, et al.
Veröffentlicht: (2025)
von: Lin, Han, et al.
Veröffentlicht: (2025)
Pixel-Space Post-Training of Latent Diffusion Models
von: Zhang, Christina, et al.
Veröffentlicht: (2024)
von: Zhang, Christina, et al.
Veröffentlicht: (2024)
Bridging Conversational and Collaborative Signals for Conversational Recommendation
von: Rabiah, Ahmad Bin, et al.
Veröffentlicht: (2024)
von: Rabiah, Ahmad Bin, et al.
Veröffentlicht: (2024)
Imagery as Inquiry: Exploring A Multimodal Dataset for Conversational Recommendation
von: Yoon, Se-eun, et al.
Veröffentlicht: (2024)
von: Yoon, Se-eun, et al.
Veröffentlicht: (2024)
Calibration-Disentangled Learning and Relevance-Prioritized Reranking for Calibrated Sequential Recommendation
von: Jeon, Hyunsik, et al.
Veröffentlicht: (2024)
von: Jeon, Hyunsik, et al.
Veröffentlicht: (2024)
Limit theorems for the number of sign and level-set clusters of the Gaussian free field
von: McAuley, Michael, et al.
Veröffentlicht: (2025)
von: McAuley, Michael, et al.
Veröffentlicht: (2025)
Educating Young Children: A Structural Approach. Routledge Library Editions: Early Years
von: McAuley, Helen, et al.
Veröffentlicht: (2022)
von: McAuley, Helen, et al.
Veröffentlicht: (2022)
Unlocking Decoding-time Controllability: Gradient-Free Multi-Objective Alignment with Contrastive Prompts
von: Fu, Tingchen, et al.
Veröffentlicht: (2024)
von: Fu, Tingchen, et al.
Veröffentlicht: (2024)
Composer Vector: Style-steering Symbolic Music Generation in a Latent Space
von: Jiang, Xunyi, et al.
Veröffentlicht: (2026)
von: Jiang, Xunyi, et al.
Veröffentlicht: (2026)
LVCHAT: Facilitating Long Video Comprehension
von: Wang, Yu, et al.
Veröffentlicht: (2024)
von: Wang, Yu, et al.
Veröffentlicht: (2024)
BiasEdit: Debiasing Stereotyped Language Models via Model Editing
von: Xu, Xin, et al.
Veröffentlicht: (2025)
von: Xu, Xin, et al.
Veröffentlicht: (2025)
When Benchmarks Age: Temporal Misalignment through Large Language Model Factuality Evaluation
von: Jiang, Xunyi, et al.
Veröffentlicht: (2025)
von: Jiang, Xunyi, et al.
Veröffentlicht: (2025)
FINEST: Stabilizing Recommendations by Rank-Preserving Fine-Tuning
von: Oh, Sejoon, et al.
Veröffentlicht: (2024)
von: Oh, Sejoon, et al.
Veröffentlicht: (2024)
ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning
von: Gu, Jiawei, et al.
Veröffentlicht: (2025)
von: Gu, Jiawei, et al.
Veröffentlicht: (2025)
Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation
von: Guo, Ziyu, et al.
Veröffentlicht: (2025)
von: Guo, Ziyu, et al.
Veröffentlicht: (2025)
Improving In-Context Learning with Reasoning Distillation
von: Sadeq, Nafis, et al.
Veröffentlicht: (2025)
von: Sadeq, Nafis, et al.
Veröffentlicht: (2025)
Explainable Chain-of-Thought Reasoning: An Empirical Analysis on State-Aware Reasoning Dynamics
von: Yu, Sheldon, et al.
Veröffentlicht: (2025)
von: Yu, Sheldon, et al.
Veröffentlicht: (2025)
MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning
von: Chen, Xinyan, et al.
Veröffentlicht: (2025)
von: Chen, Xinyan, et al.
Veröffentlicht: (2025)
The Thinking Pixel: Recursive Sparse Reasoning in Multimodal Diffusion Latents
von: Sun, Yuwei, et al.
Veröffentlicht: (2026)
von: Sun, Yuwei, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
On Generalization in Agentic Tool Calling: CoreThink Agentic Reasoner and MAVEN Dataset
von: Bhat, Vishvesh, et al.
Veröffentlicht: (2025) -
Improving Chain-of-Thought Efficiency for Autoregressive Image Generation
von: Gu, Zeqi, et al.
Veröffentlicht: (2025) -
ThinkRouter: Efficient Reasoning via Routing Thinking between Latent and Discrete Spaces
von: Xu, Xin, et al.
Veröffentlicht: (2026) -
CoreThink: A Symbolic Reasoning Layer to reason over Long Horizon Tasks with LLMs
von: Vaghasiya, Jay, et al.
Veröffentlicht: (2025) -
Code to Think, Think to Code: A Survey on Code-Enhanced Reasoning and Reasoning-Driven Code Intelligence in LLMs
von: Yang, Dayu, et al.
Veröffentlicht: (2025)