Think-Then-Generate: Reasoning-Aware Text-to-Image Diffusion with LLM Encoders
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kou, Siqi, Jin, Jiachun, Zhou, Zetong, Ma, Ye, Wang, Yugang, Chen, Quan, Jiang, Peng, Yang, Xiao, Zhu, Jun, Yu, Kai, Deng, Zhijie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads
von: Kou, Siqi, et al.
Veröffentlicht: (2024)
von: Kou, Siqi, et al.
Veröffentlicht: (2024)
LatentUM: Unleashing the Potential of Interleaved Cross-Modal Reasoning via a Latent-Space Unified Model
von: Jin, Jiachun, et al.
Veröffentlicht: (2026)
von: Jin, Jiachun, et al.
Veröffentlicht: (2026)
Thinking with Generated Images
von: Chern, Ethan, et al.
Veröffentlicht: (2025)
von: Chern, Ethan, et al.
Veröffentlicht: (2025)
ProductWebGen: Benchmarking Multimodal Product Webpage Generation
von: Liu, Zhihong, et al.
Veröffentlicht: (2026)
von: Liu, Zhihong, et al.
Veröffentlicht: (2026)
Diffusion LLMs Can Do Faster-Than-AR Inference via Discrete Diffusion Forcing
von: Wang, Xu, et al.
Veröffentlicht: (2025)
von: Wang, Xu, et al.
Veröffentlicht: (2025)
Which Data Attributes Stimulate Math and Code Reasoning? An Investigation via Influence Functions
von: Kou, Siqi, et al.
Veröffentlicht: (2025)
von: Kou, Siqi, et al.
Veröffentlicht: (2025)
BayesDiff: Estimating Pixel-wise Uncertainty in Diffusion via Bayesian Inference
von: Kou, Siqi, et al.
Veröffentlicht: (2023)
von: Kou, Siqi, et al.
Veröffentlicht: (2023)
Lost in Translation: Latent Concept Misalignment in Text-to-Image Diffusion Models
von: Zhao, Juntu, et al.
Veröffentlicht: (2024)
von: Zhao, Juntu, et al.
Veröffentlicht: (2024)
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks
von: Yang, Yi, et al.
Veröffentlicht: (2025)
von: Yang, Yi, et al.
Veröffentlicht: (2025)
CLLMs: Consistency Large Language Models
von: Kou, Siqi, et al.
Veröffentlicht: (2024)
von: Kou, Siqi, et al.
Veröffentlicht: (2024)
Diffusion Lens: Interpreting Text Encoders in Text-to-Image Pipelines
von: Toker, Michael, et al.
Veröffentlicht: (2024)
von: Toker, Michael, et al.
Veröffentlicht: (2024)
Scaling Down Text Encoders of Text-to-Image Diffusion Models
von: Wang, Lifu, et al.
Veröffentlicht: (2025)
von: Wang, Lifu, et al.
Veröffentlicht: (2025)
Aligning Text to Image in Diffusion Models is Easier Than You Think
von: Lee, Jaa-Yeon, et al.
Veröffentlicht: (2025)
von: Lee, Jaa-Yeon, et al.
Veröffentlicht: (2025)
TFRank: Think-Free Reasoning Enables Practical Pointwise LLM Ranking
von: Fan, Yongqi, et al.
Veröffentlicht: (2025)
von: Fan, Yongqi, et al.
Veröffentlicht: (2025)
Hawk: Leveraging Spatial Context for Faster Autoregressive Text-to-Image Generation
von: Chen, Zhi-Kai, et al.
Veröffentlicht: (2025)
von: Chen, Zhi-Kai, et al.
Veröffentlicht: (2025)
Language-Image Alignment with Fixed Text Encoders
von: Yang, Jingfeng, et al.
Veröffentlicht: (2025)
von: Yang, Jingfeng, et al.
Veröffentlicht: (2025)
SIFT: Grounding LLM Reasoning in Contexts via Stickers
von: Zeng, Zihao, et al.
Veröffentlicht: (2025)
von: Zeng, Zihao, et al.
Veröffentlicht: (2025)
Think When Needed: Model-Aware Reasoning Routing for LLM-based Ranking
von: Guo, Huizhong, et al.
Veröffentlicht: (2026)
von: Guo, Huizhong, et al.
Veröffentlicht: (2026)
Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting
von: Shen, Si, et al.
Veröffentlicht: (2025)
von: Shen, Si, et al.
Veröffentlicht: (2025)
BicliqueEncoder: An Efficient Method for Link Prediction in Bipartite Networks using Formal Concept Analysis and Transformer Encoder
von: Yang, Hongyuan, et al.
Veröffentlicht: (2025)
von: Yang, Hongyuan, et al.
Veröffentlicht: (2025)
ThinkRL-Edit: Thinking in Reinforcement Learning for Reasoning-Centric Image Editing
von: Li, Hengjia, et al.
Veröffentlicht: (2026)
von: Li, Hengjia, et al.
Veröffentlicht: (2026)
Think Smarter not Harder: Adaptive Reasoning with Inference Aware Optimization
von: Yu, Zishun, et al.
Veröffentlicht: (2025)
von: Yu, Zishun, et al.
Veröffentlicht: (2025)
Efficient Detection of LLM-generated Texts with a Bayesian Surrogate Model
von: Miao, Yibo, et al.
Veröffentlicht: (2023)
von: Miao, Yibo, et al.
Veröffentlicht: (2023)
MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal Projection
von: Lin, Bokai, et al.
Veröffentlicht: (2024)
von: Lin, Bokai, et al.
Veröffentlicht: (2024)
Text-Aware Image Restoration with Diffusion Models
von: Min, Jaewon, et al.
Veröffentlicht: (2025)
von: Min, Jaewon, et al.
Veröffentlicht: (2025)
SocialTraj: Two-Stage Socially-Aware Trajectory Prediction for Autonomous Driving via Conditional Diffusion Model
von: Zhou, Xiao, et al.
Veröffentlicht: (2025)
von: Zhou, Xiao, et al.
Veröffentlicht: (2025)
One-Way Ticket:Time-Independent Unified Encoder for Distilling Text-to-Image Diffusion Models
von: Li, Senmao, et al.
Veröffentlicht: (2025)
von: Li, Senmao, et al.
Veröffentlicht: (2025)
Think-Augmented Function Calling: Improving LLM Parameter Accuracy Through Embedded Reasoning
von: Wei, Lei, et al.
Veröffentlicht: (2026)
von: Wei, Lei, et al.
Veröffentlicht: (2026)
When to Think, When to Speak: Learning Disclosure Policies for LLM Reasoning
von: Wei, Jiaqi, et al.
Veröffentlicht: (2026)
von: Wei, Jiaqi, et al.
Veröffentlicht: (2026)
Reason2Attack: Jailbreaking Text-to-Image Models via LLM Reasoning
von: Zhang, Chenyu, et al.
Veröffentlicht: (2025)
von: Zhang, Chenyu, et al.
Veröffentlicht: (2025)
MIRAGE: Evaluating and Explaining Inductive Reasoning Process in Language Models
von: Li, Jiachun, et al.
Veröffentlicht: (2024)
von: Li, Jiachun, et al.
Veröffentlicht: (2024)
d3LLM: Ultra-Fast Diffusion LLM using Pseudo-Trajectory Distillation
von: Qian, Yu-Yang, et al.
Veröffentlicht: (2026)
von: Qian, Yu-Yang, et al.
Veröffentlicht: (2026)
Learn-to-Distance: Distance Learning for Detecting LLM-Generated Text
von: Zhou, Hongyi, et al.
Veröffentlicht: (2026)
von: Zhou, Hongyi, et al.
Veröffentlicht: (2026)
LOVECon: Text-driven Training-Free Long Video Editing with ControlNet
von: Liao, Zhenyi, et al.
Veröffentlicht: (2023)
von: Liao, Zhenyi, et al.
Veröffentlicht: (2023)
Ideal-LLM: Integrating Dual Encoders and Language-Adapted LLM for Multilingual Speech-to-Text
von: Xue, Hongfei, et al.
Veröffentlicht: (2024)
von: Xue, Hongfei, et al.
Veröffentlicht: (2024)
Thinking in Text and Images: Interleaved Vision--Language Reasoning Traces for Long-Horizon Robot Manipulation
von: Liu, Jinkun, et al.
Veröffentlicht: (2026)
von: Liu, Jinkun, et al.
Veröffentlicht: (2026)
Could Thinking Multilingually Empower LLM Reasoning?
von: Gao, Changjiang, et al.
Veröffentlicht: (2025)
von: Gao, Changjiang, et al.
Veröffentlicht: (2025)
Thinking with Constructions: A Benchmark and Policy Optimization for Visual-Text Interleaved Geometric Reasoning
von: Zhao, Haokun, et al.
Veröffentlicht: (2026)
von: Zhao, Haokun, et al.
Veröffentlicht: (2026)
Enhancing Spatial Reasoning through Visual and Textual Thinking
von: Liang, Xun, et al.
Veröffentlicht: (2025)
von: Liang, Xun, et al.
Veröffentlicht: (2025)
DualEdit: Mitigating Safety Fallback in LLM Backdoor Editing via Affirmation-Refusal Regulation
von: Jiang, Houcheng, et al.
Veröffentlicht: (2025)
von: Jiang, Houcheng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads
von: Kou, Siqi, et al.
Veröffentlicht: (2024) -
LatentUM: Unleashing the Potential of Interleaved Cross-Modal Reasoning via a Latent-Space Unified Model
von: Jin, Jiachun, et al.
Veröffentlicht: (2026) -
Thinking with Generated Images
von: Chern, Ethan, et al.
Veröffentlicht: (2025) -
ProductWebGen: Benchmarking Multimodal Product Webpage Generation
von: Liu, Zhihong, et al.
Veröffentlicht: (2026) -
Diffusion LLMs Can Do Faster-Than-AR Inference via Discrete Diffusion Forcing
von: Wang, Xu, et al.
Veröffentlicht: (2025)