Think-Then-Generate: Reasoning-Aware Text-to-Image Diffusion with LLM Encoders
Fuente:
arXiv
Salvato in:
| Autori principali: | Kou, Siqi, Jin, Jiachun, Zhou, Zetong, Ma, Ye, Wang, Yugang, Chen, Quan, Jiang, Peng, Yang, Xiao, Zhu, Jun, Yu, Kai, Deng, Zhijie |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads
di: Kou, Siqi, et al.
Pubblicazione: (2024)
di: Kou, Siqi, et al.
Pubblicazione: (2024)
LatentUM: Unleashing the Potential of Interleaved Cross-Modal Reasoning via a Latent-Space Unified Model
di: Jin, Jiachun, et al.
Pubblicazione: (2026)
di: Jin, Jiachun, et al.
Pubblicazione: (2026)
Thinking with Generated Images
di: Chern, Ethan, et al.
Pubblicazione: (2025)
di: Chern, Ethan, et al.
Pubblicazione: (2025)
ProductWebGen: Benchmarking Multimodal Product Webpage Generation
di: Liu, Zhihong, et al.
Pubblicazione: (2026)
di: Liu, Zhihong, et al.
Pubblicazione: (2026)
Diffusion LLMs Can Do Faster-Than-AR Inference via Discrete Diffusion Forcing
di: Wang, Xu, et al.
Pubblicazione: (2025)
di: Wang, Xu, et al.
Pubblicazione: (2025)
Which Data Attributes Stimulate Math and Code Reasoning? An Investigation via Influence Functions
di: Kou, Siqi, et al.
Pubblicazione: (2025)
di: Kou, Siqi, et al.
Pubblicazione: (2025)
BayesDiff: Estimating Pixel-wise Uncertainty in Diffusion via Bayesian Inference
di: Kou, Siqi, et al.
Pubblicazione: (2023)
di: Kou, Siqi, et al.
Pubblicazione: (2023)
Lost in Translation: Latent Concept Misalignment in Text-to-Image Diffusion Models
di: Zhao, Juntu, et al.
Pubblicazione: (2024)
di: Zhao, Juntu, et al.
Pubblicazione: (2024)
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks
di: Yang, Yi, et al.
Pubblicazione: (2025)
di: Yang, Yi, et al.
Pubblicazione: (2025)
CLLMs: Consistency Large Language Models
di: Kou, Siqi, et al.
Pubblicazione: (2024)
di: Kou, Siqi, et al.
Pubblicazione: (2024)
Diffusion Lens: Interpreting Text Encoders in Text-to-Image Pipelines
di: Toker, Michael, et al.
Pubblicazione: (2024)
di: Toker, Michael, et al.
Pubblicazione: (2024)
Scaling Down Text Encoders of Text-to-Image Diffusion Models
di: Wang, Lifu, et al.
Pubblicazione: (2025)
di: Wang, Lifu, et al.
Pubblicazione: (2025)
Aligning Text to Image in Diffusion Models is Easier Than You Think
di: Lee, Jaa-Yeon, et al.
Pubblicazione: (2025)
di: Lee, Jaa-Yeon, et al.
Pubblicazione: (2025)
TFRank: Think-Free Reasoning Enables Practical Pointwise LLM Ranking
di: Fan, Yongqi, et al.
Pubblicazione: (2025)
di: Fan, Yongqi, et al.
Pubblicazione: (2025)
Hawk: Leveraging Spatial Context for Faster Autoregressive Text-to-Image Generation
di: Chen, Zhi-Kai, et al.
Pubblicazione: (2025)
di: Chen, Zhi-Kai, et al.
Pubblicazione: (2025)
Language-Image Alignment with Fixed Text Encoders
di: Yang, Jingfeng, et al.
Pubblicazione: (2025)
di: Yang, Jingfeng, et al.
Pubblicazione: (2025)
SIFT: Grounding LLM Reasoning in Contexts via Stickers
di: Zeng, Zihao, et al.
Pubblicazione: (2025)
di: Zeng, Zihao, et al.
Pubblicazione: (2025)
Think When Needed: Model-Aware Reasoning Routing for LLM-based Ranking
di: Guo, Huizhong, et al.
Pubblicazione: (2026)
di: Guo, Huizhong, et al.
Pubblicazione: (2026)
Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting
di: Shen, Si, et al.
Pubblicazione: (2025)
di: Shen, Si, et al.
Pubblicazione: (2025)
BicliqueEncoder: An Efficient Method for Link Prediction in Bipartite Networks using Formal Concept Analysis and Transformer Encoder
di: Yang, Hongyuan, et al.
Pubblicazione: (2025)
di: Yang, Hongyuan, et al.
Pubblicazione: (2025)
ThinkRL-Edit: Thinking in Reinforcement Learning for Reasoning-Centric Image Editing
di: Li, Hengjia, et al.
Pubblicazione: (2026)
di: Li, Hengjia, et al.
Pubblicazione: (2026)
Think Smarter not Harder: Adaptive Reasoning with Inference Aware Optimization
di: Yu, Zishun, et al.
Pubblicazione: (2025)
di: Yu, Zishun, et al.
Pubblicazione: (2025)
Efficient Detection of LLM-generated Texts with a Bayesian Surrogate Model
di: Miao, Yibo, et al.
Pubblicazione: (2023)
di: Miao, Yibo, et al.
Pubblicazione: (2023)
MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal Projection
di: Lin, Bokai, et al.
Pubblicazione: (2024)
di: Lin, Bokai, et al.
Pubblicazione: (2024)
Text-Aware Image Restoration with Diffusion Models
di: Min, Jaewon, et al.
Pubblicazione: (2025)
di: Min, Jaewon, et al.
Pubblicazione: (2025)
SocialTraj: Two-Stage Socially-Aware Trajectory Prediction for Autonomous Driving via Conditional Diffusion Model
di: Zhou, Xiao, et al.
Pubblicazione: (2025)
di: Zhou, Xiao, et al.
Pubblicazione: (2025)
One-Way Ticket:Time-Independent Unified Encoder for Distilling Text-to-Image Diffusion Models
di: Li, Senmao, et al.
Pubblicazione: (2025)
di: Li, Senmao, et al.
Pubblicazione: (2025)
Think-Augmented Function Calling: Improving LLM Parameter Accuracy Through Embedded Reasoning
di: Wei, Lei, et al.
Pubblicazione: (2026)
di: Wei, Lei, et al.
Pubblicazione: (2026)
When to Think, When to Speak: Learning Disclosure Policies for LLM Reasoning
di: Wei, Jiaqi, et al.
Pubblicazione: (2026)
di: Wei, Jiaqi, et al.
Pubblicazione: (2026)
Reason2Attack: Jailbreaking Text-to-Image Models via LLM Reasoning
di: Zhang, Chenyu, et al.
Pubblicazione: (2025)
di: Zhang, Chenyu, et al.
Pubblicazione: (2025)
MIRAGE: Evaluating and Explaining Inductive Reasoning Process in Language Models
di: Li, Jiachun, et al.
Pubblicazione: (2024)
di: Li, Jiachun, et al.
Pubblicazione: (2024)
d3LLM: Ultra-Fast Diffusion LLM using Pseudo-Trajectory Distillation
di: Qian, Yu-Yang, et al.
Pubblicazione: (2026)
di: Qian, Yu-Yang, et al.
Pubblicazione: (2026)
Learn-to-Distance: Distance Learning for Detecting LLM-Generated Text
di: Zhou, Hongyi, et al.
Pubblicazione: (2026)
di: Zhou, Hongyi, et al.
Pubblicazione: (2026)
LOVECon: Text-driven Training-Free Long Video Editing with ControlNet
di: Liao, Zhenyi, et al.
Pubblicazione: (2023)
di: Liao, Zhenyi, et al.
Pubblicazione: (2023)
Ideal-LLM: Integrating Dual Encoders and Language-Adapted LLM for Multilingual Speech-to-Text
di: Xue, Hongfei, et al.
Pubblicazione: (2024)
di: Xue, Hongfei, et al.
Pubblicazione: (2024)
Thinking in Text and Images: Interleaved Vision--Language Reasoning Traces for Long-Horizon Robot Manipulation
di: Liu, Jinkun, et al.
Pubblicazione: (2026)
di: Liu, Jinkun, et al.
Pubblicazione: (2026)
Could Thinking Multilingually Empower LLM Reasoning?
di: Gao, Changjiang, et al.
Pubblicazione: (2025)
di: Gao, Changjiang, et al.
Pubblicazione: (2025)
Thinking with Constructions: A Benchmark and Policy Optimization for Visual-Text Interleaved Geometric Reasoning
di: Zhao, Haokun, et al.
Pubblicazione: (2026)
di: Zhao, Haokun, et al.
Pubblicazione: (2026)
Enhancing Spatial Reasoning through Visual and Textual Thinking
di: Liang, Xun, et al.
Pubblicazione: (2025)
di: Liang, Xun, et al.
Pubblicazione: (2025)
DualEdit: Mitigating Safety Fallback in LLM Backdoor Editing via Affirmation-Refusal Regulation
di: Jiang, Houcheng, et al.
Pubblicazione: (2025)
di: Jiang, Houcheng, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads
di: Kou, Siqi, et al.
Pubblicazione: (2024) -
LatentUM: Unleashing the Potential of Interleaved Cross-Modal Reasoning via a Latent-Space Unified Model
di: Jin, Jiachun, et al.
Pubblicazione: (2026) -
Thinking with Generated Images
di: Chern, Ethan, et al.
Pubblicazione: (2025) -
ProductWebGen: Benchmarking Multimodal Product Webpage Generation
di: Liu, Zhihong, et al.
Pubblicazione: (2026) -
Diffusion LLMs Can Do Faster-Than-AR Inference via Discrete Diffusion Forcing
di: Wang, Xu, et al.
Pubblicazione: (2025)