Decoder-Only LLMs are Better Controllers for Diffusion Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dong, Ziyi, Xiao, Yao, Wei, Pengxu, Lin, Liang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DreamArtist++: Controllable One-Shot Text-to-Image Generation via Positive-Negative Adapter
von: Dong, Ziyi, et al.
Veröffentlicht: (2022)
von: Dong, Ziyi, et al.
Veröffentlicht: (2022)
A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation
von: Wang, Andrew Z., et al.
Veröffentlicht: (2025)
von: Wang, Andrew Z., et al.
Veröffentlicht: (2025)
Streaming-dLLM: Accelerating Diffusion LLMs via Suffix Pruning and Dynamic Decoding
von: Xiao, Zhongyu, et al.
Veröffentlicht: (2026)
von: Xiao, Zhongyu, et al.
Veröffentlicht: (2026)
Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents
von: Lin, Han, et al.
Veröffentlicht: (2025)
von: Lin, Han, et al.
Veröffentlicht: (2025)
Controlling Multimodal LLMs via Reward-guided Decoding
von: Mañas, Oscar, et al.
Veröffentlicht: (2025)
von: Mañas, Oscar, et al.
Veröffentlicht: (2025)
EasyGen: Easing Multimodal Generation with BiDiffuser and LLMs
von: Zhao, Xiangyu, et al.
Veröffentlicht: (2023)
von: Zhao, Xiangyu, et al.
Veröffentlicht: (2023)
TOPA: Extending Large Language Models for Video Understanding via Text-Only Pre-Alignment
von: Li, Wei, et al.
Veröffentlicht: (2024)
von: Li, Wei, et al.
Veröffentlicht: (2024)
POP: Prefill-Only Pruning for Efficient Large Model Inference
von: He, Junhui, et al.
Veröffentlicht: (2026)
von: He, Junhui, et al.
Veröffentlicht: (2026)
V2P-Bench: Evaluating Video-Language Understanding with Visual Prompts for Better Human-Model Interaction
von: Zhao, Yiming, et al.
Veröffentlicht: (2025)
von: Zhao, Yiming, et al.
Veröffentlicht: (2025)
MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings
von: Chen, Haonan, et al.
Veröffentlicht: (2025)
von: Chen, Haonan, et al.
Veröffentlicht: (2025)
In-Situ Tweedie Discrete Diffusion Models
von: Li, Xiao, et al.
Veröffentlicht: (2025)
von: Li, Xiao, et al.
Veröffentlicht: (2025)
Autoregressive Semantic Visual Reconstruction Helps VLMs Understand Better
von: Wang, Dianyi, et al.
Veröffentlicht: (2025)
von: Wang, Dianyi, et al.
Veröffentlicht: (2025)
Delve into Visual Contrastive Decoding for Hallucination Mitigation of Large Vision-Language Models
von: Lee, Yi-Lun, et al.
Veröffentlicht: (2024)
von: Lee, Yi-Lun, et al.
Veröffentlicht: (2024)
Better Language Models Exhibit Higher Visual Alignment
von: Ruthardt, Jona, et al.
Veröffentlicht: (2024)
von: Ruthardt, Jona, et al.
Veröffentlicht: (2024)
Bridging the Missing-Modality Gap: Improving Text-Only Calibration of Vision Language Models
von: Kim, Mingyeong, et al.
Veröffentlicht: (2026)
von: Kim, Mingyeong, et al.
Veröffentlicht: (2026)
SpecVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token Pruning
von: Ji, Yicheng, et al.
Veröffentlicht: (2025)
von: Ji, Yicheng, et al.
Veröffentlicht: (2025)
Measuring Social Bias in Vision-Language Models with Face-Only Counterfactuals from Real Photos
von: Chen, Haodong, et al.
Veröffentlicht: (2026)
von: Chen, Haodong, et al.
Veröffentlicht: (2026)
Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding
von: Wang, Xintong, et al.
Veröffentlicht: (2024)
von: Wang, Xintong, et al.
Veröffentlicht: (2024)
Dyna-Mind: Learning to Simulate from Experience for Better AI Agents
von: Yu, Xiao, et al.
Veröffentlicht: (2025)
von: Yu, Xiao, et al.
Veröffentlicht: (2025)
Mixture of Decoding: An Attention-Inspired Adaptive Decoding Strategy to Mitigate Hallucinations in Large Vision-Language Models
von: Chen, Xinlong, et al.
Veröffentlicht: (2025)
von: Chen, Xinlong, et al.
Veröffentlicht: (2025)
CrossWordBench: Evaluating the Reasoning Capabilities of LLMs and LVLMs with Controllable Puzzle Generation
von: Leng, Jixuan, et al.
Veröffentlicht: (2025)
von: Leng, Jixuan, et al.
Veröffentlicht: (2025)
How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding
von: Yu, Zhuoran, et al.
Veröffentlicht: (2025)
von: Yu, Zhuoran, et al.
Veröffentlicht: (2025)
Towards Understanding the Robustness of Diffusion-Based Purification: A Stochastic Perspective
von: Liu, Yiming, et al.
Veröffentlicht: (2024)
von: Liu, Yiming, et al.
Veröffentlicht: (2024)
Bayesian Optimization for Controlled Image Editing via LLMs
von: Cai, Chengkun, et al.
Veröffentlicht: (2025)
von: Cai, Chengkun, et al.
Veröffentlicht: (2025)
Why Only Text: Empowering Vision-and-Language Navigation with Multi-modal Prompts
von: Hong, Haodong, et al.
Veröffentlicht: (2024)
von: Hong, Haodong, et al.
Veröffentlicht: (2024)
Better Safe Than Sorry? Overreaction Problem of Vision Language Models in Visual Emergency Recognition
von: Choi, Dasol, et al.
Veröffentlicht: (2025)
von: Choi, Dasol, et al.
Veröffentlicht: (2025)
Decoding fMRI Data into Captions using Prefix Language Modeling
von: Shen, Vyacheslav, et al.
Veröffentlicht: (2025)
von: Shen, Vyacheslav, et al.
Veröffentlicht: (2025)
Beyond Static Cropping: Layer-Adaptive Visual Localization and Decoding Enhancement
von: Zhu, Zipeng, et al.
Veröffentlicht: (2026)
von: Zhu, Zipeng, et al.
Veröffentlicht: (2026)
Vision Enhancing LLMs: Empowering Multimodal Knowledge Storage and Sharing in LLMs
von: Li, Yunxin, et al.
Veröffentlicht: (2023)
von: Li, Yunxin, et al.
Veröffentlicht: (2023)
MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical Code
von: Lu, Zimu, et al.
Veröffentlicht: (2024)
von: Lu, Zimu, et al.
Veröffentlicht: (2024)
LaDiC: Are Diffusion Models Really Inferior to Autoregressive Counterparts for Image-to-Text Generation?
von: Wang, Yuchi, et al.
Veröffentlicht: (2024)
von: Wang, Yuchi, et al.
Veröffentlicht: (2024)
Visually Guided Decoding: Gradient-Free Hard Prompt Inversion with Language Models
von: Kim, Donghoon, et al.
Veröffentlicht: (2025)
von: Kim, Donghoon, et al.
Veröffentlicht: (2025)
Mitigating Hallucinations in Large Vision-Language Models via Summary-Guided Decoding
von: Min, Kyungmin, et al.
Veröffentlicht: (2024)
von: Min, Kyungmin, et al.
Veröffentlicht: (2024)
Mitigating Multilingual Hallucination in Large Vision-Language Models
von: Qu, Xiaoye, et al.
Veröffentlicht: (2024)
von: Qu, Xiaoye, et al.
Veröffentlicht: (2024)
Talk Less, Interact Better: Evaluating In-context Conversational Adaptation in Multimodal LLMs
von: Hua, Yilun, et al.
Veröffentlicht: (2024)
von: Hua, Yilun, et al.
Veröffentlicht: (2024)
Alleviating Hallucination in Large Vision-Language Models with Active Retrieval Augmentation
von: Qu, Xiaoye, et al.
Veröffentlicht: (2024)
von: Qu, Xiaoye, et al.
Veröffentlicht: (2024)
Evolutionary Negative Module Pruning for Better LoRA Merging
von: Cao, Anda, et al.
Veröffentlicht: (2026)
von: Cao, Anda, et al.
Veröffentlicht: (2026)
CognArtive: Large Language Models for Automating Art Analysis and Decoding Aesthetic Elements
von: Khadangi, Afshin, et al.
Veröffentlicht: (2025)
von: Khadangi, Afshin, et al.
Veröffentlicht: (2025)
Quantifying In-Context Reasoning Effects and Memorization Effects in LLMs
von: Lou, Siyu, et al.
Veröffentlicht: (2024)
von: Lou, Siyu, et al.
Veröffentlicht: (2024)
MMEmb-R1: Reasoning-Enhanced Multimodal Embedding with Pair-Aware Selection and Adaptive Control
von: Wang, Yuchi, et al.
Veröffentlicht: (2026)
von: Wang, Yuchi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
DreamArtist++: Controllable One-Shot Text-to-Image Generation via Positive-Negative Adapter
von: Dong, Ziyi, et al.
Veröffentlicht: (2022) -
A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation
von: Wang, Andrew Z., et al.
Veröffentlicht: (2025) -
Streaming-dLLM: Accelerating Diffusion LLMs via Suffix Pruning and Dynamic Decoding
von: Xiao, Zhongyu, et al.
Veröffentlicht: (2026) -
Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents
von: Lin, Han, et al.
Veröffentlicht: (2025) -
Controlling Multimodal LLMs via Reward-guided Decoding
von: Mañas, Oscar, et al.
Veröffentlicht: (2025)