LaDiC: Are Diffusion Models Really Inferior to Autoregressive Counterparts for Image-to-Text Generation?
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Yuchi, Ren, Shuhuai, Gao, Rundong, Yao, Linli, Guo, Qingyan, An, Kaikai, Bai, Jianhong, Sun, Xu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction
di: Wang, Yuchi, et al.
Pubblicazione: (2025)
di: Wang, Yuchi, et al.
Pubblicazione: (2025)
TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
di: Ren, Shuhuai, et al.
Pubblicazione: (2023)
di: Ren, Shuhuai, et al.
Pubblicazione: (2023)
Next Block Prediction: Video Generation via Semi-Autoregressive Modeling
di: Ren, Shuhuai, et al.
Pubblicazione: (2025)
di: Ren, Shuhuai, et al.
Pubblicazione: (2025)
DiSA: Diffusion Step Annealing in Autoregressive Image Generation
di: Zhao, Qinyu, et al.
Pubblicazione: (2025)
di: Zhao, Qinyu, et al.
Pubblicazione: (2025)
Rethinking Semantic Parsing for Large Language Models: Enhancing LLM Performance with Semantic Hints
di: An, Kaikai, et al.
Pubblicazione: (2024)
di: An, Kaikai, et al.
Pubblicazione: (2024)
Bridging Continuous and Discrete Tokens for Autoregressive Visual Generation
di: Wang, Yuqing, et al.
Pubblicazione: (2025)
di: Wang, Yuqing, et al.
Pubblicazione: (2025)
InstructAvatar: Text-Guided Emotion and Motion Control for Avatar Generation
di: Wang, Yuchi, et al.
Pubblicazione: (2024)
di: Wang, Yuchi, et al.
Pubblicazione: (2024)
DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models
di: Yao, Linli, et al.
Pubblicazione: (2024)
di: Yao, Linli, et al.
Pubblicazione: (2024)
Parallelized Autoregressive Visual Generation
di: Wang, Yuqing, et al.
Pubblicazione: (2024)
di: Wang, Yuqing, et al.
Pubblicazione: (2024)
Unifying Continuous and Discrete Text Diffusion with Non-simultaneous Diffusion Processes
di: Li, Bocheng, et al.
Pubblicazione: (2025)
di: Li, Bocheng, et al.
Pubblicazione: (2025)
VITATECS: A Diagnostic Dataset for Temporal Concept Understanding of Video-Language Models
di: Li, Shicheng, et al.
Pubblicazione: (2023)
di: Li, Shicheng, et al.
Pubblicazione: (2023)
LayoutDiT: Exploring Content-Graphic Balance in Layout Generation with Diffusion Transformer
di: Li, Yu, et al.
Pubblicazione: (2024)
di: Li, Yu, et al.
Pubblicazione: (2024)
Empowering Diffusion Models on the Embedding Space for Text Generation
di: Gao, Zhujin, et al.
Pubblicazione: (2022)
di: Gao, Zhujin, et al.
Pubblicazione: (2022)
CoDi: Subject-Consistent and Pose-Diverse Text-to-Image Generation
di: Gao, Zhanxin, et al.
Pubblicazione: (2025)
di: Gao, Zhanxin, et al.
Pubblicazione: (2025)
TempCompass: Do Video LLMs Really Understand Videos?
di: Liu, Yuanxin, et al.
Pubblicazione: (2024)
di: Liu, Yuanxin, et al.
Pubblicazione: (2024)
Towards Multimodal Video Paragraph Captioning Models Robust to Missing Modality
di: Chen, Sishuo, et al.
Pubblicazione: (2024)
di: Chen, Sishuo, et al.
Pubblicazione: (2024)
Diffusion Beats Autoregressive: An Evaluation of Compositional Generation in Text-to-Image Models
di: Marioriyad, Arash, et al.
Pubblicazione: (2024)
di: Marioriyad, Arash, et al.
Pubblicazione: (2024)
Unlocking the Power of GANs in Non-Autoregressive Text Generation
di: Ren, Da, et al.
Pubblicazione: (2023)
di: Ren, Da, et al.
Pubblicazione: (2023)
HQFNN: A Compact Quantum-Fuzzy Neural Network for Accurate Image Classification
di: Yao, Jianhong, et al.
Pubblicazione: (2025)
di: Yao, Jianhong, et al.
Pubblicazione: (2025)
DiTAR: Diffusion Transformer Autoregressive Modeling for Speech Generation
di: Jia, Dongya, et al.
Pubblicazione: (2025)
di: Jia, Dongya, et al.
Pubblicazione: (2025)
Generation then Reconstruction: Accelerating Masked Autoregressive Models via Two-Stage Sampling
di: Yan, Feihong, et al.
Pubblicazione: (2025)
di: Yan, Feihong, et al.
Pubblicazione: (2025)
DiT-Air: Revisiting the Efficiency of Diffusion Model Architecture Design in Text to Image Generation
di: Chen, Chen, et al.
Pubblicazione: (2025)
di: Chen, Chen, et al.
Pubblicazione: (2025)
JeDi: Joint-Image Diffusion Models for Finetuning-Free Personalized Text-to-Image Generation
di: Zeng, Yu, et al.
Pubblicazione: (2024)
di: Zeng, Yu, et al.
Pubblicazione: (2024)
Design Your Ad: Personalized Advertising Image and Text Generation with Unified Autoregressive Models
di: Xu, Yexing, et al.
Pubblicazione: (2026)
di: Xu, Yexing, et al.
Pubblicazione: (2026)
Differences in Text Generated by Diffusion and Autoregressive Language Models
di: Zhang, Zeyang, et al.
Pubblicazione: (2026)
di: Zhang, Zeyang, et al.
Pubblicazione: (2026)
Autoregressive Diffusion Transformer for Text-to-Speech Synthesis
di: Liu, Zhijun, et al.
Pubblicazione: (2024)
di: Liu, Zhijun, et al.
Pubblicazione: (2024)
Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
di: Sun, Peize, et al.
Pubblicazione: (2024)
di: Sun, Peize, et al.
Pubblicazione: (2024)
DiCoDe: Diffusion-Compressed Deep Tokens for Autoregressive Video Generation with Language Models
di: Li, Yizhuo, et al.
Pubblicazione: (2024)
di: Li, Yizhuo, et al.
Pubblicazione: (2024)
ADP-DiT: Text-Guided Diffusion Transformer for Brain Image Generation in Alzheimer's Disease Progression
di: Lee, Juneyong, et al.
Pubblicazione: (2026)
di: Lee, Juneyong, et al.
Pubblicazione: (2026)
MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?
di: Chen, Zhaorun, et al.
Pubblicazione: (2024)
di: Chen, Zhaorun, et al.
Pubblicazione: (2024)
UniEdit: A Unified Tuning-Free Framework for Video Motion and Appearance Editing
di: Bai, Jianhong, et al.
Pubblicazione: (2024)
di: Bai, Jianhong, et al.
Pubblicazione: (2024)
Autoregressive Styled Text Image Generation, but Make it Reliable
di: Zaccagnino, Carmine, et al.
Pubblicazione: (2025)
di: Zaccagnino, Carmine, et al.
Pubblicazione: (2025)
InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation
di: Yue, Yang, et al.
Pubblicazione: (2026)
di: Yue, Yang, et al.
Pubblicazione: (2026)
DiSTAR: Diffusion over a Scalable Token Autoregressive Representation for Speech Generation
di: Song, Yakun, et al.
Pubblicazione: (2025)
di: Song, Yakun, et al.
Pubblicazione: (2025)
Stabilize the Latent Space for Image Autoregressive Modeling: A Unified Perspective
di: Zhu, Yongxin, et al.
Pubblicazione: (2024)
di: Zhu, Yongxin, et al.
Pubblicazione: (2024)
UVE: Are MLLMs Unified Evaluators for AI-Generated Videos?
di: Liu, Yuanxin, et al.
Pubblicazione: (2025)
di: Liu, Yuanxin, et al.
Pubblicazione: (2025)
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads
di: Kou, Siqi, et al.
Pubblicazione: (2024)
di: Kou, Siqi, et al.
Pubblicazione: (2024)
Video In-context Learning: Autoregressive Transformers are Zero-Shot Video Imitators
di: Zhang, Wentao, et al.
Pubblicazione: (2024)
di: Zhang, Wentao, et al.
Pubblicazione: (2024)
PixelDiT: Pixel Diffusion Transformers for Image Generation
di: Yu, Yongsheng, et al.
Pubblicazione: (2025)
di: Yu, Yongsheng, et al.
Pubblicazione: (2025)
FocusDiT: Masking Queries in Diffusion Transformers for Fine-grained Image Generation
di: Fang, Xueji, et al.
Pubblicazione: (2026)
di: Fang, Xueji, et al.
Pubblicazione: (2026)
Documenti analoghi
-
RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction
di: Wang, Yuchi, et al.
Pubblicazione: (2025) -
TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
di: Ren, Shuhuai, et al.
Pubblicazione: (2023) -
Next Block Prediction: Video Generation via Semi-Autoregressive Modeling
di: Ren, Shuhuai, et al.
Pubblicazione: (2025) -
DiSA: Diffusion Step Annealing in Autoregressive Image Generation
di: Zhao, Qinyu, et al.
Pubblicazione: (2025) -
Rethinking Semantic Parsing for Large Language Models: Enhancing LLM Performance with Semantic Hints
di: An, Kaikai, et al.
Pubblicazione: (2024)