Flexible-length Text Infilling for Discrete Diffusion Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Andrew, Sivakumar, Anushka, Tang, Chiawei, Thomas, Chris |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
In-Situ Tweedie Discrete Diffusion Models
von: Li, Xiao, et al.
Veröffentlicht: (2025)
von: Li, Xiao, et al.
Veröffentlicht: (2025)
SteerVLM: Robust Model Control through Lightweight Activation Steering for Vision Language Models
von: Sivakumar, Anushka, et al.
Veröffentlicht: (2025)
von: Sivakumar, Anushka, et al.
Veröffentlicht: (2025)
Think While You Generate: Discrete Diffusion with Planned Denoising
von: Liu, Sulin, et al.
Veröffentlicht: (2024)
von: Liu, Sulin, et al.
Veröffentlicht: (2024)
$\textit{Jump Your Steps}$: Optimizing Sampling Schedule of Discrete Diffusion Models
von: Park, Yong-Hyun, et al.
Veröffentlicht: (2024)
von: Park, Yong-Hyun, et al.
Veröffentlicht: (2024)
CRCE: Coreference-Retention Concept Erasure in Text-to-Image Diffusion Models
von: Xue, Yuyang, et al.
Veröffentlicht: (2025)
von: Xue, Yuyang, et al.
Veröffentlicht: (2025)
Self-Play Fine-Tuning of Diffusion Models for Text-to-Image Generation
von: Yuan, Huizhuo, et al.
Veröffentlicht: (2024)
von: Yuan, Huizhuo, et al.
Veröffentlicht: (2024)
SurGen: Text-Guided Diffusion Model for Surgical Video Generation
von: Cho, Joseph, et al.
Veröffentlicht: (2024)
von: Cho, Joseph, et al.
Veröffentlicht: (2024)
Match & Choose: Model Selection Framework for Fine-tuning Text-to-Image Diffusion Models
von: Lewandowski, Basile, et al.
Veröffentlicht: (2025)
von: Lewandowski, Basile, et al.
Veröffentlicht: (2025)
On Discrete Prompt Optimization for Diffusion Models
von: Wang, Ruochen, et al.
Veröffentlicht: (2024)
von: Wang, Ruochen, et al.
Veröffentlicht: (2024)
Efficient Pruning of Text-to-Image Models: Insights from Pruning Stable Diffusion
von: Ramesh, Samarth N, et al.
Veröffentlicht: (2024)
von: Ramesh, Samarth N, et al.
Veröffentlicht: (2024)
AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling
von: Zhan, Jun, et al.
Veröffentlicht: (2024)
von: Zhan, Jun, et al.
Veröffentlicht: (2024)
The Best of Both Worlds: Integrating Language Models and Diffusion Models for Video Generation
von: Yin, Aoxiong, et al.
Veröffentlicht: (2025)
von: Yin, Aoxiong, et al.
Veröffentlicht: (2025)
Pre-trained Text-to-Image Diffusion Models Are Versatile Representation Learners for Control
von: Gupta, Gunshi, et al.
Veröffentlicht: (2024)
von: Gupta, Gunshi, et al.
Veröffentlicht: (2024)
FluoroSAM: A Language-promptable Foundation Model for Flexible X-ray Image Segmentation
von: Killeen, Benjamin D., et al.
Veröffentlicht: (2024)
von: Killeen, Benjamin D., et al.
Veröffentlicht: (2024)
Maximal Matching Matters: Preventing Representation Collapse for Robust Cross-Modal Retrieval
von: Alomari, Hani, et al.
Veröffentlicht: (2025)
von: Alomari, Hani, et al.
Veröffentlicht: (2025)
Towards Visual Text Grounding of Multimodal Large Language Model
von: Li, Ming, et al.
Veröffentlicht: (2025)
von: Li, Ming, et al.
Veröffentlicht: (2025)
Interleaving Reasoning for Better Text-to-Image Generation
von: Huang, Wenxuan, et al.
Veröffentlicht: (2025)
von: Huang, Wenxuan, et al.
Veröffentlicht: (2025)
A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation
von: Wang, Andrew Z., et al.
Veröffentlicht: (2025)
von: Wang, Andrew Z., et al.
Veröffentlicht: (2025)
Pre-trained Language Models Do Not Help Auto-regressive Text-to-Image Generation
von: Zhang, Yuhui, et al.
Veröffentlicht: (2023)
von: Zhang, Yuhui, et al.
Veröffentlicht: (2023)
Test-Time Matching: Unlocking Compositional Reasoning in Multimodal Models
von: Zhu, Yinglun, et al.
Veröffentlicht: (2025)
von: Zhu, Yinglun, et al.
Veröffentlicht: (2025)
A Generalist Model for Diverse Text-Guided Medical Image Synthesis
von: Cho, Joseph, et al.
Veröffentlicht: (2024)
von: Cho, Joseph, et al.
Veröffentlicht: (2024)
Unaligning Everything: Or Aligning Any Text to Any Image in Multimodal Models
von: Salman, Shaeke, et al.
Veröffentlicht: (2024)
von: Salman, Shaeke, et al.
Veröffentlicht: (2024)
CapsFusion: Rethinking Image-Text Data at Scale
von: Yu, Qiying, et al.
Veröffentlicht: (2023)
von: Yu, Qiying, et al.
Veröffentlicht: (2023)
Gesture2Text: A Generalizable Decoder for Word-Gesture Keyboards in XR Through Trajectory Coarse Discretization and Pre-training
von: Shen, Junxiao, et al.
Veröffentlicht: (2024)
von: Shen, Junxiao, et al.
Veröffentlicht: (2024)
A Novel Framework for Automated Explain Vision Model Using Vision-Language Models
von: Nguyen, Phu-Vinh, et al.
Veröffentlicht: (2025)
von: Nguyen, Phu-Vinh, et al.
Veröffentlicht: (2025)
Visual Text Matters: Improving Text-KVQA with Visual Text Entity Knowledge-aware Large Multimodal Assistant
von: Penamakuri, Abhirama Subramanyam, et al.
Veröffentlicht: (2024)
von: Penamakuri, Abhirama Subramanyam, et al.
Veröffentlicht: (2024)
ORAL: Prompting Your Large-Scale LoRAs via Conditional Recurrent Diffusion
von: Khan, Rana Muhammad Shahroz, et al.
Veröffentlicht: (2025)
von: Khan, Rana Muhammad Shahroz, et al.
Veröffentlicht: (2025)
A Comparative Study of Machine Unlearning Techniques for Image and Text Classification Models
von: Safa, Omar M., et al.
Veröffentlicht: (2024)
von: Safa, Omar M., et al.
Veröffentlicht: (2024)
RIV: Recursive Introspection Mask Diffusion Vision Language Model
von: Li, YuQian, et al.
Veröffentlicht: (2025)
von: Li, YuQian, et al.
Veröffentlicht: (2025)
ReadBench: Measuring the Dense Text Visual Reading Ability of Vision-Language Models
von: Clavié, Benjamin, et al.
Veröffentlicht: (2025)
von: Clavié, Benjamin, et al.
Veröffentlicht: (2025)
Erasing with Precision: Evaluating Specific Concept Erasure from Text-to-Image Generative Models
von: Fuchi, Masane, et al.
Veröffentlicht: (2025)
von: Fuchi, Masane, et al.
Veröffentlicht: (2025)
Zero-Shot Vehicle Model Recognition via Text-Based Retrieval-Augmented Generation
von: Chang, Wei-Chia, et al.
Veröffentlicht: (2025)
von: Chang, Wei-Chia, et al.
Veröffentlicht: (2025)
ZigMa: A DiT-style Zigzag Mamba Diffusion Model
von: Hu, Vincent Tao, et al.
Veröffentlicht: (2024)
von: Hu, Vincent Tao, et al.
Veröffentlicht: (2024)
Preference Alignment for Diffusion Model via Explicit Denoised Distribution Estimation
von: Shi, Dingyuan, et al.
Veröffentlicht: (2024)
von: Shi, Dingyuan, et al.
Veröffentlicht: (2024)
RichSpace: Enriching Text-to-Video Prompt Space via Text Embedding Interpolation
von: Cao, Yuefan, et al.
Veröffentlicht: (2025)
von: Cao, Yuefan, et al.
Veröffentlicht: (2025)
Cross-modal RAG: Sub-dimensional Text-to-Image Retrieval-Augmented Generation
von: Zhu, Mengdan, et al.
Veröffentlicht: (2025)
von: Zhu, Mengdan, et al.
Veröffentlicht: (2025)
Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M)
von: Merchant, Nicholas, et al.
Veröffentlicht: (2025)
von: Merchant, Nicholas, et al.
Veröffentlicht: (2025)
One-Step is Enough: Sparse Autoencoders for Text-to-Image Diffusion Models
von: Surkov, Viacheslav, et al.
Veröffentlicht: (2024)
von: Surkov, Viacheslav, et al.
Veröffentlicht: (2024)
Tables as Texts or Images: Evaluating the Table Reasoning Ability of LLMs and MLLMs
von: Deng, Naihao, et al.
Veröffentlicht: (2024)
von: Deng, Naihao, et al.
Veröffentlicht: (2024)
Evaluating Text-to-Visual Generation with Image-to-Text Generation
von: Lin, Zhiqiu, et al.
Veröffentlicht: (2024)
von: Lin, Zhiqiu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
In-Situ Tweedie Discrete Diffusion Models
von: Li, Xiao, et al.
Veröffentlicht: (2025) -
SteerVLM: Robust Model Control through Lightweight Activation Steering for Vision Language Models
von: Sivakumar, Anushka, et al.
Veröffentlicht: (2025) -
Think While You Generate: Discrete Diffusion with Planned Denoising
von: Liu, Sulin, et al.
Veröffentlicht: (2024) -
$\textit{Jump Your Steps}$: Optimizing Sampling Schedule of Discrete Diffusion Models
von: Park, Yong-Hyun, et al.
Veröffentlicht: (2024) -
CRCE: Coreference-Retention Concept Erasure in Text-to-Image Diffusion Models
von: Xue, Yuyang, et al.
Veröffentlicht: (2025)