Slight Corruption in Pre-training Data Makes Better Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Hao, Han, Yujin, Misra, Diganta, Li, Xiang, Hu, Kai, Zou, Difan, Sugiyama, Masashi, Wang, Jindong, Raj, Bhiksha |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Understanding and Mitigating the Label Noise in Pre-training on Downstream Tasks
by: Chen, Hao, et al.
Published: (2023)
by: Chen, Hao, et al.
Published: (2023)
Masked Autoencoders Are Effective Tokenizers for Diffusion Models
by: Chen, Hao, et al.
Published: (2025)
by: Chen, Hao, et al.
Published: (2025)
Corruption-Aware Training of Latent Video Diffusion Models for Robust Text-to-Video Generation
by: Maduabuchi, Chika, et al.
Published: (2025)
by: Maduabuchi, Chika, et al.
Published: (2025)
Impact of Noisy Supervision in Foundation Model Learning
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
Can Diffusion Models Learn Hidden Inter-Feature Rules Behind Images?
by: Han, Yujin, et al.
Published: (2025)
by: Han, Yujin, et al.
Published: (2025)
Imprecise Label Learning: A Unified Framework for Learning with Various Imprecise Label Configurations
by: Chen, Hao, et al.
Published: (2023)
by: Chen, Hao, et al.
Published: (2023)
XQ-GAN: An Open-source Image Tokenization Framework for Autoregressive Generation
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
Weak-to-Strong Diffusion with Reflection
by: Bai, Lichen, et al.
Published: (2025)
by: Bai, Lichen, et al.
Published: (2025)
On the low-shot transferability of [V]-Mamba
by: Misra, Diganta, et al.
Published: (2024)
by: Misra, Diganta, et al.
Published: (2024)
ImageFolder: Autoregressive Image Generation with Folded Tokens
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
GPT4Image: Large Pre-trained Models Help Vision Models Learn Better on Perception Task
by: Ding, Ning, et al.
Published: (2023)
by: Ding, Ning, et al.
Published: (2023)
SIDE: Surrogate Conditional Data Extraction from Diffusion Models
by: Chen, Yunhao, et al.
Published: (2024)
by: Chen, Yunhao, et al.
Published: (2024)
ControlVAR: Exploring Controllable Visual Autoregressive Modeling
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
Extracting Training Data from Unconditional Diffusion Models
by: Chen, Yunhao, et al.
Published: (2024)
by: Chen, Yunhao, et al.
Published: (2024)
Completing Visual Objects via Bridging Generation and Segmentation
by: Li, Xiang, et al.
Published: (2023)
by: Li, Xiang, et al.
Published: (2023)
Retaining and Enhancing Pre-trained Knowledge in Vision-Language Models with Prompt Ensembling
by: Kim, Donggeun, et al.
Published: (2024)
by: Kim, Donggeun, et al.
Published: (2024)
ED-SAM: An Efficient Diffusion Sampling Approach to Domain Generalization in Vision-Language Foundation Models
by: Truong, Thanh-Dat, et al.
Published: (2024)
by: Truong, Thanh-Dat, et al.
Published: (2024)
Routing Matters in MoE: Scaling Diffusion Transformers with Explicit Routing Guidance
by: Wei, Yujie, et al.
Published: (2025)
by: Wei, Yujie, et al.
Published: (2025)
Corrupted but Not Broken: Understanding and Mitigating the Negative Impacts of Corrupted Data in Visual Instruction Tuning
by: Gou, Yunhao, et al.
Published: (2025)
by: Gou, Yunhao, et al.
Published: (2025)
MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings
by: Chen, Haonan, et al.
Published: (2025)
by: Chen, Haonan, et al.
Published: (2025)
An Embarrassingly Simple Baseline for Imbalanced Semi-Supervised Learning
by: Chen, Hao, et al.
Published: (2022)
by: Chen, Hao, et al.
Published: (2022)
$\text{R}^2$-Bench: Benchmarking the Robustness of Referring Perception Models under Perturbations
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
A Category-theoretical Meta-analysis of Definitions of Disentanglement
by: Zhang, Yivan, et al.
Published: (2023)
by: Zhang, Yivan, et al.
Published: (2023)
Unsupervised Pre-training with Language-Vision Prompts for Low-Data Instance Segmentation
by: Zhang, Dingwen, et al.
Published: (2024)
by: Zhang, Dingwen, et al.
Published: (2024)
SoftVQ-VAE: Efficient 1-Dimensional Continuous Tokenizer
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
(Almost) Free Modality Stitching of Foundation Models
by: Singh, Jaisidh, et al.
Published: (2025)
by: Singh, Jaisidh, et al.
Published: (2025)
Pre-Training Multimodal Hallucination Detectors with Corrupted Grounding Data
by: Whitehead, Spencer, et al.
Published: (2024)
by: Whitehead, Spencer, et al.
Published: (2024)
Parallelized Autoregressive Visual Generation
by: Wang, Yuqing, et al.
Published: (2024)
by: Wang, Yuqing, et al.
Published: (2024)
Towards Better Cephalometric Landmark Detection with Diffusion Data Generation
by: Guo, Dongqian, et al.
Published: (2025)
by: Guo, Dongqian, et al.
Published: (2025)
AesRM: Improving Video Aesthetics with Expert-Level Feedback
by: Han, Yujin, et al.
Published: (2026)
by: Han, Yujin, et al.
Published: (2026)
Image Tokenizer Needs Post-Training
by: Qiu, Kai, et al.
Published: (2025)
by: Qiu, Kai, et al.
Published: (2025)
Bootstrapping Diffusion: Diffusion Model Training Leveraging Partial and Corrupted Data
by: Ma, Xudong
Published: (2025)
by: Ma, Xudong
Published: (2025)
What Makes "Good" Distractors for Object Hallucination Evaluation in Large Vision-Language Models?
by: Xie, Ming-Kun, et al.
Published: (2025)
by: Xie, Ming-Kun, et al.
Published: (2025)
Slightly Shift New Classes to Remember Old Classes for Video Class-Incremental Learning
by: Jiao, Jian, et al.
Published: (2024)
by: Jiao, Jian, et al.
Published: (2024)
Linear Combination of Saved Checkpoints Makes Consistency and Diffusion Models Better
by: Liu, Enshu, et al.
Published: (2024)
by: Liu, Enshu, et al.
Published: (2024)
MaskDiffusion: Exploiting Pre-trained Diffusion Models for Semantic Segmentation
by: Kawano, Yasufumi, et al.
Published: (2024)
by: Kawano, Yasufumi, et al.
Published: (2024)
QDFormer: Towards Robust Audiovisual Segmentation in Complex Environments with Quantization-based Semantic Decomposition
by: Li, Xiang, et al.
Published: (2023)
by: Li, Xiang, et al.
Published: (2023)
Is Pre-training Truly Better Than Meta-Learning?
by: Miranda, Brando, et al.
Published: (2023)
by: Miranda, Brando, et al.
Published: (2023)
Shortcutting Pre-trained Flow Matching Diffusion Models is Almost Free Lunch
by: Cai, Xu, et al.
Published: (2025)
by: Cai, Xu, et al.
Published: (2025)
Uncovering the Hidden Cost of Model Compression
by: Misra, Diganta, et al.
Published: (2023)
by: Misra, Diganta, et al.
Published: (2023)
Similar Items
-
Understanding and Mitigating the Label Noise in Pre-training on Downstream Tasks
by: Chen, Hao, et al.
Published: (2023) -
Masked Autoencoders Are Effective Tokenizers for Diffusion Models
by: Chen, Hao, et al.
Published: (2025) -
Corruption-Aware Training of Latent Video Diffusion Models for Robust Text-to-Video Generation
by: Maduabuchi, Chika, et al.
Published: (2025) -
Impact of Noisy Supervision in Foundation Model Learning
by: Chen, Hao, et al.
Published: (2024) -
Can Diffusion Models Learn Hidden Inter-Feature Rules Behind Images?
by: Han, Yujin, et al.
Published: (2025)