A Survey of Mix-based Data Augmentation: Taxonomy, Methods, Applications, and Explainability
Fuente:
arXiv
Saved in:
| Main Authors: | Cao, Chengtai, Zhou, Fan, Dai, Yurou, Wang, Jianping, Zhang, Kunpeng |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Survey on Data Augmentation in Large Model Era
by: Zhou, Yue, et al.
Published: (2024)
by: Zhou, Yue, et al.
Published: (2024)
Vision Mamba: A Comprehensive Survey and Taxonomy
by: Liu, Xiao, et al.
Published: (2024)
by: Liu, Xiao, et al.
Published: (2024)
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation
by: Yin, Shukang, et al.
Published: (2024)
by: Yin, Shukang, et al.
Published: (2024)
Enhancing Document Key Information Localization Through Data Augmentation
by: Dai, Yue
Published: (2025)
by: Dai, Yue
Published: (2025)
Hyperbolic Multimodal Representation Learning for Biological Taxonomies
by: Gong, ZeMing, et al.
Published: (2025)
by: Gong, ZeMing, et al.
Published: (2025)
TiMix: Text-aware Image Mixing for Effective Vision-Language Pre-training
by: Jiang, Chaoya, et al.
Published: (2023)
by: Jiang, Chaoya, et al.
Published: (2023)
Towards Transparent AI: A Survey on Explainable Large Language Models
by: Palikhe, Avash, et al.
Published: (2025)
by: Palikhe, Avash, et al.
Published: (2025)
Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities
by: Yang, Enneng, et al.
Published: (2024)
by: Yang, Enneng, et al.
Published: (2024)
Surveying the Landscape of Image Captioning Evaluation: A Comprehensive Taxonomy, Trends and Metrics Analysis
by: Berger, Uri, et al.
Published: (2024)
by: Berger, Uri, et al.
Published: (2024)
A Language Anchor-Guided Method for Robust Noisy Domain Generalization
by: Dai, Zilin, et al.
Published: (2025)
by: Dai, Zilin, et al.
Published: (2025)
Deep Augmentation: Dropout as Augmentation for Self-Supervised Learning
by: Brüel-Gabrielsson, Rickard, et al.
Published: (2023)
by: Brüel-Gabrielsson, Rickard, et al.
Published: (2023)
Adaptive Data Augmentation with Multi-armed Bandit: Sample-Efficient Embedding Calibration for Implicit Pattern Recognition
by: Tang, Minxue, et al.
Published: (2026)
by: Tang, Minxue, et al.
Published: (2026)
Contrastive Visual Data Augmentation
by: Zhou, Yu, et al.
Published: (2025)
by: Zhou, Yu, et al.
Published: (2025)
C2-Evo: Co-Evolving Multimodal Data and Model for Self-Improving Reasoning
by: Chen, Xiuwei, et al.
Published: (2025)
by: Chen, Xiuwei, et al.
Published: (2025)
A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations
by: Lan, Tian, et al.
Published: (2025)
by: Lan, Tian, et al.
Published: (2025)
Learning Self-Correction in Vision-Language Models via Rollout Augmentation
by: Ding, Yi, et al.
Published: (2026)
by: Ding, Yi, et al.
Published: (2026)
OS Agents: A Survey on MLLM-based Agents for General Computing Devices Use
by: Hu, Xueyu, et al.
Published: (2025)
by: Hu, Xueyu, et al.
Published: (2025)
CapsFusion: Rethinking Image-Text Data at Scale
by: Yu, Qiying, et al.
Published: (2023)
by: Yu, Qiying, et al.
Published: (2023)
DeepCoT: Deep Continual Transformers for Real-Time Inference on Data Streams
by: Picón, Ginés Carreto, et al.
Published: (2025)
by: Picón, Ginés Carreto, et al.
Published: (2025)
Large Vision-Language Model Alignment and Misalignment: A Survey Through the Lens of Explainability
by: Shu, Dong, et al.
Published: (2025)
by: Shu, Dong, et al.
Published: (2025)
Mamba in Vision: A Comprehensive Survey of Techniques and Applications
by: Rahman, Md Maklachur, et al.
Published: (2024)
by: Rahman, Md Maklachur, et al.
Published: (2024)
MixCut:A Data Augmentation Method for Facial Expression Recognition
by: Yu, Jiaxiang, et al.
Published: (2024)
by: Yu, Jiaxiang, et al.
Published: (2024)
A Survey on Transformer Compression
by: Tang, Yehui, et al.
Published: (2024)
by: Tang, Yehui, et al.
Published: (2024)
Few-shot Adaptation to Distribution Shifts By Mixing Source and Target Embeddings
by: Xue, Yihao, et al.
Published: (2023)
by: Xue, Yihao, et al.
Published: (2023)
ViTmiX: Vision Transformer Explainability Augmented by Mixed Visualization Methods
by: Hogea, Eduard, et al.
Published: (2024)
by: Hogea, Eduard, et al.
Published: (2024)
Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding
by: Gao, Sensen, et al.
Published: (2025)
by: Gao, Sensen, et al.
Published: (2025)
ICM-Assistant: Instruction-tuning Multimodal Large Language Models for Rule-based Explainable Image Content Moderation
by: Wu, Mengyang, et al.
Published: (2024)
by: Wu, Mengyang, et al.
Published: (2024)
The Cost of Context: Mitigating Textual Bias in Multimodal Retrieval-Augmented Generation
by: Jung, Hoin, et al.
Published: (2026)
by: Jung, Hoin, et al.
Published: (2026)
A Survey on Hallucination in Large Vision-Language Models
by: Liu, Hanchao, et al.
Published: (2024)
by: Liu, Hanchao, et al.
Published: (2024)
GeoMix: Towards Geometry-Aware Data Augmentation
by: Zhao, Wentao, et al.
Published: (2024)
by: Zhao, Wentao, et al.
Published: (2024)
RadZero: Similarity-Based Cross-Attention for Explainable Vision-Language Alignment in Chest X-ray with Zero-Shot Multi-Task Capability
by: Park, Jonggwon, et al.
Published: (2025)
by: Park, Jonggwon, et al.
Published: (2025)
Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization
by: Wang, Weiyun, et al.
Published: (2024)
by: Wang, Weiyun, et al.
Published: (2024)
MM-Verify: Enhancing Multimodal Reasoning with Chain-of-Thought Verification
by: Sun, Linzhuang, et al.
Published: (2025)
by: Sun, Linzhuang, et al.
Published: (2025)
Survey of Large Multimodal Model Datasets, Application Categories and Taxonomy
by: Pattnayak, Priyaranjan, et al.
Published: (2024)
by: Pattnayak, Priyaranjan, et al.
Published: (2024)
Residual-based Language Models are Free Boosters for Biomedical Imaging
by: Lai, Zhixin, et al.
Published: (2024)
by: Lai, Zhixin, et al.
Published: (2024)
UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents
by: Xiao, Han, et al.
Published: (2025)
by: Xiao, Han, et al.
Published: (2025)
RA-RRG: Multimodal Retrieval-Augmented Radiology Report Generation with Key Phrase Extraction
by: Park, Jonggwon, et al.
Published: (2025)
by: Park, Jonggwon, et al.
Published: (2025)
LaMI: Augmenting Large Language Models via Late Multi-Image Fusion
by: Yariv, Guy, et al.
Published: (2024)
by: Yariv, Guy, et al.
Published: (2024)
Survey of Video Diffusion Models: Foundations, Implementations, and Applications
by: Wang, Yimu, et al.
Published: (2025)
by: Wang, Yimu, et al.
Published: (2025)
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
by: McKinzie, Brandon, et al.
Published: (2024)
by: McKinzie, Brandon, et al.
Published: (2024)
Similar Items
-
A Survey on Data Augmentation in Large Model Era
by: Zhou, Yue, et al.
Published: (2024) -
Vision Mamba: A Comprehensive Survey and Taxonomy
by: Liu, Xiao, et al.
Published: (2024) -
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation
by: Yin, Shukang, et al.
Published: (2024) -
Enhancing Document Key Information Localization Through Data Augmentation
by: Dai, Yue
Published: (2025) -
Hyperbolic Multimodal Representation Learning for Biological Taxonomies
by: Gong, ZeMing, et al.
Published: (2025)