A Survey on Data Augmentation in Large Model Era
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Yue, Guo, Chenlu, Wang, Xu, Chang, Yi, Wu, Yuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Survey of Mix-based Data Augmentation: Taxonomy, Methods, Applications, and Explainability
von: Cao, Chengtai, et al.
Veröffentlicht: (2022)
von: Cao, Chengtai, et al.
Veröffentlicht: (2022)
A Survey on Hallucination in Large Vision-Language Models
von: Liu, Hanchao, et al.
Veröffentlicht: (2024)
von: Liu, Hanchao, et al.
Veröffentlicht: (2024)
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey
von: Qi, Yayun, et al.
Veröffentlicht: (2024)
von: Qi, Yayun, et al.
Veröffentlicht: (2024)
Hierarchical Safety Realignment: Lightweight Restoration of Safety in Pruned Large Vision-Language Models
von: Li, Yue, et al.
Veröffentlicht: (2025)
von: Li, Yue, et al.
Veröffentlicht: (2025)
Matryoshka Query Transformer for Large Vision-Language Models
von: Hu, Wenbo, et al.
Veröffentlicht: (2024)
von: Hu, Wenbo, et al.
Veröffentlicht: (2024)
A Survey on Multimodal Large Language Models
von: Yin, Shukang, et al.
Veröffentlicht: (2023)
von: Yin, Shukang, et al.
Veröffentlicht: (2023)
A Survey on Transformer Compression
von: Tang, Yehui, et al.
Veröffentlicht: (2024)
von: Tang, Yehui, et al.
Veröffentlicht: (2024)
Learning Self-Correction in Vision-Language Models via Rollout Augmentation
von: Ding, Yi, et al.
Veröffentlicht: (2026)
von: Ding, Yi, et al.
Veröffentlicht: (2026)
Towards Efficient Large Vision-Language Models: A Comprehensive Survey on Inference Strategies
von: Pathak, Surendra, et al.
Veröffentlicht: (2026)
von: Pathak, Surendra, et al.
Veröffentlicht: (2026)
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation
von: Yin, Shukang, et al.
Veröffentlicht: (2024)
von: Yin, Shukang, et al.
Veröffentlicht: (2024)
Contrastive Visual Data Augmentation
von: Zhou, Yu, et al.
Veröffentlicht: (2025)
von: Zhou, Yu, et al.
Veröffentlicht: (2025)
LaMI: Augmenting Large Language Models via Late Multi-Image Fusion
von: Yariv, Guy, et al.
Veröffentlicht: (2024)
von: Yariv, Guy, et al.
Veröffentlicht: (2024)
A Survey of Reasoning with Foundation Models
von: Sun, Jiankai, et al.
Veröffentlicht: (2023)
von: Sun, Jiankai, et al.
Veröffentlicht: (2023)
mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
von: Ye, Qinghao, et al.
Veröffentlicht: (2023)
von: Ye, Qinghao, et al.
Veröffentlicht: (2023)
DREAM: Improving Video-Text Retrieval Through Relevance-Based Augmentation Using Large Foundation Models
von: Wang, Yimu, et al.
Veröffentlicht: (2024)
von: Wang, Yimu, et al.
Veröffentlicht: (2024)
Multi-Task Model Merging via Adaptive Weight Disentanglement
von: Xiong, Feng, et al.
Veröffentlicht: (2024)
von: Xiong, Feng, et al.
Veröffentlicht: (2024)
Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models
von: Ma, Chuofan, et al.
Veröffentlicht: (2024)
von: Ma, Chuofan, et al.
Veröffentlicht: (2024)
RAP: Retrieval-Augmented Personalization for Multimodal Large Language Models
von: Hao, Haoran, et al.
Veröffentlicht: (2024)
von: Hao, Haoran, et al.
Veröffentlicht: (2024)
Deep Augmentation: Dropout as Augmentation for Self-Supervised Learning
von: Brüel-Gabrielsson, Rickard, et al.
Veröffentlicht: (2023)
von: Brüel-Gabrielsson, Rickard, et al.
Veröffentlicht: (2023)
C2-Evo: Co-Evolving Multimodal Data and Model for Self-Improving Reasoning
von: Chen, Xiuwei, et al.
Veröffentlicht: (2025)
von: Chen, Xiuwei, et al.
Veröffentlicht: (2025)
AI Security in the Foundation Model Era: A Comprehensive Survey from a Unified Perspective
von: Wang, Zhenyi, et al.
Veröffentlicht: (2026)
von: Wang, Zhenyi, et al.
Veröffentlicht: (2026)
GenSim: Generating Robotic Simulation Tasks via Large Language Models
von: Wang, Lirui, et al.
Veröffentlicht: (2023)
von: Wang, Lirui, et al.
Veröffentlicht: (2023)
Diffusion-RPO: Aligning Diffusion Models through Relative Preference Optimization
von: Gu, Yi, et al.
Veröffentlicht: (2024)
von: Gu, Yi, et al.
Veröffentlicht: (2024)
GUARD: Role-playing to Generate Natural-language Jailbreakings to Test Guideline Adherence of Large Language Models
von: Jin, Haibo, et al.
Veröffentlicht: (2024)
von: Jin, Haibo, et al.
Veröffentlicht: (2024)
Prophet: Prompting Large Language Models with Complementary Answer Heuristics for Knowledge-based Visual Question Answering
von: Yu, Zhou, et al.
Veröffentlicht: (2023)
von: Yu, Zhou, et al.
Veröffentlicht: (2023)
DeepCoT: Deep Continual Transformers for Real-Time Inference on Data Streams
von: Picón, Ginés Carreto, et al.
Veröffentlicht: (2025)
von: Picón, Ginés Carreto, et al.
Veröffentlicht: (2025)
FlagEval Findings Report: A Preliminary Evaluation of Large Reasoning Models on Automatically Verifiable Textual and Visual Questions
von: Qin, Bowen, et al.
Veröffentlicht: (2025)
von: Qin, Bowen, et al.
Veröffentlicht: (2025)
MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models
von: Xia, Peng, et al.
Veröffentlicht: (2024)
von: Xia, Peng, et al.
Veröffentlicht: (2024)
Lever LM: Configuring In-Context Sequence to Lever Large Vision Language Models
von: Yang, Xu, et al.
Veröffentlicht: (2023)
von: Yang, Xu, et al.
Veröffentlicht: (2023)
Logic-RAG: Augmenting Large Multimodal Models with Visual-Spatial Knowledge for Road Scene Understanding
von: Kabir, Imran, et al.
Veröffentlicht: (2025)
von: Kabir, Imran, et al.
Veröffentlicht: (2025)
OS Agents: A Survey on MLLM-based Agents for General Computing Devices Use
von: Hu, Xueyu, et al.
Veröffentlicht: (2025)
von: Hu, Xueyu, et al.
Veröffentlicht: (2025)
Vision-and-Language Navigation Today and Tomorrow: A Survey in the Era of Foundation Models
von: Zhang, Yue, et al.
Veröffentlicht: (2024)
von: Zhang, Yue, et al.
Veröffentlicht: (2024)
Aligning Modalities in Vision Large Language Models via Preference Fine-tuning
von: Zhou, Yiyang, et al.
Veröffentlicht: (2024)
von: Zhou, Yiyang, et al.
Veröffentlicht: (2024)
Adaptive Data Augmentation with Multi-armed Bandit: Sample-Efficient Embedding Calibration for Implicit Pattern Recognition
von: Tang, Minxue, et al.
Veröffentlicht: (2026)
von: Tang, Minxue, et al.
Veröffentlicht: (2026)
TLDR: Token-Level Detective Reward Model for Large Vision Language Models
von: Fu, Deqing, et al.
Veröffentlicht: (2024)
von: Fu, Deqing, et al.
Veröffentlicht: (2024)
Analyzing and Mitigating Object Hallucination in Large Vision-Language Models
von: Zhou, Yiyang, et al.
Veröffentlicht: (2023)
von: Zhou, Yiyang, et al.
Veröffentlicht: (2023)
Investigating Video Reasoning Capability of Large Language Models with Tropes in Movies
von: Su, Hung-Ting, et al.
Veröffentlicht: (2024)
von: Su, Hung-Ting, et al.
Veröffentlicht: (2024)
S-GRPO: Unified Post-Training for Large Vision-Language Models
von: Yan, Yuming, et al.
Veröffentlicht: (2026)
von: Yan, Yuming, et al.
Veröffentlicht: (2026)
SEASON: Mitigating Temporal Hallucination in Video Large Language Models via Self-Diagnostic Contrastive Decoding
von: Wu, Chang-Hsun, et al.
Veröffentlicht: (2025)
von: Wu, Chang-Hsun, et al.
Veröffentlicht: (2025)
MoTe: Learning Motion-Text Diffusion Model for Multiple Generation Tasks
von: Wu, Yiming, et al.
Veröffentlicht: (2024)
von: Wu, Yiming, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
A Survey of Mix-based Data Augmentation: Taxonomy, Methods, Applications, and Explainability
von: Cao, Chengtai, et al.
Veröffentlicht: (2022) -
A Survey on Hallucination in Large Vision-Language Models
von: Liu, Hanchao, et al.
Veröffentlicht: (2024) -
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey
von: Qi, Yayun, et al.
Veröffentlicht: (2024) -
Hierarchical Safety Realignment: Lightweight Restoration of Safety in Pruned Large Vision-Language Models
von: Li, Yue, et al.
Veröffentlicht: (2025) -
Matryoshka Query Transformer for Large Vision-Language Models
von: Hu, Wenbo, et al.
Veröffentlicht: (2024)