Batch Augmentation with Unimodal Fine-tuning for Multimodal Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Kabir, H M Dipu, Mondal, Subrota Kumar, Moni, Mohammad Ali |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reduction of Class Activation Uncertainty with Background Information
by: Kabir, H M Dipu
Published: (2023)
by: Kabir, H M Dipu
Published: (2023)
Hybrid Quantum-MambaVision: A Quantum-enhanced State Space Model for Calibrated Mixed-type Wafer Defect Detection
by: Sahoo, Satwik Sai Prakash, et al.
Published: (2026)
by: Sahoo, Satwik Sai Prakash, et al.
Published: (2026)
Multimodal Representation Learning by Alternating Unimodal Adaptation
by: Zhang, Xiaohui, et al.
Published: (2023)
by: Zhang, Xiaohui, et al.
Published: (2023)
Harnessing Frozen Unimodal Encoders for Flexible Multimodal Alignment
by: Maniparambil, Mayug, et al.
Published: (2024)
by: Maniparambil, Mayug, et al.
Published: (2024)
Robo-MUTUAL: Robotic Multimodal Task Specification via Unimodal Learning
by: Li, Jianxiong, et al.
Published: (2024)
by: Li, Jianxiong, et al.
Published: (2024)
UMCL: Unimodal-generated Multimodal Contrastive Learning for Cross-compression-rate Deepfake Detection
by: Lai, Ching-Yi, et al.
Published: (2025)
by: Lai, Ching-Yi, et al.
Published: (2025)
Robust Defense Strategies for Multimodal Contrastive Learning: Efficient Fine-tuning Against Backdoor Attacks
by: Hossain, Md. Iqbal, et al.
Published: (2025)
by: Hossain, Md. Iqbal, et al.
Published: (2025)
Localization vs. Semantics: Visual Representations in Unimodal and Multimodal Models
by: Li, Zhuowan, et al.
Published: (2022)
by: Li, Zhuowan, et al.
Published: (2022)
Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions
by: Li, Juncheng, et al.
Published: (2023)
by: Li, Juncheng, et al.
Published: (2023)
Assessing and Learning Alignment of Unimodal Vision and Language Models
by: Zhang, Le, et al.
Published: (2024)
by: Zhang, Le, et al.
Published: (2024)
A Unified Framework with Multimodal Fine-tuning for Remote Sensing Semantic Segmentation
by: Ma, Xianping, et al.
Published: (2024)
by: Ma, Xianping, et al.
Published: (2024)
MMPareto: Boosting Multimodal Learning with Innocent Unimodal Assistance
by: Wei, Yake, et al.
Published: (2024)
by: Wei, Yake, et al.
Published: (2024)
Learning Robust Anymodal Segmentor with Unimodal and Cross-modal Distillation
by: Zheng, Xu, et al.
Published: (2024)
by: Zheng, Xu, et al.
Published: (2024)
Missing Modality Prediction for Unpaired Multimodal Learning via Joint Embedding of Unimodal Models
by: Kim, Donggeun, et al.
Published: (2024)
by: Kim, Donggeun, et al.
Published: (2024)
Fine-Grained VLM Fine-tuning via Latent Hierarchical Adapter Learning
by: Zhao, Yumiao, et al.
Published: (2025)
by: Zhao, Yumiao, et al.
Published: (2025)
Foundation Models in Remote Sensing: Evolving from Unimodality to Multimodality
by: Hong, Danfeng, et al.
Published: (2026)
by: Hong, Danfeng, et al.
Published: (2026)
Adversarial Batch Representation Augmentation for Batch Correction in High-Content Cellular Screening
by: Tong, Lei, et al.
Published: (2026)
by: Tong, Lei, et al.
Published: (2026)
Connect Later: Improving Fine-tuning for Robustness with Targeted Augmentations
by: Qu, Helen, et al.
Published: (2024)
by: Qu, Helen, et al.
Published: (2024)
CoVFT: Context-aware Visual Fine-tuning for Multimodal Large Language Models
by: Zhou, Nan, et al.
Published: (2026)
by: Zhou, Nan, et al.
Published: (2026)
Unimodal and Multimodal Static Facial Expression Recognition for Virtual Reality Users with EmoHeVRDB
by: Ortmann, Thorben, et al.
Published: (2024)
by: Ortmann, Thorben, et al.
Published: (2024)
Quantifying and Mitigating Unimodal Biases in Multimodal Large Language Models: A Causal Perspective
by: Chen, Meiqi, et al.
Published: (2024)
by: Chen, Meiqi, et al.
Published: (2024)
DaMo: Data Mixing Optimizer in Fine-tuning Multimodal LLMs for Mobile Phone Agents
by: Shi, Kai, et al.
Published: (2025)
by: Shi, Kai, et al.
Published: (2025)
Attention-guided Fine-tuning of Multimodal Large Language Models Improves Chain-of-Thought Reasoning
by: Sinha, Sanchit, et al.
Published: (2026)
by: Sinha, Sanchit, et al.
Published: (2026)
Better Together: Leveraging Unpaired Multimodal Data for Stronger Unimodal Models
by: Gupta, Sharut, et al.
Published: (2025)
by: Gupta, Sharut, et al.
Published: (2025)
Enhancing Fine-grained Image Classification through Attentive Batch Training
by: Le, Duy M., et al.
Published: (2024)
by: Le, Duy M., et al.
Published: (2024)
Fine-tuning Vision Classifiers On A Budget
by: Kumar, Sunil, et al.
Published: (2024)
by: Kumar, Sunil, et al.
Published: (2024)
Semantic-aware Adversarial Fine-tuning for CLIP
by: Zhang, Jiacheng, et al.
Published: (2026)
by: Zhang, Jiacheng, et al.
Published: (2026)
Using Computer Vision for Skin Disease Diagnosis in Bangladesh Enhancing Interpretability and Transparency in Deep Learning Models for Skin Cancer Classification
by: Islam, Rafiul, et al.
Published: (2025)
by: Islam, Rafiul, et al.
Published: (2025)
Coffee: Controllable Diffusion Fine-tuning
by: Zeng, Ziyao, et al.
Published: (2025)
by: Zeng, Ziyao, et al.
Published: (2025)
Singular Value Fine-tuning for Few-Shot Class-Incremental Learning
by: Wang, Zhiwu, et al.
Published: (2025)
by: Wang, Zhiwu, et al.
Published: (2025)
Semantic Alignment of Unimodal Medical Text and Vision Representations
by: Di Folco, Maxime, et al.
Published: (2025)
by: Di Folco, Maxime, et al.
Published: (2025)
Beyond Unimodal Boundaries: Generative Recommendation with Multimodal Semantics
by: Zhu, Jing, et al.
Published: (2025)
by: Zhu, Jing, et al.
Published: (2025)
Prototypical Contrastive Learning-based CLIP Fine-tuning for Object Re-identification
by: Li, Jiachen, et al.
Published: (2023)
by: Li, Jiachen, et al.
Published: (2023)
MultiModal Fine-tuning with Synthetic Captions
by: Enomoto, Shohei, et al.
Published: (2026)
by: Enomoto, Shohei, et al.
Published: (2026)
Beyond Augmentation: Leveraging Inter-Instance Relation in Self-Supervised Representation Learning
by: Javidani, Ali, et al.
Published: (2025)
by: Javidani, Ali, et al.
Published: (2025)
Less Is More: An Explainable AI Framework for Lightweight Malaria Classification
by: Kafi, Md Abdullah Al, et al.
Published: (2025)
by: Kafi, Md Abdullah Al, et al.
Published: (2025)
Patch-Wise Self-Supervised Visual Representation Learning: A Fine-Grained Approach
by: Javidani, Ali, et al.
Published: (2023)
by: Javidani, Ali, et al.
Published: (2023)
FedMLLM: Federated Fine-tuning MLLM on Multimodal Heterogeneity Data
by: Xu, Binqian, et al.
Published: (2024)
by: Xu, Binqian, et al.
Published: (2024)
KDC-Diff: A Latent-Aware Diffusion Model with Knowledge Retention for Memory-Efficient Image Generation
by: Borno, Md. Naimur Asif, et al.
Published: (2025)
by: Borno, Md. Naimur Asif, et al.
Published: (2025)
Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment
by: Hong, Joanna, et al.
Published: (2025)
by: Hong, Joanna, et al.
Published: (2025)
Similar Items
-
Reduction of Class Activation Uncertainty with Background Information
by: Kabir, H M Dipu
Published: (2023) -
Hybrid Quantum-MambaVision: A Quantum-enhanced State Space Model for Calibrated Mixed-type Wafer Defect Detection
by: Sahoo, Satwik Sai Prakash, et al.
Published: (2026) -
Multimodal Representation Learning by Alternating Unimodal Adaptation
by: Zhang, Xiaohui, et al.
Published: (2023) -
Harnessing Frozen Unimodal Encoders for Flexible Multimodal Alignment
by: Maniparambil, Mayug, et al.
Published: (2024) -
Robo-MUTUAL: Robotic Multimodal Task Specification via Unimodal Learning
by: Li, Jianxiong, et al.
Published: (2024)