Enhanced Continual Learning of Vision-Language Models with Model Fusion
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gao, Haoyuan, Zhang, Zicong, Wei, Yuqi, Zhao, Linglan, Li, Guilin, Li, Yexin, Wang, Bo, Kong, Linghe, Huang, Weiran |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
IDER: IDempotent Experience Replay for Reliable Continual Learning
von: Liu, Zhanwang, et al.
Veröffentlicht: (2026)
von: Liu, Zhanwang, et al.
Veröffentlicht: (2026)
SAFE: Slow and Fast Parameter-Efficient Tuning for Continual Learning with Pre-Trained Models
von: Zhao, Linglan, et al.
Veröffentlicht: (2024)
von: Zhao, Linglan, et al.
Veröffentlicht: (2024)
Revisiting Visual Understanding in Multimodal Reasoning through a Lens of Image Perturbation
von: Li, Yuting, et al.
Veröffentlicht: (2025)
von: Li, Yuting, et al.
Veröffentlicht: (2025)
TransMed: Large Language Models Enhance Vision Transformer for Biomedical Image Classification
von: Zheng, Kaipeng, et al.
Veröffentlicht: (2023)
von: Zheng, Kaipeng, et al.
Veröffentlicht: (2023)
Generalized Category Discovery via Reciprocal Learning and Class-Wise Distribution Regularization
von: Liu, Duo, et al.
Veröffentlicht: (2025)
von: Liu, Duo, et al.
Veröffentlicht: (2025)
VEQ: Modality-Adaptive Quantization for MoE Vision-Language Models
von: Qin, Guangshuo, et al.
Veröffentlicht: (2026)
von: Qin, Guangshuo, et al.
Veröffentlicht: (2026)
MLLM-Enhanced Face Forgery Detection: A Vision-Language Fusion Solution
von: Peng, Siran, et al.
Veröffentlicht: (2025)
von: Peng, Siran, et al.
Veröffentlicht: (2025)
SynArtifact: Classifying and Alleviating Artifacts in Synthetic Images via Vision-Language Model
von: Cao, Bin, et al.
Veröffentlicht: (2024)
von: Cao, Bin, et al.
Veröffentlicht: (2024)
First SFT, Second RL, Third UPT: Continual Improving Multi-Modal LLM Reasoning via Unsupervised Post-Training
von: Wei, Lai, et al.
Veröffentlicht: (2025)
von: Wei, Lai, et al.
Veröffentlicht: (2025)
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion
von: Chen, Jiuhai, et al.
Veröffentlicht: (2024)
von: Chen, Jiuhai, et al.
Veröffentlicht: (2024)
Understanding Degradation with Vision Language Model
von: Lan, Guanzhou, et al.
Veröffentlicht: (2026)
von: Lan, Guanzhou, et al.
Veröffentlicht: (2026)
Branch, or Layer? Zeroth-Order Optimization for Continual Learning of Vision-Language Models
von: Liu, Ziwei, et al.
Veröffentlicht: (2025)
von: Liu, Ziwei, et al.
Veröffentlicht: (2025)
Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start
von: Wei, Lai, et al.
Veröffentlicht: (2025)
von: Wei, Lai, et al.
Veröffentlicht: (2025)
Hierarchical Fusion and Joint Aggregation: A Multi-Level Feature Representation Method for AIGC Image Quality Assessment
von: Meng, Linghe, et al.
Veröffentlicht: (2025)
von: Meng, Linghe, et al.
Veröffentlicht: (2025)
AutoEval-Video: An Automatic Benchmark for Assessing Large Vision Language Models in Open-Ended Video Question Answering
von: Chen, Xiuyuan, et al.
Veröffentlicht: (2023)
von: Chen, Xiuyuan, et al.
Veröffentlicht: (2023)
Let's Reward Step-by-Step: Step-Aware Contrastive Alignment for Vision-Language Navigation in Continuous Environments
von: Li, Haoyuan, et al.
Veröffentlicht: (2026)
von: Li, Haoyuan, et al.
Veröffentlicht: (2026)
Unveiling the Tapestry of Consistency in Large Vision-Language Models
von: Zhang, Yuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuan, et al.
Veröffentlicht: (2024)
Enhancing Continual Learning of Vision-Language Models via Dynamic Prefix Weighting
von: Jang, Hyeonseo, et al.
Veröffentlicht: (2026)
von: Jang, Hyeonseo, et al.
Veröffentlicht: (2026)
iGSP:Implicit Gradient Subspace Projection for Efficient Continual Learning of Vision-Language Models
von: Cui, Xuezhi, et al.
Veröffentlicht: (2026)
von: Cui, Xuezhi, et al.
Veröffentlicht: (2026)
Enhancing Robustness of Vision-Language Models through Orthogonality Learning and Self-Regularization
von: Li, Jinlong, et al.
Veröffentlicht: (2024)
von: Li, Jinlong, et al.
Veröffentlicht: (2024)
Image Fusion via Vision-Language Model
von: Zhao, Zixiang, et al.
Veröffentlicht: (2024)
von: Zhao, Zixiang, et al.
Veröffentlicht: (2024)
MM-Skin: Enhancing Dermatology Vision-Language Model with an Image-Text Dataset Derived from Textbooks
von: Zeng, Wenqi, et al.
Veröffentlicht: (2025)
von: Zeng, Wenqi, et al.
Veröffentlicht: (2025)
Unified Vision-Language-Action Model
von: Wang, Yuqi, et al.
Veröffentlicht: (2025)
von: Wang, Yuqi, et al.
Veröffentlicht: (2025)
DualDiff: Dual-branch Diffusion Model for Autonomous Driving with Semantic Fusion
von: Li, Haoteng, et al.
Veröffentlicht: (2025)
von: Li, Haoteng, et al.
Veröffentlicht: (2025)
Mind the Interference: Retaining Pre-trained Knowledge in Parameter Efficient Continual Learning of Vision-Language Models
von: Tang, Longxiang, et al.
Veröffentlicht: (2024)
von: Tang, Longxiang, et al.
Veröffentlicht: (2024)
Improving SAM for Camouflaged Object Detection via Dual Stream Adapters
von: Liu, Jiaming, et al.
Veröffentlicht: (2025)
von: Liu, Jiaming, et al.
Veröffentlicht: (2025)
Hierarchical Dual-Subspace Decoupling for Continual Learning in Vision-Language Models
von: Qin, Mengxin, et al.
Veröffentlicht: (2026)
von: Qin, Mengxin, et al.
Veröffentlicht: (2026)
GLAD: Generalizable Tuning for Vision-Language Models
von: Peng, Yuqi, et al.
Veröffentlicht: (2025)
von: Peng, Yuqi, et al.
Veröffentlicht: (2025)
Modular Prompt Learning Improves Vision-Language Models
von: Huang, Zhenhan, et al.
Veröffentlicht: (2025)
von: Huang, Zhenhan, et al.
Veröffentlicht: (2025)
Fose: Fusion of One-Step Diffusion and End-to-End Network for Pansharpening
von: Liu, Kai, et al.
Veröffentlicht: (2025)
von: Liu, Kai, et al.
Veröffentlicht: (2025)
MERGETUNE: Continued Fine-Tuning of Vision-Language Models
von: Wang, Wenqing, et al.
Veröffentlicht: (2026)
von: Wang, Wenqing, et al.
Veröffentlicht: (2026)
From Captions to Rewards (CAREVL): Leveraging Large Language Model Experts for Enhanced Reward Modeling in Large Vision-Language Models
von: Dai, Muzhi, et al.
Veröffentlicht: (2025)
von: Dai, Muzhi, et al.
Veröffentlicht: (2025)
Evolution-based Region Adversarial Prompt Learning for Robustness Enhancement in Vision-Language Models
von: Jia, Xiaojun, et al.
Veröffentlicht: (2025)
von: Jia, Xiaojun, et al.
Veröffentlicht: (2025)
ReCAD: Reinforcement Learning Enhanced Parametric CAD Model Generation with Vision-Language Models
von: Li, Jiahao, et al.
Veröffentlicht: (2025)
von: Li, Jiahao, et al.
Veröffentlicht: (2025)
Visual-Advantage On-Policy Distillation for Vision-Language Models
von: Liu, Ruiqi, et al.
Veröffentlicht: (2026)
von: Liu, Ruiqi, et al.
Veröffentlicht: (2026)
Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
von: Li, Lei, et al.
Veröffentlicht: (2024)
von: Li, Lei, et al.
Veröffentlicht: (2024)
Active Learning via Vision-Language Model Adaptation with Open Data
von: Wang, Tong, et al.
Veröffentlicht: (2025)
von: Wang, Tong, et al.
Veröffentlicht: (2025)
CoViPAL: Layer-wise Contextualized Visual Token Pruning for Large Vision-Language Models
von: Tang, Zicong, et al.
Veröffentlicht: (2025)
von: Tang, Zicong, et al.
Veröffentlicht: (2025)
Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision
von: Han, Xiaofeng, et al.
Veröffentlicht: (2025)
von: Han, Xiaofeng, et al.
Veröffentlicht: (2025)
ECMF: Enhanced Cross-Modal Fusion for Multimodal Emotion Recognition in MER-SEMI Challenge
von: Hu, Juewen, et al.
Veröffentlicht: (2025)
von: Hu, Juewen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
IDER: IDempotent Experience Replay for Reliable Continual Learning
von: Liu, Zhanwang, et al.
Veröffentlicht: (2026) -
SAFE: Slow and Fast Parameter-Efficient Tuning for Continual Learning with Pre-Trained Models
von: Zhao, Linglan, et al.
Veröffentlicht: (2024) -
Revisiting Visual Understanding in Multimodal Reasoning through a Lens of Image Perturbation
von: Li, Yuting, et al.
Veröffentlicht: (2025) -
TransMed: Large Language Models Enhance Vision Transformer for Biomedical Image Classification
von: Zheng, Kaipeng, et al.
Veröffentlicht: (2023) -
Generalized Category Discovery via Reciprocal Learning and Class-Wise Distribution Regularization
von: Liu, Duo, et al.
Veröffentlicht: (2025)