Continual Learning in Vision-Language Models via Aligned Model Merging
Fuente:
arXiv
Saved in:
| Main Authors: | Sokar, Ghada, Dziugaite, Gintare Karolina, Arnab, Anurag, Iscen, Ahmet, Castro, Pablo Samuel, Schmid, Cordelia |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Retrieval-Enhanced Contrastive Vision-Text Models
by: Iscen, Ahmet, et al.
Published: (2023)
by: Iscen, Ahmet, et al.
Published: (2023)
Web-Scale Visual Entity Recognition: An LLM-Driven Data Approach
by: Caron, Mathilde, et al.
Published: (2024)
by: Caron, Mathilde, et al.
Published: (2024)
A Generative Approach for Wikipedia-Scale Visual Entity Recognition
by: Caron, Mathilde, et al.
Published: (2024)
by: Caron, Mathilde, et al.
Published: (2024)
OVFact: Measuring and Improving Open-Vocabulary Factuality for Long Caption Models
by: Wysoczańska, Monika, et al.
Published: (2025)
by: Wysoczańska, Monika, et al.
Published: (2025)
Memory-Modular Classification: Learning to Generalize with Memory Replacement
by: Kang, Dahyun, et al.
Published: (2025)
by: Kang, Dahyun, et al.
Published: (2025)
Dense Video Object Captioning from Disjoint Supervision
by: Zhou, Xingyi, et al.
Published: (2023)
by: Zhou, Xingyi, et al.
Published: (2023)
VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
by: Bousselham, Walid, et al.
Published: (2025)
by: Bousselham, Walid, et al.
Published: (2025)
Identifying Spurious Biases Early in Training through the Lens of Simplicity Bias
by: Yang, Yu, et al.
Published: (2023)
by: Yang, Yu, et al.
Published: (2023)
CAViAR: Critic-Augmented Video Agentic Reasoning
by: Menon, Sachit, et al.
Published: (2025)
by: Menon, Sachit, et al.
Published: (2025)
Time-, Memory- and Parameter-Efficient Visual Adaptation
by: Mercea, Otniel-Bogdan, et al.
Published: (2024)
by: Mercea, Otniel-Bogdan, et al.
Published: (2024)
PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction
by: Chen, Shizhe, et al.
Published: (2026)
by: Chen, Shizhe, et al.
Published: (2026)
What Are You Doing? A Closer Look at Controllable Human Video Generation
by: Bugliarello, Emanuele, et al.
Published: (2025)
by: Bugliarello, Emanuele, et al.
Published: (2025)
Towards Generalizable Vision-Language Robotic Manipulation: A Benchmark and LLM-guided 3D Policy
by: Garcia, Ricardo, et al.
Published: (2024)
by: Garcia, Ricardo, et al.
Published: (2024)
Learning Correlation Structures for Vision Transformers
by: Kim, Manjin, et al.
Published: (2024)
by: Kim, Manjin, et al.
Published: (2024)
Audiovisual Masked Autoencoders
by: Georgescu, Mariana-Iuliana, et al.
Published: (2022)
by: Georgescu, Mariana-Iuliana, et al.
Published: (2022)
Scaling Cross-Environment Failure Reasoning Data for Vision-Language Robotic Manipulation
by: Pacaud, Paul, et al.
Published: (2025)
by: Pacaud, Paul, et al.
Published: (2025)
VoCap: Video Object Captioning and Segmentation from Any Prompt
by: Uijlings, Jasper, et al.
Published: (2025)
by: Uijlings, Jasper, et al.
Published: (2025)
Streaming Dense Video Captioning
by: Zhou, Xingyi, et al.
Published: (2024)
by: Zhou, Xingyi, et al.
Published: (2024)
Learning text-to-video retrieval from image captioning
by: Ventura, Lucas, et al.
Published: (2024)
by: Ventura, Lucas, et al.
Published: (2024)
Gondola: Grounded Vision Language Planning for Generalizable Robotic Manipulation
by: Chen, Shizhe, et al.
Published: (2025)
by: Chen, Shizhe, et al.
Published: (2025)
MergeSlide: Continual Model Merging and Task-to-Class Prompt-Aligned Inference for Lifelong Learning on Whole Slide Images
by: Bui, Doanh C., et al.
Published: (2025)
by: Bui, Doanh C., et al.
Published: (2025)
BrickNet: Graph-Backed Generative Brick Assembly
by: Kulits, Peter, et al.
Published: (2026)
by: Kulits, Peter, et al.
Published: (2026)
From Dormant to Deleted: Tamper-Resistant Unlearning Through Weight-Space Regularization
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2025)
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2025)
SceneCraft: An LLM Agent for Synthesizing 3D Scene as Blender Code
by: Hu, Ziniu, et al.
Published: (2024)
by: Hu, Ziniu, et al.
Published: (2024)
Convolutional Prompting meets Language Models for Continual Learning
by: Roy, Anurag, et al.
Published: (2024)
by: Roy, Anurag, et al.
Published: (2024)
Grounded Video Caption Generation
by: Kazakos, Evangelos, et al.
Published: (2024)
by: Kazakos, Evangelos, et al.
Published: (2024)
ComposeAnything: Composite Object Priors for Text-to-Image Generation
by: Khan, Zeeshan, et al.
Published: (2025)
by: Khan, Zeeshan, et al.
Published: (2025)
Large-scale Pre-training for Grounded Video Caption Generation
by: Kazakos, Evangelos, et al.
Published: (2025)
by: Kazakos, Evangelos, et al.
Published: (2025)
Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames
by: Arnab, Anurag, et al.
Published: (2025)
by: Arnab, Anurag, et al.
Published: (2025)
AMES: Asymmetric and Memory-Efficient Similarity Estimation for Instance-level Retrieval
by: Suma, Pavel, et al.
Published: (2024)
by: Suma, Pavel, et al.
Published: (2024)
ViViDex: Learning Vision-based Dexterous Manipulation from Human Videos
by: Chen, Zerui, et al.
Published: (2024)
by: Chen, Zerui, et al.
Published: (2024)
Interpretable Zero-Shot Learning with Locally-Aligned Vision-Language Model
by: Chen, Shiming, et al.
Published: (2025)
by: Chen, Shiming, et al.
Published: (2025)
MINERVA: Evaluating Complex Video Reasoning
by: Nagrani, Arsha, et al.
Published: (2025)
by: Nagrani, Arsha, et al.
Published: (2025)
Dense Optical Tracking: Connecting the Dots
by: Moing, Guillaume Le, et al.
Published: (2023)
by: Moing, Guillaume Le, et al.
Published: (2023)
Enhanced Continual Learning of Vision-Language Models with Model Fusion
by: Gao, Haoyuan, et al.
Published: (2025)
by: Gao, Haoyuan, et al.
Published: (2025)
Enhancing Continual Learning of Vision-Language Models via Dynamic Prefix Weighting
by: Jang, Hyeonseo, et al.
Published: (2026)
by: Jang, Hyeonseo, et al.
Published: (2026)
Boosting Continual Learning of Vision-Language Models via Mixture-of-Experts Adapters
by: Yu, Jiazuo, et al.
Published: (2024)
by: Yu, Jiazuo, et al.
Published: (2024)
Towards Efficient Vision State Space Models via Token Merging
by: Park, Jinyoung, et al.
Published: (2025)
by: Park, Jinyoung, et al.
Published: (2025)
SUGAR: Pre-training 3D Visual Representations for Robotics
by: Chen, Shizhe, et al.
Published: (2024)
by: Chen, Shizhe, et al.
Published: (2024)
Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs
by: Ventura, Lucas, et al.
Published: (2025)
by: Ventura, Lucas, et al.
Published: (2025)
Similar Items
-
Retrieval-Enhanced Contrastive Vision-Text Models
by: Iscen, Ahmet, et al.
Published: (2023) -
Web-Scale Visual Entity Recognition: An LLM-Driven Data Approach
by: Caron, Mathilde, et al.
Published: (2024) -
A Generative Approach for Wikipedia-Scale Visual Entity Recognition
by: Caron, Mathilde, et al.
Published: (2024) -
OVFact: Measuring and Improving Open-Vocabulary Factuality for Long Caption Models
by: Wysoczańska, Monika, et al.
Published: (2025) -
Memory-Modular Classification: Learning to Generalize with Memory Replacement
by: Kang, Dahyun, et al.
Published: (2025)