Mema: Memory-Augmented Adapter for Enhanced Vision-Language Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Ying, Han, Yudong, Shi, Kean, Pan, Liyuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding
von: Han, Yudong, et al.
Veröffentlicht: (2024)
von: Han, Yudong, et al.
Veröffentlicht: (2024)
Test-time Alignment-Enhanced Adapter for Vision-Language Models
von: Tong, Baoshun, et al.
Veröffentlicht: (2024)
von: Tong, Baoshun, et al.
Veröffentlicht: (2024)
Visual Enhanced Depth Scaling for Multimodal Latent Reasoning
von: Han, Yudong, et al.
Veröffentlicht: (2026)
von: Han, Yudong, et al.
Veröffentlicht: (2026)
Meta-Adapter: An Online Few-shot Learner for Vision-Language Model
von: Cheng, Cheng, et al.
Veröffentlicht: (2023)
von: Cheng, Cheng, et al.
Veröffentlicht: (2023)
CLIP-Adapter: Better Vision-Language Models with Feature Adapters
von: Gao, Peng, et al.
Veröffentlicht: (2021)
von: Gao, Peng, et al.
Veröffentlicht: (2021)
T-Gated Adapter: A Lightweight Temporal Adapter for Vision-Language Medical Segmentation
von: Khadka, Pranjal
Veröffentlicht: (2026)
von: Khadka, Pranjal
Veröffentlicht: (2026)
Robust Calibration of Large Vision-Language Adapters
von: Murugesan, Balamurali, et al.
Veröffentlicht: (2024)
von: Murugesan, Balamurali, et al.
Veröffentlicht: (2024)
ScVLM: Enhancing Vision-Language Model for Safety-Critical Event Understanding
von: Shi, Liang, et al.
Veröffentlicht: (2024)
von: Shi, Liang, et al.
Veröffentlicht: (2024)
BrepLLM: Native Boundary Representation Understanding with Large Language Models
von: Deng, Liyuan, et al.
Veröffentlicht: (2025)
von: Deng, Liyuan, et al.
Veröffentlicht: (2025)
Memory-Augmented Vision-Language Agents for Persistent and Semantically Consistent Object Captioning
von: Galliena, Tommaso, et al.
Veröffentlicht: (2026)
von: Galliena, Tommaso, et al.
Veröffentlicht: (2026)
HeGraphAdapter: Tuning Multi-Modal Vision-Language Models with Heterogeneous Graph Adapter
von: Zhao, Yumiao, et al.
Veröffentlicht: (2024)
von: Zhao, Yumiao, et al.
Veröffentlicht: (2024)
ArtGPT-4: Towards Artistic-understanding Large Vision-Language Models with Enhanced Adapter
von: Yuan, Zhengqing, et al.
Veröffentlicht: (2023)
von: Yuan, Zhengqing, et al.
Veröffentlicht: (2023)
Vision Language Models for Spreadsheet Understanding: Challenges and Opportunities
von: Xia, Shiyu, et al.
Veröffentlicht: (2024)
von: Xia, Shiyu, et al.
Veröffentlicht: (2024)
HeBA: Heterogeneous Bottleneck Adapters for Robust Vision-Language Models
von: Islam, Md Jahidul
Veröffentlicht: (2026)
von: Islam, Md Jahidul
Veröffentlicht: (2026)
Beyond CLIP Generalization: Against Forward&Backward Forgetting Adapter for Continual Learning of Vision-Language Models
von: Dong, Songlin, et al.
Veröffentlicht: (2025)
von: Dong, Songlin, et al.
Veröffentlicht: (2025)
Memory Efficient Transformer Adapter for Dense Predictions
von: Zhang, Dong, et al.
Veröffentlicht: (2025)
von: Zhang, Dong, et al.
Veröffentlicht: (2025)
RE-VLM: Event-Augmented Vision-Language Model for Scene Understanding
von: Liu, Hanqing, et al.
Veröffentlicht: (2026)
von: Liu, Hanqing, et al.
Veröffentlicht: (2026)
Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes
von: Rivera, Antonio Carlos, et al.
Veröffentlicht: (2024)
von: Rivera, Antonio Carlos, et al.
Veröffentlicht: (2024)
MARC: Memory-Augmented RL Token Compression for Efficient Video Understanding
von: Wu, Peiran, et al.
Veröffentlicht: (2025)
von: Wu, Peiran, et al.
Veröffentlicht: (2025)
Vision-Language Memory for Spatial Reasoning
von: Liu, Zuntao, et al.
Veröffentlicht: (2025)
von: Liu, Zuntao, et al.
Veröffentlicht: (2025)
NODE-Adapter: Neural Ordinary Differential Equations for Better Vision-Language Reasoning
von: Zhang, Yi, et al.
Veröffentlicht: (2024)
von: Zhang, Yi, et al.
Veröffentlicht: (2024)
Boosting Continual Learning of Vision-Language Models via Mixture-of-Experts Adapters
von: Yu, Jiazuo, et al.
Veröffentlicht: (2024)
von: Yu, Jiazuo, et al.
Veröffentlicht: (2024)
Event Camera Data Dense Pre-training
von: Yang, Yan, et al.
Veröffentlicht: (2023)
von: Yang, Yan, et al.
Veröffentlicht: (2023)
Memory-enhanced Retrieval Augmentation for Long Video Understanding
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
Enhancing Vision-Language Compositional Understanding with Multimodal Synthetic Data
von: Li, Haoxin, et al.
Veröffentlicht: (2025)
von: Li, Haoxin, et al.
Veröffentlicht: (2025)
CAD: Memory Efficient Convolutional Adapter for Segment Anything
von: Kim, Joohyeok, et al.
Veröffentlicht: (2024)
von: Kim, Joohyeok, et al.
Veröffentlicht: (2024)
Enhancing Fine-Grained Vision-Language Pretraining with Negative Augmented Samples
von: Wang, Yeyuan, et al.
Veröffentlicht: (2024)
von: Wang, Yeyuan, et al.
Veröffentlicht: (2024)
O-MARC: Omni Memory-Augmented Compression Distillation for Efficient Video Understanding
von: Wu, Peiran, et al.
Veröffentlicht: (2026)
von: Wu, Peiran, et al.
Veröffentlicht: (2026)
Cross-Modal Adapter for Vision-Language Retrieval
von: Jiang, Haojun, et al.
Veröffentlicht: (2022)
von: Jiang, Haojun, et al.
Veröffentlicht: (2022)
DIMoE-Adapters: Dynamic Expert Evolution for Continual Learning in Vision-Language Models
von: Qin, Mengxin, et al.
Veröffentlicht: (2026)
von: Qin, Mengxin, et al.
Veröffentlicht: (2026)
Skip Tuning: Pre-trained Vision-Language Models are Effective and Efficient Adapters Themselves
von: Wu, Shihan, et al.
Veröffentlicht: (2024)
von: Wu, Shihan, et al.
Veröffentlicht: (2024)
BoostAdapter: Improving Vision-Language Test-Time Adaptation via Regional Bootstrapping
von: Zhang, Taolin, et al.
Veröffentlicht: (2024)
von: Zhang, Taolin, et al.
Veröffentlicht: (2024)
Memory-Augmented Query Intent Understanding for Efficient Chat-based Image Retrieval
von: Chen, Xianke, et al.
Veröffentlicht: (2026)
von: Chen, Xianke, et al.
Veröffentlicht: (2026)
MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
von: Shi, Hao, et al.
Veröffentlicht: (2025)
von: Shi, Hao, et al.
Veröffentlicht: (2025)
QueryAdapter: Rapid Adaptation of Vision-Language Models in Response to Natural Language Queries
von: Chapman, Nicolas Harvey, et al.
Veröffentlicht: (2025)
von: Chapman, Nicolas Harvey, et al.
Veröffentlicht: (2025)
Effectively Enhancing Vision Language Large Models by Prompt Augmentation and Caption Utilization
von: Zhao, Minyi, et al.
Veröffentlicht: (2024)
von: Zhao, Minyi, et al.
Veröffentlicht: (2024)
Memory-based Adapters for Online 3D Scene Perception
von: Xu, Xiuwei, et al.
Veröffentlicht: (2024)
von: Xu, Xiuwei, et al.
Veröffentlicht: (2024)
ResAdapter: Domain Consistent Resolution Adapter for Diffusion Models
von: Cheng, Jiaxiang, et al.
Veröffentlicht: (2024)
von: Cheng, Jiaxiang, et al.
Veröffentlicht: (2024)
MAP-VLA: Memory-Augmented Prompting for Vision-Language-Action Model in Robotic Manipulation
von: Li, Runhao, et al.
Veröffentlicht: (2025)
von: Li, Runhao, et al.
Veröffentlicht: (2025)
ESceme: Vision-and-Language Navigation with Episodic Scene Memory
von: Zheng, Qi, et al.
Veröffentlicht: (2023)
von: Zheng, Qi, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding
von: Han, Yudong, et al.
Veröffentlicht: (2024) -
Test-time Alignment-Enhanced Adapter for Vision-Language Models
von: Tong, Baoshun, et al.
Veröffentlicht: (2024) -
Visual Enhanced Depth Scaling for Multimodal Latent Reasoning
von: Han, Yudong, et al.
Veröffentlicht: (2026) -
Meta-Adapter: An Online Few-shot Learner for Vision-Language Model
von: Cheng, Cheng, et al.
Veröffentlicht: (2023) -
CLIP-Adapter: Better Vision-Language Models with Feature Adapters
von: Gao, Peng, et al.
Veröffentlicht: (2021)