AIM: Let Any Multi-modal Large Language Models Embrace Efficient In-Context Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Gao, Jun, Qiao, Qian, Cao, Ziqiang, Wang, Zili, Li, Wenjie |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Not All Attention is Needed: Parameter and Computation Efficient Transfer Learning for Multi-modal Large Language Models
por: Wu, Qiong, et al.
Publicado: (2024)
por: Wu, Qiong, et al.
Publicado: (2024)
A Benchmark and Robustness Study of In-Context-Learning with Large Language Models in Music Entity Detection
por: Hachmeier, Simon, et al.
Publicado: (2024)
por: Hachmeier, Simon, et al.
Publicado: (2024)
Collaborative Evolution: Multi-Round Learning Between Large and Small Language Models for Emergent Fake News Detection
por: Zhou, Ziyi, et al.
Publicado: (2025)
por: Zhou, Ziyi, et al.
Publicado: (2025)
Evaluating Multimodal Large Language Models on Spoken Sarcasm Understanding
por: Li, Zhu, et al.
Publicado: (2025)
por: Li, Zhu, et al.
Publicado: (2025)
HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language Models
por: Wang, Xiao, et al.
Publicado: (2025)
por: Wang, Xiao, et al.
Publicado: (2025)
MMPKUBase: A Comprehensive and High-quality Chinese Multi-modal Knowledge Graph
por: Yi, Xuan, et al.
Publicado: (2024)
por: Yi, Xuan, et al.
Publicado: (2024)
IBMEA: Exploring Variational Information Bottleneck for Multi-modal Entity Alignment
por: Su, Taoyu, et al.
Publicado: (2024)
por: Su, Taoyu, et al.
Publicado: (2024)
Mixture-of-Prompt-Experts for Multi-modal Semantic Understanding
por: Wu, Zichen, et al.
Publicado: (2024)
por: Wu, Zichen, et al.
Publicado: (2024)
Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models
por: Ye, Weihao, et al.
Publicado: (2024)
por: Ye, Weihao, et al.
Publicado: (2024)
mPLUG-PaperOwl: Scientific Diagram Analysis with the Multimodal Large Language Model
por: Hu, Anwen, et al.
Publicado: (2023)
por: Hu, Anwen, et al.
Publicado: (2023)
Dynamic Self-adaptive Multiscale Distillation from Pre-trained Multimodal Large Model for Efficient Cross-modal Representation Learning
por: Liang, Zhengyang, et al.
Publicado: (2024)
por: Liang, Zhengyang, et al.
Publicado: (2024)
LoginMEA: Local-to-Global Interaction Network for Multi-modal Entity Alignment
por: Su, Taoyu, et al.
Publicado: (2024)
por: Su, Taoyu, et al.
Publicado: (2024)
OmnixR: Evaluating Omni-modality Language Models on Reasoning across Modalities
por: Chen, Lichang, et al.
Publicado: (2024)
por: Chen, Lichang, et al.
Publicado: (2024)
KMMMU: Evaluation of Massive Multi-discipline Multimodal Understanding in Korean Language and Context
por: Lee, Nahyun, et al.
Publicado: (2026)
por: Lee, Nahyun, et al.
Publicado: (2026)
Distilling Implicit Multimodal Knowledge into Large Language Models for Zero-Resource Dialogue Generation
por: Zhang, Bo, et al.
Publicado: (2024)
por: Zhang, Bo, et al.
Publicado: (2024)
MolCA: Molecular Graph-Language Modeling with Cross-Modal Projector and Uni-Modal Adapter
por: Liu, Zhiyuan, et al.
Publicado: (2023)
por: Liu, Zhiyuan, et al.
Publicado: (2023)
ChartAdapter: Large Vision-Language Model for Chart Summarization
por: Xu, Peixin, et al.
Publicado: (2024)
por: Xu, Peixin, et al.
Publicado: (2024)
Multi-MLLM Knowledge Distillation for Out-of-Context News Detection
por: Gu, Yimeng, et al.
Publicado: (2025)
por: Gu, Yimeng, et al.
Publicado: (2025)
Learning Domain-Invariant Features for Out-of-Context News Detection
por: Gu, Yimeng, et al.
Publicado: (2024)
por: Gu, Yimeng, et al.
Publicado: (2024)
SoMeLVLM: A Large Vision Language Model for Social Media Processing
por: Zhang, Xinnong, et al.
Publicado: (2024)
por: Zhang, Xinnong, et al.
Publicado: (2024)
MIND Your Reasoning: A Meta-Cognitive Intuitive-Reflective Network for Dual-Reasoning in Multimodal Stance Detection
por: Wang, Bingbing, et al.
Publicado: (2025)
por: Wang, Bingbing, et al.
Publicado: (2025)
Mutual Information-based Representations Disentanglement for Unaligned Multimodal Language Sequences
por: Qian, Fan, et al.
Publicado: (2024)
por: Qian, Fan, et al.
Publicado: (2024)
MAC-SLU: Multi-Intent Automotive Cabin Spoken Language Understanding Benchmark
por: Peng, Yuezhang, et al.
Publicado: (2025)
por: Peng, Yuezhang, et al.
Publicado: (2025)
SLAM-LLM: A Modular, Open-Source Multimodal Large Language Model Framework and Best Practice for Speech, Language, Audio and Music Processing
por: Ma, Ziyang, et al.
Publicado: (2026)
por: Ma, Ziyang, et al.
Publicado: (2026)
RealBench: A Chinese Multi-image Understanding Benchmark Close to Real-world Scenarios
por: Zhao, Fei, et al.
Publicado: (2025)
por: Zhao, Fei, et al.
Publicado: (2025)
CMMU: A Benchmark for Chinese Multi-modal Multi-type Question Understanding and Reasoning
por: He, Zheqi, et al.
Publicado: (2024)
por: He, Zheqi, et al.
Publicado: (2024)
Mitigating Modality Bias in Multi-modal Entity Alignment from a Causal Perspective
por: Su, Taoyu, et al.
Publicado: (2025)
por: Su, Taoyu, et al.
Publicado: (2025)
Traj-MLLM: Can Multimodal Large Language Models Reform Trajectory Data Mining?
por: Liu, Shuo, et al.
Publicado: (2025)
por: Liu, Shuo, et al.
Publicado: (2025)
Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark
por: Zhang, Hanlei, et al.
Publicado: (2025)
por: Zhang, Hanlei, et al.
Publicado: (2025)
NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow Matching
por: Luo, Run, et al.
Publicado: (2025)
por: Luo, Run, et al.
Publicado: (2025)
MIntRec2.0: A Large-scale Benchmark Dataset for Multimodal Intent Recognition and Out-of-scope Detection in Conversations
por: Zhang, Hanlei, et al.
Publicado: (2024)
por: Zhang, Hanlei, et al.
Publicado: (2024)
Hallu-PI: Evaluating Hallucination in Multi-modal Large Language Models within Perturbed Inputs
por: Ding, Peng, et al.
Publicado: (2024)
por: Ding, Peng, et al.
Publicado: (2024)
MIDI-LLM: Adapting Large Language Models for Text-to-MIDI Music Generation
por: Wu, Shih-Lun, et al.
Publicado: (2025)
por: Wu, Shih-Lun, et al.
Publicado: (2025)
Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs
por: Song, Dingjie, et al.
Publicado: (2024)
por: Song, Dingjie, et al.
Publicado: (2024)
SelfCP: Compressing Over-Limit Prompt via the Frozen Large Language Model Itself
por: Gao, Jun, et al.
Publicado: (2024)
por: Gao, Jun, et al.
Publicado: (2024)
Knowledge Acquisition Disentanglement for Knowledge-based Visual Question Answering with Large Language Models
por: An, Wenbin, et al.
Publicado: (2024)
por: An, Wenbin, et al.
Publicado: (2024)
MMAPS: End-to-End Multi-Grained Multi-Modal Attribute-Aware Product Summarization
por: Chen, Tao, et al.
Publicado: (2023)
por: Chen, Tao, et al.
Publicado: (2023)
Towards Unified Multi-Modal Personalization: Large Vision-Language Models for Generative Recommendation and Beyond
por: Wei, Tianxin, et al.
Publicado: (2024)
por: Wei, Tianxin, et al.
Publicado: (2024)
Listening to the Unspoken: Exploring "365" Aspects of Multimodal Interview Performance Assessment
por: Li, Jia, et al.
Publicado: (2025)
por: Li, Jia, et al.
Publicado: (2025)
Verifying Cross-modal Entity Consistency in News using Vision-language Models
por: Tahmasebi, Sahar, et al.
Publicado: (2025)
por: Tahmasebi, Sahar, et al.
Publicado: (2025)
Ejemplares similares
-
Not All Attention is Needed: Parameter and Computation Efficient Transfer Learning for Multi-modal Large Language Models
por: Wu, Qiong, et al.
Publicado: (2024) -
A Benchmark and Robustness Study of In-Context-Learning with Large Language Models in Music Entity Detection
por: Hachmeier, Simon, et al.
Publicado: (2024) -
Collaborative Evolution: Multi-Round Learning Between Large and Small Language Models for Emergent Fake News Detection
por: Zhou, Ziyi, et al.
Publicado: (2025) -
Evaluating Multimodal Large Language Models on Spoken Sarcasm Understanding
por: Li, Zhu, et al.
Publicado: (2025) -
HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language Models
por: Wang, Xiao, et al.
Publicado: (2025)