What Factors Affect Multi-Modal In-Context Learning? An In-Depth Exploration
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Qin, Libo, Chen, Qiguang, Fei, Hao, Chen, Zhi, Li, Min, Che, Wanxiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought
von: Chen, Qiguang, et al.
Veröffentlicht: (2024)
von: Chen, Qiguang, et al.
Veröffentlicht: (2024)
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models
von: Zhang, Yongheng, et al.
Veröffentlicht: (2025)
von: Zhang, Yongheng, et al.
Veröffentlicht: (2025)
RBF++: Quantifying and Optimizing Reasoning Boundaries across Measurable and Unmeasurable Capabilities for Chain-of-Thought Reasoning
von: Chen, Qiguang, et al.
Veröffentlicht: (2025)
von: Chen, Qiguang, et al.
Veröffentlicht: (2025)
CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models
von: Cheng, Zihui, et al.
Veröffentlicht: (2024)
von: Cheng, Zihui, et al.
Veröffentlicht: (2024)
OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model
von: Chen, Qiguang, et al.
Veröffentlicht: (2026)
von: Chen, Qiguang, et al.
Veröffentlicht: (2026)
MPCC: A Novel Benchmark for Multimodal Planning with Complex Constraints in Multimodal Large Language Models
von: Ji, Yiyan, et al.
Veröffentlicht: (2025)
von: Ji, Yiyan, et al.
Veröffentlicht: (2025)
Visual Thoughts: A Unified Perspective of Understanding Multimodal Chain-of-Thought
von: Cheng, Zihui, et al.
Veröffentlicht: (2025)
von: Cheng, Zihui, et al.
Veröffentlicht: (2025)
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs
von: Xu, Xiao, et al.
Veröffentlicht: (2025)
von: Xu, Xiao, et al.
Veröffentlicht: (2025)
Exploring Multi-Grained Concept Annotations for Multimodal Large Language Models
von: Xu, Xiao, et al.
Veröffentlicht: (2024)
von: Xu, Xiao, et al.
Veröffentlicht: (2024)
What are the Essential Factors in Crafting Effective Long Context Multi-Hop Instruction Datasets? Insights and Best Practices
von: Chen, Zhi, et al.
Veröffentlicht: (2024)
von: Chen, Zhi, et al.
Veröffentlicht: (2024)
MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
von: Zhao, Haozhe, et al.
Veröffentlicht: (2023)
von: Zhao, Haozhe, et al.
Veröffentlicht: (2023)
Let's Think with Images Efficiently! An Interleaved-Modal Chain-of-Thought Reasoning Framework with Dynamic and Precise Visual Thoughts
von: Liu, Xu, et al.
Veröffentlicht: (2026)
von: Liu, Xu, et al.
Veröffentlicht: (2026)
MMLongCite: A Benchmark for Evaluating Fidelity of Long-Context Vision-Language Models
von: Zhou, Keyan, et al.
Veröffentlicht: (2025)
von: Zhou, Keyan, et al.
Veröffentlicht: (2025)
ChartREG++: Towards Benchmarking and Improving Chart Referring Expression Grounding under Diverse referring clues and Multi-Target Referring
von: Niu, Tianhao, et al.
Veröffentlicht: (2026)
von: Niu, Tianhao, et al.
Veröffentlicht: (2026)
How Does the Textual Information Affect the Retrieval of Multimodal In-Context Learning?
von: Luo, Yang, et al.
Veröffentlicht: (2024)
von: Luo, Yang, et al.
Veröffentlicht: (2024)
CVLUE: A New Benchmark Dataset for Chinese Vision-Language Understanding Evaluation
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
Revisiting Multi-Modal LLM Evaluation
von: Lu, Jian, et al.
Veröffentlicht: (2024)
von: Lu, Jian, et al.
Veröffentlicht: (2024)
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models
von: Huang, Chengyue, et al.
Veröffentlicht: (2025)
von: Huang, Chengyue, et al.
Veröffentlicht: (2025)
Multi-Prompt with Depth Partitioned Cross-Modal Learning
von: Tian, Yingjie, et al.
Veröffentlicht: (2023)
von: Tian, Yingjie, et al.
Veröffentlicht: (2023)
Otter: A Multi-Modal Model with In-Context Instruction Tuning
von: Li, Bo, et al.
Veröffentlicht: (2023)
von: Li, Bo, et al.
Veröffentlicht: (2023)
APLe: Token-Wise Adaptive for Multi-Modal Prompt Learning
von: Cao, Guiming, et al.
Veröffentlicht: (2024)
von: Cao, Guiming, et al.
Veröffentlicht: (2024)
Generative Sign-description Prompts with Multi-positive Contrastive Learning for Sign Language Recognition
von: Liang, Siyu, et al.
Veröffentlicht: (2025)
von: Liang, Siyu, et al.
Veröffentlicht: (2025)
Multi-Modal Language Models as Text-to-Image Model Evaluators
von: Chen, Jiahui, et al.
Veröffentlicht: (2025)
von: Chen, Jiahui, et al.
Veröffentlicht: (2025)
SAM4MLLM: Enhance Multi-Modal Large Language Model for Referring Expression Segmentation
von: Chen, Yi-Chia, et al.
Veröffentlicht: (2024)
von: Chen, Yi-Chia, et al.
Veröffentlicht: (2024)
AutoCAP: Towards Automatic Cross-lingual Alignment Planning for Zero-shot Chain-of-Thought
von: Zhang, Yongheng, et al.
Veröffentlicht: (2024)
von: Zhang, Yongheng, et al.
Veröffentlicht: (2024)
Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models
von: Zhu, Tinghui, et al.
Veröffentlicht: (2024)
von: Zhu, Tinghui, et al.
Veröffentlicht: (2024)
In-Context Meta LoRA Generation
von: Shao, Yihua, et al.
Veröffentlicht: (2025)
von: Shao, Yihua, et al.
Veröffentlicht: (2025)
MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI
von: Yao, Huanjin, et al.
Veröffentlicht: (2025)
von: Yao, Huanjin, et al.
Veröffentlicht: (2025)
Learning What Matters: Dynamic Dimension Selection and Aggregation for Interpretable Vision-Language Reward Modeling
von: Chen, Qiyuan, et al.
Veröffentlicht: (2026)
von: Chen, Qiyuan, et al.
Veröffentlicht: (2026)
From Introspection to Best Practices: Principled Analysis of Demonstrations in Multimodal In-Context Learning
von: Xu, Nan, et al.
Veröffentlicht: (2024)
von: Xu, Nan, et al.
Veröffentlicht: (2024)
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models
von: Nguyen, Pha, et al.
Veröffentlicht: (2025)
von: Nguyen, Pha, et al.
Veröffentlicht: (2025)
DeepSight: Bridging Depth Maps and Language with a Depth-Driven Multimodal Model
von: Yang, Hao, et al.
Veröffentlicht: (2026)
von: Yang, Hao, et al.
Veröffentlicht: (2026)
From Compound Figures to Composite Understanding: Developing a Multi-Modal LLM from Biomedical Literature with Medical Multiple-Image Benchmarking and Validation
von: Chen, Zhen, et al.
Veröffentlicht: (2025)
von: Chen, Zhen, et al.
Veröffentlicht: (2025)
TeamCraft: A Benchmark for Multi-Modal Multi-Agent Systems in Minecraft
von: Long, Qian, et al.
Veröffentlicht: (2024)
von: Long, Qian, et al.
Veröffentlicht: (2024)
Core Knowledge Deficits in Multi-Modal Language Models
von: Li, Yijiang, et al.
Veröffentlicht: (2024)
von: Li, Yijiang, et al.
Veröffentlicht: (2024)
MileBench: Benchmarking MLLMs in Long Context
von: Song, Dingjie, et al.
Veröffentlicht: (2024)
von: Song, Dingjie, et al.
Veröffentlicht: (2024)
MM-R5: MultiModal Reasoning-Enhanced ReRanker via Reinforcement Learning for Document Retrieval
von: Xu, Mingjun, et al.
Veröffentlicht: (2025)
von: Xu, Mingjun, et al.
Veröffentlicht: (2025)
Calibrating Uncertainty Quantification of Multi-Modal LLMs using Grounding
von: Padhi, Trilok, et al.
Veröffentlicht: (2025)
von: Padhi, Trilok, et al.
Veröffentlicht: (2025)
Dermacen Analytica: A Novel Methodology Integrating Multi-Modal Large Language Models with Machine Learning in tele-dermatology
von: Panagoulias, Dimitrios P., et al.
Veröffentlicht: (2024)
von: Panagoulias, Dimitrios P., et al.
Veröffentlicht: (2024)
VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models
von: Li, Zejun, et al.
Veröffentlicht: (2024)
von: Li, Zejun, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought
von: Chen, Qiguang, et al.
Veröffentlicht: (2024) -
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models
von: Zhang, Yongheng, et al.
Veröffentlicht: (2025) -
RBF++: Quantifying and Optimizing Reasoning Boundaries across Measurable and Unmeasurable Capabilities for Chain-of-Thought Reasoning
von: Chen, Qiguang, et al.
Veröffentlicht: (2025) -
CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models
von: Cheng, Zihui, et al.
Veröffentlicht: (2024) -
OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model
von: Chen, Qiguang, et al.
Veröffentlicht: (2026)