How Does the Textual Information Affect the Retrieval of Multimodal In-Context Learning?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Luo, Yang, Zheng, Zangwei, Zhu, Zirui, You, Yang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
What Factors Affect Multi-Modal In-Context Learning? An In-Depth Exploration
von: Qin, Libo, et al.
Veröffentlicht: (2024)
von: Qin, Libo, et al.
Veröffentlicht: (2024)
Retrieving Counterfactuals Improves Visual In-Context Learning
von: Xiong, Guangzhi, et al.
Veröffentlicht: (2026)
von: Xiong, Guangzhi, et al.
Veröffentlicht: (2026)
Cross-Modal Projection in Multimodal LLMs Doesn't Really Project Visual Attributes to Textual Space
von: Verma, Gaurav, et al.
Veröffentlicht: (2024)
von: Verma, Gaurav, et al.
Veröffentlicht: (2024)
Decoding by Perturbation: Mitigating MLLM Hallucinations via Dynamic Textual Perturbation
von: Jia, Sihang, et al.
Veröffentlicht: (2026)
von: Jia, Sihang, et al.
Veröffentlicht: (2026)
Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
von: Yue, Yang, et al.
Veröffentlicht: (2025)
von: Yue, Yang, et al.
Veröffentlicht: (2025)
Mitigating Behavioral Hallucination in Multimodal Large Language Models for Sequential Images
von: You, Liangliang, et al.
Veröffentlicht: (2025)
von: You, Liangliang, et al.
Veröffentlicht: (2025)
Why LVLMs Are More Prone to Hallucinations in Longer Responses: The Role of Context
von: Zheng, Ge, et al.
Veröffentlicht: (2025)
von: Zheng, Ge, et al.
Veröffentlicht: (2025)
Nearest Neighbor Normalization Improves Multimodal Retrieval
von: Chowdhury, Neil, et al.
Veröffentlicht: (2024)
von: Chowdhury, Neil, et al.
Veröffentlicht: (2024)
Advancing Multimodal In-Context Learning in Large Vision-Language Models with Task-aware Demonstrations
von: Li, Yanshu
Veröffentlicht: (2025)
von: Li, Yanshu
Veröffentlicht: (2025)
List Items One by One: A New Data Source and Learning Paradigm for Multimodal LLMs
von: Yan, An, et al.
Veröffentlicht: (2024)
von: Yan, An, et al.
Veröffentlicht: (2024)
LFTR: Learning-Free Token Reduction for Multimodal Large Language Models
von: Zhao, Zihui, et al.
Veröffentlicht: (2025)
von: Zhao, Zihui, et al.
Veröffentlicht: (2025)
See It from My Perspective: How Language Affects Cultural Bias in Image Understanding
von: Ananthram, Amith, et al.
Veröffentlicht: (2024)
von: Ananthram, Amith, et al.
Veröffentlicht: (2024)
Model Composition for Multimodal Large Language Models
von: Chen, Chi, et al.
Veröffentlicht: (2024)
von: Chen, Chi, et al.
Veröffentlicht: (2024)
LAViTeR: Learning Aligned Visual and Textual Representations Assisted by Image and Caption Generation
von: Hashemi, Mohammad Abuzar, et al.
Veröffentlicht: (2021)
von: Hashemi, Mohammad Abuzar, et al.
Veröffentlicht: (2021)
MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation
von: Li, Yan, et al.
Veröffentlicht: (2026)
von: Li, Yan, et al.
Veröffentlicht: (2026)
Multimodal Task Vectors Enable Many-Shot Multimodal In-Context Learning
von: Huang, Brandon, et al.
Veröffentlicht: (2024)
von: Huang, Brandon, et al.
Veröffentlicht: (2024)
MRAG-Bench: Vision-Centric Evaluation for Retrieval-Augmented Multimodal Models
von: Hu, Wenbo, et al.
Veröffentlicht: (2024)
von: Hu, Wenbo, et al.
Veröffentlicht: (2024)
Multimodal Retrieval-Augmented Generation with Large Language Models for Medical VQA
von: Karim, A H M Rezaul, et al.
Veröffentlicht: (2025)
von: Karim, A H M Rezaul, et al.
Veröffentlicht: (2025)
How to Train Your Long-Context Visual Document Model
von: Veselka, Austin
Veröffentlicht: (2026)
von: Veselka, Austin
Veröffentlicht: (2026)
Image-Text Out-Of-Context Detection Using Synthetic Multimodal Misinformation
von: Shalabi, Fatma, et al.
Veröffentlicht: (2024)
von: Shalabi, Fatma, et al.
Veröffentlicht: (2024)
MERIT: Maximum-normalized Element-wise Ratio for Language Model Large-batch Training
von: Luo, Yang, et al.
Veröffentlicht: (2025)
von: Luo, Yang, et al.
Veröffentlicht: (2025)
Progressive Multimodal Reasoning via Active Retrieval
von: Dong, Guanting, et al.
Veröffentlicht: (2024)
von: Dong, Guanting, et al.
Veröffentlicht: (2024)
Explaining How Visual, Textual and Multimodal Encoders Share Concepts
von: Cornet, Clément, et al.
Veröffentlicht: (2025)
von: Cornet, Clément, et al.
Veröffentlicht: (2025)
MERRIN: A Benchmark for Multimodal Evidence Retrieval and Reasoning in Noisy Web Environments
von: Wang, Han, et al.
Veröffentlicht: (2026)
von: Wang, Han, et al.
Veröffentlicht: (2026)
Efficient Personalized Text-to-image Generation by Leveraging Textual Subspace
von: Du, Shian, et al.
Veröffentlicht: (2024)
von: Du, Shian, et al.
Veröffentlicht: (2024)
Think Visually, Reason Textually: Vision-Language Synergy in ARC
von: Zhang, Beichen, et al.
Veröffentlicht: (2025)
von: Zhang, Beichen, et al.
Veröffentlicht: (2025)
MULTI: Multimodal Understanding Leaderboard with Text and Images
von: Zhu, Zichen, et al.
Veröffentlicht: (2024)
von: Zhu, Zichen, et al.
Veröffentlicht: (2024)
E$^2$AT: Multimodal Jailbreak Defense via Dynamic Joint Optimization for Multimodal Large Language Models
von: Lu, Liming, et al.
Veröffentlicht: (2025)
von: Lu, Liming, et al.
Veröffentlicht: (2025)
MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings
von: Chen, Haonan, et al.
Veröffentlicht: (2025)
von: Chen, Haonan, et al.
Veröffentlicht: (2025)
Many-Shot In-Context Learning in Multimodal Foundation Models
von: Jiang, Yixing, et al.
Veröffentlicht: (2024)
von: Jiang, Yixing, et al.
Veröffentlicht: (2024)
Recurrence Meets Transformers for Universal Multimodal Retrieval
von: Caffagni, Davide, et al.
Veröffentlicht: (2025)
von: Caffagni, Davide, et al.
Veröffentlicht: (2025)
UniCode: Learning a Unified Codebook for Multimodal Large Language Models
von: Zheng, Sipeng, et al.
Veröffentlicht: (2024)
von: Zheng, Sipeng, et al.
Veröffentlicht: (2024)
LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
von: Zhang, Shaolei, et al.
Veröffentlicht: (2025)
von: Zhang, Shaolei, et al.
Veröffentlicht: (2025)
MLLM-CL: Continual Learning for Multimodal Large Language Models
von: Zhao, Hongbo, et al.
Veröffentlicht: (2025)
von: Zhao, Hongbo, et al.
Veröffentlicht: (2025)
Generative Universal Verifier as Multimodal Meta-Reasoner
von: Zhang, Xinchen, et al.
Veröffentlicht: (2025)
von: Zhang, Xinchen, et al.
Veröffentlicht: (2025)
MLLMReID: Multimodal Large Language Model-based Person Re-identification
von: Yang, Shan, et al.
Veröffentlicht: (2024)
von: Yang, Shan, et al.
Veröffentlicht: (2024)
Robust Adaptation of Large Multimodal Models for Retrieval Augmented Hateful Meme Detection
von: Mei, Jingbiao, et al.
Veröffentlicht: (2025)
von: Mei, Jingbiao, et al.
Veröffentlicht: (2025)
Figuring out Figures: Using Textual References to Caption Scientific Figures
von: Cao, Stanley, et al.
Veröffentlicht: (2024)
von: Cao, Stanley, et al.
Veröffentlicht: (2024)
Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation
von: Guo, Ziyu, et al.
Veröffentlicht: (2025)
von: Guo, Ziyu, et al.
Veröffentlicht: (2025)
VisionGraph: Leveraging Large Multimodal Models for Graph Theory Problems in Visual Context
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
What Factors Affect Multi-Modal In-Context Learning? An In-Depth Exploration
von: Qin, Libo, et al.
Veröffentlicht: (2024) -
Retrieving Counterfactuals Improves Visual In-Context Learning
von: Xiong, Guangzhi, et al.
Veröffentlicht: (2026) -
Cross-Modal Projection in Multimodal LLMs Doesn't Really Project Visual Attributes to Textual Space
von: Verma, Gaurav, et al.
Veröffentlicht: (2024) -
Decoding by Perturbation: Mitigating MLLM Hallucinations via Dynamic Textual Perturbation
von: Jia, Sihang, et al.
Veröffentlicht: (2026) -
Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
von: Yue, Yang, et al.
Veröffentlicht: (2025)