Multimodal Task Vectors Enable Many-Shot Multimodal In-Context Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Brandon, Mitra, Chancharik, Arbelle, Assaf, Karlinsky, Leonid, Darrell, Trevor, Herzig, Roei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features
von: Mitra, Chancharik, et al.
Veröffentlicht: (2024)
von: Mitra, Chancharik, et al.
Veröffentlicht: (2024)
Compositional Chain-of-Thought Prompting for Large Multimodal Models
von: Mitra, Chancharik, et al.
Veröffentlicht: (2023)
von: Mitra, Chancharik, et al.
Veröffentlicht: (2023)
Activation Reward Models for Few-Shot Model Alignment
von: Chai, Tianning, et al.
Veröffentlicht: (2025)
von: Chai, Tianning, et al.
Veröffentlicht: (2025)
Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
von: Mitra, Chancharik, et al.
Veröffentlicht: (2025)
von: Mitra, Chancharik, et al.
Veröffentlicht: (2025)
Latent Implicit Visual Reasoning
von: Li, Kelvin, et al.
Veröffentlicht: (2025)
von: Li, Kelvin, et al.
Veröffentlicht: (2025)
Towards Multimodal In-Context Learning for Vision & Language Models
von: Doveh, Sivan, et al.
Veröffentlicht: (2024)
von: Doveh, Sivan, et al.
Veröffentlicht: (2024)
Recursive Visual Programming
von: Ge, Jiaxin, et al.
Veröffentlicht: (2023)
von: Ge, Jiaxin, et al.
Veröffentlicht: (2023)
Which One? Leveraging Context Between Objects and Multiple Views for Language Grounding
von: Mitra, Chancharik, et al.
Veröffentlicht: (2023)
von: Mitra, Chancharik, et al.
Veröffentlicht: (2023)
TraveLER: A Modular Multi-LMM Agent Framework for Video Question-Answering
von: Shang, Chuyi, et al.
Veröffentlicht: (2024)
von: Shang, Chuyi, et al.
Veröffentlicht: (2024)
DAVE: A VLM Vision Encoder for Document Understanding and Web Agents
von: Huang, Brandon, et al.
Veröffentlicht: (2025)
von: Huang, Brandon, et al.
Veröffentlicht: (2025)
ConMe: Rethinking Evaluation of Compositional Reasoning for Modern VLMs
von: Huang, Irene, et al.
Veröffentlicht: (2024)
von: Huang, Irene, et al.
Veröffentlicht: (2024)
Many-Shot In-Context Learning in Multimodal Foundation Models
von: Jiang, Yixing, et al.
Veröffentlicht: (2024)
von: Jiang, Yixing, et al.
Veröffentlicht: (2024)
MAEDAY: MAE for few and zero shot AnomalY-Detection
von: Schwartz, Eli, et al.
Veröffentlicht: (2022)
von: Schwartz, Eli, et al.
Veröffentlicht: (2022)
Vision-Language Models Create Cross-Modal Task Representations
von: Luo, Grace, et al.
Veröffentlicht: (2024)
von: Luo, Grace, et al.
Veröffentlicht: (2024)
TULIP: Towards Unified Language-Image Pretraining
von: Tang, Zineng, et al.
Veröffentlicht: (2025)
von: Tang, Zineng, et al.
Veröffentlicht: (2025)
Finding Visual Task Vectors
von: Hojel, Alberto, et al.
Veröffentlicht: (2024)
von: Hojel, Alberto, et al.
Veröffentlicht: (2024)
In-Context Learning Enables Robot Action Prediction in LLMs
von: Yin, Yida, et al.
Veröffentlicht: (2024)
von: Yin, Yida, et al.
Veröffentlicht: (2024)
Evaluating Linguistic Capabilities of Multimodal LLMs in the Lens of Few-Shot Learning
von: Dogan, Mustafa, et al.
Veröffentlicht: (2024)
von: Dogan, Mustafa, et al.
Veröffentlicht: (2024)
TACO: Enhancing Multimodal In-context Learning via Task Mapping-Guided Sequence Configuration
von: Li, Yanshu, et al.
Veröffentlicht: (2025)
von: Li, Yanshu, et al.
Veröffentlicht: (2025)
Make LVLMs Focus: Context-Aware Attention Modulation for Better Multimodal In-Context Learning
von: Li, Yanshu, et al.
Veröffentlicht: (2025)
von: Li, Yanshu, et al.
Veröffentlicht: (2025)
Large Multilingual Models Pivot Zero-Shot Multimodal Learning across Languages
von: Hu, Jinyi, et al.
Veröffentlicht: (2023)
von: Hu, Jinyi, et al.
Veröffentlicht: (2023)
Advancing Multimodal In-Context Learning in Large Vision-Language Models with Task-aware Demonstrations
von: Li, Yanshu
Veröffentlicht: (2025)
von: Li, Yanshu
Veröffentlicht: (2025)
LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning
von: Niu, Dantong, et al.
Veröffentlicht: (2024)
von: Niu, Dantong, et al.
Veröffentlicht: (2024)
Kosmos-G: Generating Images in Context with Multimodal Large Language Models
von: Pan, Xichen, et al.
Veröffentlicht: (2023)
von: Pan, Xichen, et al.
Veröffentlicht: (2023)
ZALM3: Zero-Shot Enhancement of Vision-Language Alignment via In-Context Information in Multi-Turn Multimodal Medical Dialogue
von: Li, Zhangpu, et al.
Veröffentlicht: (2024)
von: Li, Zhangpu, et al.
Veröffentlicht: (2024)
From Generated Human Videos to Physically Plausible Robot Trajectories
von: Ni, James, et al.
Veröffentlicht: (2025)
von: Ni, James, et al.
Veröffentlicht: (2025)
Adaptive Memory Replay for Continual Learning
von: Smith, James Seale, et al.
Veröffentlicht: (2024)
von: Smith, James Seale, et al.
Veröffentlicht: (2024)
LiveXiv -- A Multi-Modal Live Benchmark Based on Arxiv Papers Content
von: Shabtay, Nimrod, et al.
Veröffentlicht: (2024)
von: Shabtay, Nimrod, et al.
Veröffentlicht: (2024)
Zero-Shot Scene Understanding with Multimodal Large Language Models for Automated Vehicles
von: Elhenawy, Mohammed, et al.
Veröffentlicht: (2025)
von: Elhenawy, Mohammed, et al.
Veröffentlicht: (2025)
Towards Zero-Shot Anomaly Detection and Reasoning with Multimodal Large Language Models
von: Xu, Jiacong, et al.
Veröffentlicht: (2025)
von: Xu, Jiacong, et al.
Veröffentlicht: (2025)
Context-Aware Multimodal Pretraining
von: Roth, Karsten, et al.
Veröffentlicht: (2024)
von: Roth, Karsten, et al.
Veröffentlicht: (2024)
LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference
von: Wan, Zhongwei, et al.
Veröffentlicht: (2024)
von: Wan, Zhongwei, et al.
Veröffentlicht: (2024)
Shared and Private Information Learning in Multimodal Sentiment Analysis with Deep Modal Alignment and Self-supervised Multi-Task Learning
von: Lai, Songning, et al.
Veröffentlicht: (2023)
von: Lai, Songning, et al.
Veröffentlicht: (2023)
LLaVA-Critic: Learning to Evaluate Multimodal Models
von: Xiong, Tianyi, et al.
Veröffentlicht: (2024)
von: Xiong, Tianyi, et al.
Veröffentlicht: (2024)
Teaching VLMs to Localize Specific Objects from In-context Examples
von: Doveh, Sivan, et al.
Veröffentlicht: (2024)
von: Doveh, Sivan, et al.
Veröffentlicht: (2024)
Common Objects Out of Context (COOCo): Investigating Multimodal Context and Semantic Scene Violations in Referential Communication
von: Merlo, Filippo, et al.
Veröffentlicht: (2025)
von: Merlo, Filippo, et al.
Veröffentlicht: (2025)
The Power of Many: Multi-Agent Multimodal Models for Cultural Image Captioning
von: Bai, Longju, et al.
Veröffentlicht: (2024)
von: Bai, Longju, et al.
Veröffentlicht: (2024)
Bridging the Visual Gap: Fine-Tuning Multimodal Models with Knowledge-Adapted Captions
von: Yanuka, Moran, et al.
Veröffentlicht: (2024)
von: Yanuka, Moran, et al.
Veröffentlicht: (2024)
Learning from Fine-Grained Visual Discrepancies: Mitigating Multimodal Hallucinations via In-Context Visual Contrastive Optimization
von: Deng, Haolin, et al.
Veröffentlicht: (2026)
von: Deng, Haolin, et al.
Veröffentlicht: (2026)
Culture-Aware Humorous Captioning: Multimodal Humor Generation across Cultural Contexts
von: Xu, Run, et al.
Veröffentlicht: (2026)
von: Xu, Run, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features
von: Mitra, Chancharik, et al.
Veröffentlicht: (2024) -
Compositional Chain-of-Thought Prompting for Large Multimodal Models
von: Mitra, Chancharik, et al.
Veröffentlicht: (2023) -
Activation Reward Models for Few-Shot Model Alignment
von: Chai, Tianning, et al.
Veröffentlicht: (2025) -
Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
von: Mitra, Chancharik, et al.
Veröffentlicht: (2025) -
Latent Implicit Visual Reasoning
von: Li, Kelvin, et al.
Veröffentlicht: (2025)