ContextNav: Towards Agentic Multimodal In-Context Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Fu, Honghao, Ouyang, Yuan, Chang, Kai-Wei, Wang, Yiwei, Huang, Zi, Cai, Yujun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG
by: Fu, Honghao, et al.
Published: (2026)
by: Fu, Honghao, et al.
Published: (2026)
FrameMind: Frame-Interleaved Video Reasoning via Reinforcement Learning
by: Ge, Haonan, et al.
Published: (2025)
by: Ge, Haonan, et al.
Published: (2025)
VistaWise: Building Cost-Effective Agent with Cross-Modal Knowledge Graph for Minecraft
by: Fu, Honghao, et al.
Published: (2025)
by: Fu, Honghao, et al.
Published: (2025)
Detecting Multimodal Situations with Insufficient Context and Abstaining from Baseless Predictions
by: Liu, Junzhang, et al.
Published: (2024)
by: Liu, Junzhang, et al.
Published: (2024)
Token-Efficient Prompt Injection Attack: Provoking Cessation in LLM Reasoning via Adaptive Token Compression
by: Cui, Yu, et al.
Published: (2025)
by: Cui, Yu, et al.
Published: (2025)
Mimic In-Context Learning for Multimodal Tasks
by: Jiang, Yuchu, et al.
Published: (2025)
by: Jiang, Yuchu, et al.
Published: (2025)
An Agentic Model Context Protocol Framework for Medical Concept Standardization
by: Ahn, Jaerong, et al.
Published: (2025)
by: Ahn, Jaerong, et al.
Published: (2025)
Cure or Poison? Embedding Instructions Visually Alters Hallucination in Vision-Language Models
by: Wang, Zhaochen, et al.
Published: (2025)
by: Wang, Zhaochen, et al.
Published: (2025)
To Retrieve or To Think? An Agentic Approach for Context Evolution
by: Chen, Rubing, et al.
Published: (2026)
by: Chen, Rubing, et al.
Published: (2026)
Process or Result? Manipulated Ending Tokens Can Mislead Reasoning LLMs to Ignore the Correct Reasoning Steps
by: Cui, Yu, et al.
Published: (2025)
by: Cui, Yu, et al.
Published: (2025)
BRIGHT+: Upgrading the BRIGHT Benchmark with MARCUS, a Multi-Agent RAG Clean-Up Suite
by: Chen, Liyang, et al.
Published: (2025)
by: Chen, Liyang, et al.
Published: (2025)
Haystack Engineering: Context Engineering for Heterogeneous and Agentic Long-Context Evaluation
by: Li, Mufei, et al.
Published: (2025)
by: Li, Mufei, et al.
Published: (2025)
Toward Autonomous Computational Catalysis Research via Agentic Systems
by: Chen, Honghao, et al.
Published: (2026)
by: Chen, Honghao, et al.
Published: (2026)
ChainMPQ: Interleaved Text-Image Reasoning Chains for Mitigating Relation Hallucinations
by: Wu, Yike, et al.
Published: (2025)
by: Wu, Yike, et al.
Published: (2025)
Text Speaks Louder than Vision: ASCII Art Reveals Textual Biases in Vision-Language Models
by: Wang, Zhaochen, et al.
Published: (2025)
by: Wang, Zhaochen, et al.
Published: (2025)
Differentially Private Multimodal In-Context Learning
by: Ngong, Ivoline C., et al.
Published: (2026)
by: Ngong, Ivoline C., et al.
Published: (2026)
Context Engineering 2.0: The Context of Context Engineering
by: Hua, Qishuo, et al.
Published: (2025)
by: Hua, Qishuo, et al.
Published: (2025)
Structured Attention Matters to Multimodal LLMs in Document Understanding
by: Liu, Chang, et al.
Published: (2025)
by: Liu, Chang, et al.
Published: (2025)
AudioMotionBench: Evaluating Auditory Motion Perception in Audio LLMs
by: Sun, Zhe, et al.
Published: (2025)
by: Sun, Zhe, et al.
Published: (2025)
OpenNav: Open-World Navigation with Multimodal Large Language Models
by: Yuan, Mingfeng, et al.
Published: (2025)
by: Yuan, Mingfeng, et al.
Published: (2025)
Multimodal Contrastive In-Context Learning
by: Miyanishi, Yosuke, et al.
Published: (2024)
by: Miyanishi, Yosuke, et al.
Published: (2024)
Demonstration Selection for In-Context Learning via Reinforcement Learning
by: Wang, Xubin, et al.
Published: (2024)
by: Wang, Xubin, et al.
Published: (2024)
M$^3$-ACE: Rectifying Visual Perception in Multimodal Math Reasoning via Multi-Agentic Context Engineering
by: Xie, Peijin, et al.
Published: (2026)
by: Xie, Peijin, et al.
Published: (2026)
CLEAR: Context Augmentation from Contrastive Learning of Experience via Agentic Reflection
by: Liu, Linbo, et al.
Published: (2026)
by: Liu, Linbo, et al.
Published: (2026)
Where and What Matters: Sensitivity-Aware Task Vectors for Many-Shot Multimodal In-Context Learning
by: Ma, Ziyu, et al.
Published: (2025)
by: Ma, Ziyu, et al.
Published: (2025)
FactGuard: Agentic Video Misinformation Detection via Reinforcement Learning
by: Li, Zehao, et al.
Published: (2026)
by: Li, Zehao, et al.
Published: (2026)
M$^2$IV: Towards Efficient and Fine-grained Multimodal In-Context Learning via Representation Engineering
by: Li, Yanshu, et al.
Published: (2025)
by: Li, Yanshu, et al.
Published: (2025)
Think Carefully and Check Again! Meta-Generation Unlocking LLMs for Low-Resource Cross-Lingual Summarization
by: Li, Zhecheng, et al.
Published: (2024)
by: Li, Zhecheng, et al.
Published: (2024)
Context-DPO: Aligning Language Models for Context-Faithfulness
by: Bi, Baolong, et al.
Published: (2024)
by: Bi, Baolong, et al.
Published: (2024)
How Few-Shot Examples Add Up: A Causal Decomposition of Function Vectors in In-Context Learning
by: Wang, Entang, et al.
Published: (2026)
by: Wang, Entang, et al.
Published: (2026)
Self-Taught Agentic Long Context Understanding
by: Zhuang, Yufan, et al.
Published: (2025)
by: Zhuang, Yufan, et al.
Published: (2025)
True Multimodal In-Context Learning Needs Attention to the Visual Context
by: Chen, Shuo, et al.
Published: (2025)
by: Chen, Shuo, et al.
Published: (2025)
ConTextual: Evaluating Context-Sensitive Text-Rich Visual Reasoning in Large Multimodal Models
by: Wadhawan, Rohan, et al.
Published: (2024)
by: Wadhawan, Rohan, et al.
Published: (2024)
Towards Monotonic Improvement in In-Context Reinforcement Learning
by: Zhang, Wenhao, et al.
Published: (2025)
by: Zhang, Wenhao, et al.
Published: (2025)
Knowledgeable In-Context Tuning: Exploring and Exploiting Factual Knowledge for In-Context Learning
by: Wang, Jianing, et al.
Published: (2023)
by: Wang, Jianing, et al.
Published: (2023)
Context-level Language Modeling by Learning Predictive Context Embeddings
by: Dai, Beiya, et al.
Published: (2025)
by: Dai, Beiya, et al.
Published: (2025)
Rationale-Grounded In-Context Learning for Time Series Reasoning with Multimodal Large Language Models
by: Liu, Qingxiang, et al.
Published: (2026)
by: Liu, Qingxiang, et al.
Published: (2026)
AutoContext: Instance-Level Context Learning for LLM Agents
by: Cai, Kuntai, et al.
Published: (2025)
by: Cai, Kuntai, et al.
Published: (2025)
Agentic-MME: What Agentic Capability Really Brings to Multimodal Intelligence?
by: Wei, Qianshan, et al.
Published: (2026)
by: Wei, Qianshan, et al.
Published: (2026)
Timeline-based Sentence Decomposition with In-Context Learning for Temporal Fact Extraction
by: Chen, Jianhao, et al.
Published: (2024)
by: Chen, Jianhao, et al.
Published: (2024)
Similar Items
-
VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG
by: Fu, Honghao, et al.
Published: (2026) -
FrameMind: Frame-Interleaved Video Reasoning via Reinforcement Learning
by: Ge, Haonan, et al.
Published: (2025) -
VistaWise: Building Cost-Effective Agent with Cross-Modal Knowledge Graph for Minecraft
by: Fu, Honghao, et al.
Published: (2025) -
Detecting Multimodal Situations with Insufficient Context and Abstaining from Baseless Predictions
by: Liu, Junzhang, et al.
Published: (2024) -
Token-Efficient Prompt Injection Attack: Provoking Cessation in LLM Reasoning via Adaptive Token Compression
by: Cui, Yu, et al.
Published: (2025)