Guardado en:
| Autores principales: | Pei, Rongcan, Li, Huan, Guo, Fang, Zhu, Qi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2602.10146 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LongVILA: Scaling Long-Context Visual Language Models for Long Videos
por: Chen, Yukang, et al.
Publicado: (2024)
por: Chen, Yukang, et al.
Publicado: (2024)
Internalized Reasoning for Long-Context Visual Document Understanding
por: Veselka, Austin
Publicado: (2026)
por: Veselka, Austin
Publicado: (2026)
From Text to Pixel: Advancing Long-Context Understanding in MLLMs
por: Lu, Yujie, et al.
Publicado: (2024)
por: Lu, Yujie, et al.
Publicado: (2024)
No Tokens Wasted: Leveraging Long Context in Biomedical Vision-Language Models
por: Sun, Min Woo, et al.
Publicado: (2025)
por: Sun, Min Woo, et al.
Publicado: (2025)
Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding
por: Gao, Sensen, et al.
Publicado: (2025)
por: Gao, Sensen, et al.
Publicado: (2025)
RAVENEA: A Benchmark for Multimodal Retrieval-Augmented Visual Culture Understanding
por: Li, Jiaang, et al.
Publicado: (2025)
por: Li, Jiaang, et al.
Publicado: (2025)
DocLens : A Tool-Augmented Multi-Agent Framework for Long Visual Document Understanding
por: Zhu, Dawei, et al.
Publicado: (2025)
por: Zhu, Dawei, et al.
Publicado: (2025)
EviMem: Evidence-Gap-Driven Iterative Retrieval for Long-Term Conversational Memory
por: Li, Yuyang, et al.
Publicado: (2026)
por: Li, Yuyang, et al.
Publicado: (2026)
OmChat: A Recipe to Train Multimodal Language Models with Strong Long Context and Video Understanding
por: Zhao, Tiancheng, et al.
Publicado: (2024)
por: Zhao, Tiancheng, et al.
Publicado: (2024)
Understanding Retrieval Robustness for Retrieval-Augmented Image Captioning
por: Li, Wenyan, et al.
Publicado: (2024)
por: Li, Wenyan, et al.
Publicado: (2024)
VimRAG: Navigating Massive Visual Context in Retrieval-Augmented Generation via Multimodal Memory Graph
por: Wang, Qiuchen, et al.
Publicado: (2026)
por: Wang, Qiuchen, et al.
Publicado: (2026)
MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations
por: Ma, Yubo, et al.
Publicado: (2024)
por: Ma, Yubo, et al.
Publicado: (2024)
Retrieving Counterfactuals Improves Visual In-Context Learning
por: Xiong, Guangzhi, et al.
Publicado: (2026)
por: Xiong, Guangzhi, et al.
Publicado: (2026)
Support or Refute: Analyzing the Stance of Evidence to Detect Out-of-Context Mis- and Disinformation
por: Yuan, Xin, et al.
Publicado: (2023)
por: Yuan, Xin, et al.
Publicado: (2023)
Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models
por: Zhou, Yucheng, et al.
Publicado: (2024)
por: Zhou, Yucheng, et al.
Publicado: (2024)
Cracking the Code of Hallucination in LVLMs with Vision-aware Head Divergence
por: He, Jinghan, et al.
Publicado: (2024)
por: He, Jinghan, et al.
Publicado: (2024)
Can Retrieval Heads See Images? Multimodal Retrieval Heads in Long-Context Vision-Language Models
por: Li, Aaron Branson Cigres, et al.
Publicado: (2026)
por: Li, Aaron Branson Cigres, et al.
Publicado: (2026)
VisionGraph: Leveraging Large Multimodal Models for Graph Theory Problems in Visual Context
por: Li, Yunxin, et al.
Publicado: (2024)
por: Li, Yunxin, et al.
Publicado: (2024)
VTCBench: Can Vision-Language Models Understand Long Context with Vision-Text Compression?
por: Zhao, Hongbo, et al.
Publicado: (2025)
por: Zhao, Hongbo, et al.
Publicado: (2025)
ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding
por: Ma, David, et al.
Publicado: (2025)
por: Ma, David, et al.
Publicado: (2025)
VisRAG 2.0: Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation
por: Sun, Yubo, et al.
Publicado: (2025)
por: Sun, Yubo, et al.
Publicado: (2025)
GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models
por: Li, Mukai, et al.
Publicado: (2024)
por: Li, Mukai, et al.
Publicado: (2024)
EmoGist: Efficient In-Context Learning for Visual Emotion Understanding
por: Seoh, Ronald, et al.
Publicado: (2025)
por: Seoh, Ronald, et al.
Publicado: (2025)
Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding
por: Wang, Ziyang, et al.
Publicado: (2025)
por: Wang, Ziyang, et al.
Publicado: (2025)
How to Train Your Long-Context Visual Document Model
por: Veselka, Austin
Publicado: (2026)
por: Veselka, Austin
Publicado: (2026)
DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding
por: Zhang, Hongzhi, et al.
Publicado: (2025)
por: Zhang, Hongzhi, et al.
Publicado: (2025)
Efficient End-to-End Visual Document Understanding with Rationale Distillation
por: Zhu, Wang, et al.
Publicado: (2023)
por: Zhu, Wang, et al.
Publicado: (2023)
Visual In-Context Learning for Large Vision-Language Models
por: Zhou, Yucheng, et al.
Publicado: (2024)
por: Zhou, Yucheng, et al.
Publicado: (2024)
LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference
por: Wan, Zhongwei, et al.
Publicado: (2024)
por: Wan, Zhongwei, et al.
Publicado: (2024)
Mitigating GenAI-powered Evidence Pollution for Out-of-Context Multimodal Misinformation Detection
por: Yan, Zehong, et al.
Publicado: (2025)
por: Yan, Zehong, et al.
Publicado: (2025)
XL-HeadTags: Leveraging Multimodal Retrieval Augmentation for the Multilingual Generation of News Headlines and Tags
por: Shohan, Faisal Tareque, et al.
Publicado: (2024)
por: Shohan, Faisal Tareque, et al.
Publicado: (2024)
Hierarchical Multimodal Pre-training for Visually Rich Webpage Understanding
por: Xu, Hongshen, et al.
Publicado: (2024)
por: Xu, Hongshen, et al.
Publicado: (2024)
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding
por: Wang, Zhaokai, et al.
Publicado: (2025)
por: Wang, Zhaokai, et al.
Publicado: (2025)
Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval
por: Sun, Hao, et al.
Publicado: (2026)
por: Sun, Hao, et al.
Publicado: (2026)
Seeing is Believing: Rich-Context Hallucination Detection for MLLMs via Backward Visual Grounding
por: Guo, Pinxue, et al.
Publicado: (2025)
por: Guo, Pinxue, et al.
Publicado: (2025)
Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries
por: Wu, Yin, et al.
Publicado: (2025)
por: Wu, Yin, et al.
Publicado: (2025)
ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding
por: Wang, Xiao, et al.
Publicado: (2024)
por: Wang, Xiao, et al.
Publicado: (2024)
DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies
por: Song, Wei, et al.
Publicado: (2025)
por: Song, Wei, et al.
Publicado: (2025)
Bi-VLDoc: Bidirectional Vision-Language Modeling for Visually-Rich Document Understanding
por: Luo, Chuwei, et al.
Publicado: (2022)
por: Luo, Chuwei, et al.
Publicado: (2022)
Context-aware Visual Storytelling with Visual Prefix Tuning and Contrastive Learning
por: Song, Yingjin, et al.
Publicado: (2024)
por: Song, Yingjin, et al.
Publicado: (2024)
Ejemplares similares
-
LongVILA: Scaling Long-Context Visual Language Models for Long Videos
por: Chen, Yukang, et al.
Publicado: (2024) -
Internalized Reasoning for Long-Context Visual Document Understanding
por: Veselka, Austin
Publicado: (2026) -
From Text to Pixel: Advancing Long-Context Understanding in MLLMs
por: Lu, Yujie, et al.
Publicado: (2024) -
No Tokens Wasted: Leveraging Long Context in Biomedical Vision-Language Models
por: Sun, Min Woo, et al.
Publicado: (2025) -
Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding
por: Gao, Sensen, et al.
Publicado: (2025)