LITTA: Late-Interaction and Test-Time Alignment for Visually-Grounded Multimodal Retrieval
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Kim, Seonok |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction
von: Xiao, Zilin, et al.
Veröffentlicht: (2025)
von: Xiao, Zilin, et al.
Veröffentlicht: (2025)
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval
von: Wan, David, et al.
Veröffentlicht: (2025)
von: Wan, David, et al.
Veröffentlicht: (2025)
Progressive Multimodal Reasoning via Active Retrieval
von: Dong, Guanting, et al.
Veröffentlicht: (2024)
von: Dong, Guanting, et al.
Veröffentlicht: (2024)
MMDocIR: Benchmarking Multimodal Retrieval for Long Documents
von: Dong, Kuicai, et al.
Veröffentlicht: (2025)
von: Dong, Kuicai, et al.
Veröffentlicht: (2025)
M3DR: Towards Universal Multilingual Multimodal Document Retrieval
von: Kolavi, Adithya S, et al.
Veröffentlicht: (2025)
von: Kolavi, Adithya S, et al.
Veröffentlicht: (2025)
MegaRAG: Multimodal Knowledge Graph-Based Retrieval Augmented Generation
von: Hsiao, Chi-Hsiang, et al.
Veröffentlicht: (2025)
von: Hsiao, Chi-Hsiang, et al.
Veröffentlicht: (2025)
VDocRAG: Retrieval-Augmented Generation over Visually-Rich Documents
von: Tanaka, Ryota, et al.
Veröffentlicht: (2025)
von: Tanaka, Ryota, et al.
Veröffentlicht: (2025)
Multi-Grained Query-Guided Set Prediction Network for Grounded Multimodal Named Entity Recognition
von: Tang, Jielong, et al.
Veröffentlicht: (2024)
von: Tang, Jielong, et al.
Veröffentlicht: (2024)
Utility-Oriented Visual Evidence Selection for Multimodal Retrieval-Augmented Generation
von: Luo, Weiqing, et al.
Veröffentlicht: (2026)
von: Luo, Weiqing, et al.
Veröffentlicht: (2026)
Chain of Evidence: Pixel-Level Visual Attribution for Iterative Retrieval-Augmented Generation
von: Liu, Peiyang, et al.
Veröffentlicht: (2026)
von: Liu, Peiyang, et al.
Veröffentlicht: (2026)
XL-HeadTags: Leveraging Multimodal Retrieval Augmentation for the Multilingual Generation of News Headlines and Tags
von: Shohan, Faisal Tareque, et al.
Veröffentlicht: (2024)
von: Shohan, Faisal Tareque, et al.
Veröffentlicht: (2024)
Generative Cross-Modal Retrieval: Memorizing Images in Multimodal Language Models for Retrieval and Beyond
von: Li, Yongqi, et al.
Veröffentlicht: (2024)
von: Li, Yongqi, et al.
Veröffentlicht: (2024)
ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents
von: Wang, Qiuchen, et al.
Veröffentlicht: (2025)
von: Wang, Qiuchen, et al.
Veröffentlicht: (2025)
VideoComp: Advancing Fine-Grained Compositional and Temporal Alignment in Video-Text Models
von: Kim, Dahun, et al.
Veröffentlicht: (2025)
von: Kim, Dahun, et al.
Veröffentlicht: (2025)
MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs
von: Lin, Sheng-Chieh, et al.
Veröffentlicht: (2024)
von: Lin, Sheng-Chieh, et al.
Veröffentlicht: (2024)
RE-TRIANGLE: Does TRIANGLE Enable Multimodal Alignment Beyond Cosine Similarity in Retrieval?
von: Ghosh, Arijit, et al.
Veröffentlicht: (2026)
von: Ghosh, Arijit, et al.
Veröffentlicht: (2026)
MIRe: Enhancing Multimodal Queries Representation via Fusion-Free Modality Interaction for Multimodal Retrieval
von: Ju, Yeong-Joon, et al.
Veröffentlicht: (2024)
von: Ju, Yeong-Joon, et al.
Veröffentlicht: (2024)
Attention Grounded Enhancement for Visual Document Retrieval
von: Cui, Wanqing, et al.
Veröffentlicht: (2025)
von: Cui, Wanqing, et al.
Veröffentlicht: (2025)
HV-Attack: Hierarchical Visual Attack for Multimodal Retrieval Augmented Generation
von: Luo, Linyin, et al.
Veröffentlicht: (2025)
von: Luo, Linyin, et al.
Veröffentlicht: (2025)
Composite Sketch+Text Queries for Retrieving Objects with Elusive Names and Complex Interactions
von: Gatti, Prajwal, et al.
Veröffentlicht: (2025)
von: Gatti, Prajwal, et al.
Veröffentlicht: (2025)
DSRAG: A Domain-Specific Retrieval Framework Based on Document-derived Multimodal Knowledge Graph
von: Yang, Mengzheng, et al.
Veröffentlicht: (2025)
von: Yang, Mengzheng, et al.
Veröffentlicht: (2025)
Visual Late Chunking: An Empirical Study of Contextual Chunking for Efficient Visual Document Retrieval
von: Yan, Yibo, et al.
Veröffentlicht: (2026)
von: Yan, Yibo, et al.
Veröffentlicht: (2026)
Personalized Multimodal Large Language Models: A Survey
von: Wu, Junda, et al.
Veröffentlicht: (2024)
von: Wu, Junda, et al.
Veröffentlicht: (2024)
GPT-4V(ision) is a Generalist Web Agent, if Grounded
von: Zheng, Boyuan, et al.
Veröffentlicht: (2024)
von: Zheng, Boyuan, et al.
Veröffentlicht: (2024)
VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings
von: Giahi, Ramin, et al.
Veröffentlicht: (2025)
von: Giahi, Ramin, et al.
Veröffentlicht: (2025)
INQUIRE: A Natural World Text-to-Image Retrieval Benchmark
von: Vendrow, Edward, et al.
Veröffentlicht: (2024)
von: Vendrow, Edward, et al.
Veröffentlicht: (2024)
Beyond Unimodal Boundaries: Generative Recommendation with Multimodal Semantics
von: Zhu, Jing, et al.
Veröffentlicht: (2025)
von: Zhu, Jing, et al.
Veröffentlicht: (2025)
VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents
von: Yu, Shi, et al.
Veröffentlicht: (2024)
von: Yu, Shi, et al.
Veröffentlicht: (2024)
Unified Interactive Multimodal Moment Retrieval via Cascaded Embedding-Reranking and Temporal-Aware Score Fusion
von: Thanh, Toan Le Ngo, et al.
Veröffentlicht: (2025)
von: Thanh, Toan Le Ngo, et al.
Veröffentlicht: (2025)
Learning Visual Composition through Improved Semantic Guidance
von: Stone, Austin, et al.
Veröffentlicht: (2024)
von: Stone, Austin, et al.
Veröffentlicht: (2024)
From Videos to Indexed Knowledge Graphs -- Framework to Marry Methods for Multimodal Content Analysis and Understanding
von: Rizk, Basem, et al.
Veröffentlicht: (2025)
von: Rizk, Basem, et al.
Veröffentlicht: (2025)
ViBERTgrid BiLSTM-CRF: Multimodal Key Information Extraction from Unstructured Financial Documents
von: Pala, Furkan, et al.
Veröffentlicht: (2024)
von: Pala, Furkan, et al.
Veröffentlicht: (2024)
Towards Universal Video Retrieval: Generalizing Video Embedding via Synthesized Multimodal Pyramid Curriculum
von: Guo, Zhuoning, et al.
Veröffentlicht: (2025)
von: Guo, Zhuoning, et al.
Veröffentlicht: (2025)
ITSELF: Attention Guided Fine-Grained Alignment for Vision-Language Retrieval
von: Nguyen, Tien-Huy, et al.
Veröffentlicht: (2026)
von: Nguyen, Tien-Huy, et al.
Veröffentlicht: (2026)
VideoRAG: Retrieval-Augmented Generation over Video Corpus
von: Jeong, Soyeong, et al.
Veröffentlicht: (2025)
von: Jeong, Soyeong, et al.
Veröffentlicht: (2025)
Diffusion Augmented Retrieval: A Training-Free Approach to Interactive Text-to-Image Retrieval
von: Long, Zijun, et al.
Veröffentlicht: (2025)
von: Long, Zijun, et al.
Veröffentlicht: (2025)
Reversed in Time: A Novel Temporal-Emphasized Benchmark for Cross-Modal Video-Text Retrieval
von: Du, Yang, et al.
Veröffentlicht: (2024)
von: Du, Yang, et al.
Veröffentlicht: (2024)
Smart Routing for Multimodal Video Retrieval: When to Search What
von: Rosa, Kevin Dela
Veröffentlicht: (2025)
von: Rosa, Kevin Dela
Veröffentlicht: (2025)
OMGM: Orchestrate Multiple Granularities and Modalities for Efficient Multimodal Retrieval
von: Yang, Wei, et al.
Veröffentlicht: (2025)
von: Yang, Wei, et al.
Veröffentlicht: (2025)
MagicLens: Self-Supervised Image Retrieval with Open-Ended Instructions
von: Zhang, Kai, et al.
Veröffentlicht: (2024)
von: Zhang, Kai, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction
von: Xiao, Zilin, et al.
Veröffentlicht: (2025) -
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval
von: Wan, David, et al.
Veröffentlicht: (2025) -
Progressive Multimodal Reasoning via Active Retrieval
von: Dong, Guanting, et al.
Veröffentlicht: (2024) -
MMDocIR: Benchmarking Multimodal Retrieval for Long Documents
von: Dong, Kuicai, et al.
Veröffentlicht: (2025) -
M3DR: Towards Universal Multilingual Multimodal Document Retrieval
von: Kolavi, Adithya S, et al.
Veröffentlicht: (2025)