Nearest Neighbor Normalization Improves Multimodal Retrieval
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chowdhury, Neil, Wang, Franklin, Shenoy, Sumedh, Kiela, Douwe, Schwettmann, Sarah, Thrush, Tristan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Multimodal Automated Interpretability Agent
von: Shaham, Tamar Rott, et al.
Veröffentlicht: (2024)
von: Shaham, Tamar Rott, et al.
Veröffentlicht: (2024)
I am a Strange Dataset: Metalinguistic Tests for Language Models
von: Thrush, Tristan, et al.
Veröffentlicht: (2024)
von: Thrush, Tristan, et al.
Veröffentlicht: (2024)
Automatic Discovery of Visual Circuits
von: Rajaram, Achyuta, et al.
Veröffentlicht: (2024)
von: Rajaram, Achyuta, et al.
Veröffentlicht: (2024)
Leveraging Diffusion Perturbations for Measuring Fairness in Computer Vision
von: Lui, Nicholas, et al.
Veröffentlicht: (2023)
von: Lui, Nicholas, et al.
Veröffentlicht: (2023)
ColorSwap: A Color and Word Order Dataset for Multimodal Evaluation
von: Burapacheep, Jirayu, et al.
Veröffentlicht: (2024)
von: Burapacheep, Jirayu, et al.
Veröffentlicht: (2024)
Line of Sight: On Linear Representations in VLLMs
von: Rajaram, Achyuta, et al.
Veröffentlicht: (2025)
von: Rajaram, Achyuta, et al.
Veröffentlicht: (2025)
Retrieving Counterfactuals Improves Visual In-Context Learning
von: Xiong, Guangzhi, et al.
Veröffentlicht: (2026)
von: Xiong, Guangzhi, et al.
Veröffentlicht: (2026)
Efficient approximation of Earth Mover's Distance Based on Nearest Neighbor Search
von: Meng, Guangyu, et al.
Veröffentlicht: (2024)
von: Meng, Guangyu, et al.
Veröffentlicht: (2024)
MERRIN: A Benchmark for Multimodal Evidence Retrieval and Reasoning in Noisy Web Environments
von: Wang, Han, et al.
Veröffentlicht: (2026)
von: Wang, Han, et al.
Veröffentlicht: (2026)
MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs
von: Kil, Jihyung, et al.
Veröffentlicht: (2024)
von: Kil, Jihyung, et al.
Veröffentlicht: (2024)
MRAG-Bench: Vision-Centric Evaluation for Retrieval-Augmented Multimodal Models
von: Hu, Wenbo, et al.
Veröffentlicht: (2024)
von: Hu, Wenbo, et al.
Veröffentlicht: (2024)
Multimodal Retrieval-Augmented Generation with Large Language Models for Medical VQA
von: Karim, A H M Rezaul, et al.
Veröffentlicht: (2025)
von: Karim, A H M Rezaul, et al.
Veröffentlicht: (2025)
How Does the Textual Information Affect the Retrieval of Multimodal In-Context Learning?
von: Luo, Yang, et al.
Veröffentlicht: (2024)
von: Luo, Yang, et al.
Veröffentlicht: (2024)
Recurrence Meets Transformers for Universal Multimodal Retrieval
von: Caffagni, Davide, et al.
Veröffentlicht: (2025)
von: Caffagni, Davide, et al.
Veröffentlicht: (2025)
PP-DocBee2: Improved Baselines with Efficient Data for Multimodal Document Understanding
von: Huang, Kui, et al.
Veröffentlicht: (2025)
von: Huang, Kui, et al.
Veröffentlicht: (2025)
Harnessing PDF Data for Improving Japanese Large Multimodal Models
von: Baek, Jeonghun, et al.
Veröffentlicht: (2025)
von: Baek, Jeonghun, et al.
Veröffentlicht: (2025)
PP-DocBee: Improving Multimodal Document Understanding Through a Bag of Tricks
von: Ni, Feng, et al.
Veröffentlicht: (2025)
von: Ni, Feng, et al.
Veröffentlicht: (2025)
CoProNN: Concept-based Prototypical Nearest Neighbors for Explaining Vision Models
von: Chiaburu, Teodor, et al.
Veröffentlicht: (2024)
von: Chiaburu, Teodor, et al.
Veröffentlicht: (2024)
SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-Rewards
von: Hong, Jixiang, et al.
Veröffentlicht: (2025)
von: Hong, Jixiang, et al.
Veröffentlicht: (2025)
SPD-Faith Bench: Diagnosing and Improving Faithfulness in Chain-of-Thought for Multimodal Large Language Models
von: Lv, Weijiang, et al.
Veröffentlicht: (2026)
von: Lv, Weijiang, et al.
Veröffentlicht: (2026)
mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Data
von: Chen, Haonan, et al.
Veröffentlicht: (2025)
von: Chen, Haonan, et al.
Veröffentlicht: (2025)
ASPIRE: Language-Guided Data Augmentation for Improving Robustness Against Spurious Correlations
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2023)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2023)
When to Trust the Answer: Question-Aligned Semantic Nearest Neighbor Entropy for Safer Surgical VQA
von: Carlini, Luca, et al.
Veröffentlicht: (2025)
von: Carlini, Luca, et al.
Veröffentlicht: (2025)
MERLIN: Multimodal Embedding Refinement via LLM-based Iterative Navigation for Text-Video Retrieval-Rerank Pipeline
von: Han, Donghoon, et al.
Veröffentlicht: (2024)
von: Han, Donghoon, et al.
Veröffentlicht: (2024)
Recurrence-Enhanced Vision-and-Language Transformers for Robust Multimodal Document Retrieval
von: Caffagni, Davide, et al.
Veröffentlicht: (2025)
von: Caffagni, Davide, et al.
Veröffentlicht: (2025)
Wiki-LLaVA: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs
von: Caffagni, Davide, et al.
Veröffentlicht: (2024)
von: Caffagni, Davide, et al.
Veröffentlicht: (2024)
Progressive Multimodal Reasoning via Active Retrieval
von: Dong, Guanting, et al.
Veröffentlicht: (2024)
von: Dong, Guanting, et al.
Veröffentlicht: (2024)
Exploiting Mixture-of-Experts Redundancy Unlocks Multimodal Generative Abilities
von: Dutt, Raman, et al.
Veröffentlicht: (2025)
von: Dutt, Raman, et al.
Veröffentlicht: (2025)
RaTA-Tool: Retrieval-based Tool Selection with Multimodal Large Language Models
von: Mattioli, Gabriele, et al.
Veröffentlicht: (2026)
von: Mattioli, Gabriele, et al.
Veröffentlicht: (2026)
MegaRAG: Multimodal Knowledge Graph-Based Retrieval Augmented Generation
von: Hsiao, Chi-Hsiang, et al.
Veröffentlicht: (2025)
von: Hsiao, Chi-Hsiang, et al.
Veröffentlicht: (2025)
MMDocIR: Benchmarking Multimodal Retrieval for Long Documents
von: Dong, Kuicai, et al.
Veröffentlicht: (2025)
von: Dong, Kuicai, et al.
Veröffentlicht: (2025)
Multimodal Fact-Level Attribution for Verifiable Reasoning
von: Wan, David, et al.
Veröffentlicht: (2026)
von: Wan, David, et al.
Veröffentlicht: (2026)
Generative Cross-Modal Retrieval: Memorizing Images in Multimodal Language Models for Retrieval and Beyond
von: Li, Yongqi, et al.
Veröffentlicht: (2024)
von: Li, Yongqi, et al.
Veröffentlicht: (2024)
Gemini in Reasoning: Unveiling Commonsense in Multimodal Large Language Models
von: Wang, Yuqing, et al.
Veröffentlicht: (2023)
von: Wang, Yuqing, et al.
Veröffentlicht: (2023)
Mask-aware Text-to-Image Retrieval: Referring Expression Segmentation Meets Cross-modal Retrieval
von: Shen, Li-Cheng, et al.
Veröffentlicht: (2025)
von: Shen, Li-Cheng, et al.
Veröffentlicht: (2025)
Unsupervised Memorability Modeling from Tip-of-the-Tongue Retrieval Queries
von: Bhattacharyya, Sree, et al.
Veröffentlicht: (2025)
von: Bhattacharyya, Sree, et al.
Veröffentlicht: (2025)
Toward Robust Multimodal Learning using Multimodal Foundational Models
von: Zhao, Xianbing, et al.
Veröffentlicht: (2024)
von: Zhao, Xianbing, et al.
Veröffentlicht: (2024)
Keyword-Oriented Multimodal Modeling for Euphemism Identification
von: Hu, Yuxue, et al.
Veröffentlicht: (2025)
von: Hu, Yuxue, et al.
Veröffentlicht: (2025)
MLLM-CL: Continual Learning for Multimodal Large Language Models
von: Zhao, Hongbo, et al.
Veröffentlicht: (2025)
von: Zhao, Hongbo, et al.
Veröffentlicht: (2025)
M3DR: Towards Universal Multilingual Multimodal Document Retrieval
von: Kolavi, Adithya S, et al.
Veröffentlicht: (2025)
von: Kolavi, Adithya S, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Multimodal Automated Interpretability Agent
von: Shaham, Tamar Rott, et al.
Veröffentlicht: (2024) -
I am a Strange Dataset: Metalinguistic Tests for Language Models
von: Thrush, Tristan, et al.
Veröffentlicht: (2024) -
Automatic Discovery of Visual Circuits
von: Rajaram, Achyuta, et al.
Veröffentlicht: (2024) -
Leveraging Diffusion Perturbations for Measuring Fairness in Computer Vision
von: Lui, Nicholas, et al.
Veröffentlicht: (2023) -
ColorSwap: A Color and Word Order Dataset for Multimodal Evaluation
von: Burapacheep, Jirayu, et al.
Veröffentlicht: (2024)