MRAG-Bench: Vision-Centric Evaluation for Retrieval-Augmented Multimodal Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hu, Wenbo, Gu, Jia-Chen, Dou, Zi-Yi, Fayyaz, Mohsen, Lu, Pan, Chang, Kai-Wei, Peng, Nanyun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VALOR-EVAL: Holistic Coverage and Faithfulness Evaluation of Large Vision-Language Models
von: Qiu, Haoyi, et al.
Veröffentlicht: (2024)
von: Qiu, Haoyi, et al.
Veröffentlicht: (2024)
Matryoshka Query Transformer for Large Vision-Language Models
von: Hu, Wenbo, et al.
Veröffentlicht: (2024)
von: Hu, Wenbo, et al.
Veröffentlicht: (2024)
TemMed-Bench: Evaluating Temporal Medical Image Reasoning in Vision-Language Models
von: Zhang, Junyi, et al.
Veröffentlicht: (2025)
von: Zhang, Junyi, et al.
Veröffentlicht: (2025)
Collapse of Dense Retrievers: Short, Early, and Literal Biases Outranking Factual Evidence
von: Fayyaz, Mohsen, et al.
Veröffentlicht: (2025)
von: Fayyaz, Mohsen, et al.
Veröffentlicht: (2025)
Synchronous Faithfulness Monitoring for Trustworthy Retrieval-Augmented Generation
von: Wu, Di, et al.
Veröffentlicht: (2024)
von: Wu, Di, et al.
Veröffentlicht: (2024)
Medical Vision-Language Pre-Training for Brain Abnormalities
von: Monajatipoor, Masoud, et al.
Veröffentlicht: (2024)
von: Monajatipoor, Masoud, et al.
Veröffentlicht: (2024)
MRAG: Benchmarking Retrieval-Augmented Generation for Bio-medicine
von: Li, Liz, et al.
Veröffentlicht: (2026)
von: Li, Liz, et al.
Veröffentlicht: (2026)
Evaluating Human Alignment and Model Faithfulness of LLM Rationale
von: Fayyaz, Mohsen, et al.
Veröffentlicht: (2024)
von: Fayyaz, Mohsen, et al.
Veröffentlicht: (2024)
Self-Routing RAG: Binding Selective Retrieval with Knowledge Verbalization
von: Wu, Di, et al.
Veröffentlicht: (2025)
von: Wu, Di, et al.
Veröffentlicht: (2025)
MRAG-Suite: A Diagnostic Evaluation Platform for Visual Retrieval-Augmented Generation
von: Ji, Yuelyu, et al.
Veröffentlicht: (2025)
von: Ji, Yuelyu, et al.
Veröffentlicht: (2025)
BRIEF: Bridging Retrieval and Inference for Multi-hop Reasoning via Compression
von: Li, Yuankai, et al.
Veröffentlicht: (2024)
von: Li, Yuankai, et al.
Veröffentlicht: (2024)
Re-ReST: Reflection-Reinforced Self-Training for Language Agents
von: Dou, Zi-Yi, et al.
Veröffentlicht: (2024)
von: Dou, Zi-Yi, et al.
Veröffentlicht: (2024)
Multilingual Routing in Mixture-of-Experts
von: Bandarkar, Lucas, et al.
Veröffentlicht: (2025)
von: Bandarkar, Lucas, et al.
Veröffentlicht: (2025)
OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks
von: Hu, Wenbo, et al.
Veröffentlicht: (2026)
von: Hu, Wenbo, et al.
Veröffentlicht: (2026)
ConTextual: Evaluating Context-Sensitive Text-Rich Visual Reasoning in Large Multimodal Models
von: Wadhawan, Rohan, et al.
Veröffentlicht: (2024)
von: Wadhawan, Rohan, et al.
Veröffentlicht: (2024)
On Asymmetric Optimization of Reasoning and Perception in Vision-Language Model Post-Training
von: Wu, Xueqing, et al.
Veröffentlicht: (2026)
von: Wu, Xueqing, et al.
Veröffentlicht: (2026)
MRAG: A Modular Retrieval Framework for Time-Sensitive Question Answering
von: Siyue, Zhang, et al.
Veröffentlicht: (2024)
von: Siyue, Zhang, et al.
Veröffentlicht: (2024)
Poisoned-MRAG: Knowledge Poisoning Attacks to Multimodal Retrieval Augmented Generation
von: Liu, Yinuo, et al.
Veröffentlicht: (2025)
von: Liu, Yinuo, et al.
Veröffentlicht: (2025)
OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles
von: Deng, Yihe, et al.
Veröffentlicht: (2025)
von: Deng, Yihe, et al.
Veröffentlicht: (2025)
LongMemEval-V2: Evaluating Long-Term Agent Memory Toward Experienced Colleagues
von: Wu, Di, et al.
Veröffentlicht: (2026)
von: Wu, Di, et al.
Veröffentlicht: (2026)
METAL: A Multi-Agent Framework for Chart Generation with Test-Time Scaling
von: Li, Bingxuan, et al.
Veröffentlicht: (2025)
von: Li, Bingxuan, et al.
Veröffentlicht: (2025)
Verbalized Representation Learning for Interpretable Few-Shot Generalization
von: Yang, Cheng-Fu, et al.
Veröffentlicht: (2024)
von: Yang, Cheng-Fu, et al.
Veröffentlicht: (2024)
Model Editing Harms General Abilities of Large Language Models: Regularization to the Rescue
von: Gu, Jia-Chen, et al.
Veröffentlicht: (2024)
von: Gu, Jia-Chen, et al.
Veröffentlicht: (2024)
BRIEF-Pro: Universal Context Compression with Short-to-Long Synthesis for Fast and Accurate Multi-Hop Reasoning
von: Gu, Jia-Chen, et al.
Veröffentlicht: (2025)
von: Gu, Jia-Chen, et al.
Veröffentlicht: (2025)
PhonologyBench: Evaluating Phonological Skills of Large Language Models
von: Suvarna, Ashima, et al.
Veröffentlicht: (2024)
von: Suvarna, Ashima, et al.
Veröffentlicht: (2024)
VaPR -- Vision-language Preference alignment for Reasoning
von: Wadhawan, Rohan, et al.
Veröffentlicht: (2025)
von: Wadhawan, Rohan, et al.
Veröffentlicht: (2025)
RAD-Bench: Evaluating Large Language Models Capabilities in Retrieval Augmented Dialogues
von: Kuo, Tzu-Lin, et al.
Veröffentlicht: (2024)
von: Kuo, Tzu-Lin, et al.
Veröffentlicht: (2024)
DialectGen: Benchmarking and Improving Dialect Robustness in Multimodal Generation
von: Zhou, Yu, et al.
Veröffentlicht: (2025)
von: Zhou, Yu, et al.
Veröffentlicht: (2025)
VISCO: Benchmarking Fine-Grained Critique and Correction Towards Self-Improvement in Visual Reasoning
von: Wu, Xueqing, et al.
Veröffentlicht: (2024)
von: Wu, Xueqing, et al.
Veröffentlicht: (2024)
Steering MoE LLMs via Expert (De)Activation
von: Fayyaz, Mohsen, et al.
Veröffentlicht: (2025)
von: Fayyaz, Mohsen, et al.
Veröffentlicht: (2025)
VDebugger: Harnessing Execution Feedback for Debugging Visual Programs
von: Wu, Xueqing, et al.
Veröffentlicht: (2024)
von: Wu, Xueqing, et al.
Veröffentlicht: (2024)
DeepEdit: Knowledge Editing as Decoding with Constraints
von: Wang, Yiwei, et al.
Veröffentlicht: (2024)
von: Wang, Yiwei, et al.
Veröffentlicht: (2024)
Fact or Guesswork? Evaluating Large Language Models' Medical Knowledge with Structured One-Hop Judgments
von: Li, Jiaxi, et al.
Veröffentlicht: (2025)
von: Li, Jiaxi, et al.
Veröffentlicht: (2025)
VLR-Bench: Multilingual Benchmark Dataset for Vision-Language Retrieval Augmented Generation
von: Lim, Hyeonseok, et al.
Veröffentlicht: (2024)
von: Lim, Hyeonseok, et al.
Veröffentlicht: (2024)
GenEARL: A Training-Free Generative Framework for Multimodal Event Argument Role Labeling
von: Bansal, Hritik, et al.
Veröffentlicht: (2024)
von: Bansal, Hritik, et al.
Veröffentlicht: (2024)
Multimodal Cultural Safety: Evaluation Framework and Alignment Strategies
von: Qiu, Haoyi, et al.
Veröffentlicht: (2025)
von: Qiu, Haoyi, et al.
Veröffentlicht: (2025)
QUDSELECT: Selective Decoding for Questions Under Discussion Parsing
von: Suvarna, Ashima, et al.
Veröffentlicht: (2024)
von: Suvarna, Ashima, et al.
Veröffentlicht: (2024)
Control Large Language Models via Divide and Conquer
von: Li, Bingxuan, et al.
Veröffentlicht: (2024)
von: Li, Bingxuan, et al.
Veröffentlicht: (2024)
IW-Bench: Evaluating Large Multimodal Models for Converting Image-to-Web
von: Guo, Hongcheng, et al.
Veröffentlicht: (2024)
von: Guo, Hongcheng, et al.
Veröffentlicht: (2024)
DocSLM: A Small Vision-Language Model for Long Multimodal Document Understanding
von: Hannan, Tanveer, et al.
Veröffentlicht: (2025)
von: Hannan, Tanveer, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
VALOR-EVAL: Holistic Coverage and Faithfulness Evaluation of Large Vision-Language Models
von: Qiu, Haoyi, et al.
Veröffentlicht: (2024) -
Matryoshka Query Transformer for Large Vision-Language Models
von: Hu, Wenbo, et al.
Veröffentlicht: (2024) -
TemMed-Bench: Evaluating Temporal Medical Image Reasoning in Vision-Language Models
von: Zhang, Junyi, et al.
Veröffentlicht: (2025) -
Collapse of Dense Retrievers: Short, Early, and Literal Biases Outranking Factual Evidence
von: Fayyaz, Mohsen, et al.
Veröffentlicht: (2025) -
Synchronous Faithfulness Monitoring for Trustworthy Retrieval-Augmented Generation
von: Wu, Di, et al.
Veröffentlicht: (2024)