CRAB: A Benchmark for Evaluating Curation of Retrieval-Augmented LLMs in Biomedicine
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhong, Hanmeng, Chen, Linqing, Wu, Wentao, Wang, Weilei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Semantic Bridge: Universal Multi-Hop Question Generation via AMR-Driven Graph Synthesis
von: Chen, Linqing, et al.
Veröffentlicht: (2025)
von: Chen, Linqing, et al.
Veröffentlicht: (2025)
Hierarchical Multi-Label Generation with Probabilistic Level-Constraint
von: Chen, Linqing, et al.
Veröffentlicht: (2025)
von: Chen, Linqing, et al.
Veröffentlicht: (2025)
Large Language Model's Multi-Capability Alignment in Biomedical Domain
von: Wu, Wentao, et al.
Veröffentlicht: (2025)
von: Wu, Wentao, et al.
Veröffentlicht: (2025)
BiomedRAG: A Retrieval Augmented Large Language Model for Biomedicine
von: Li, Mingchen, et al.
Veröffentlicht: (2024)
von: Li, Mingchen, et al.
Veröffentlicht: (2024)
Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs
von: Son, Guijin, et al.
Veröffentlicht: (2026)
von: Son, Guijin, et al.
Veröffentlicht: (2026)
LLMs are Biased Evaluators But Not Biased for Retrieval Augmented Generation
von: Chen, Yen-Shan, et al.
Veröffentlicht: (2024)
von: Chen, Yen-Shan, et al.
Veröffentlicht: (2024)
On LLMs-Driven Synthetic Data Generation, Curation, and Evaluation: A Survey
von: Long, Lin, et al.
Veröffentlicht: (2024)
von: Long, Lin, et al.
Veröffentlicht: (2024)
Streamlining Biomedical Research with Specialized LLMs
von: Chen, Linqing, et al.
Veröffentlicht: (2025)
von: Chen, Linqing, et al.
Veröffentlicht: (2025)
RAGCap-Bench: Benchmarking Capabilities of LLMs in Agentic Retrieval Augmented Generation Systems
von: Lin, Jingru, et al.
Veröffentlicht: (2025)
von: Lin, Jingru, et al.
Veröffentlicht: (2025)
LLMs in Biomedicine: A study on clinical Named Entity Recognition
von: Monajatipoor, Masoud, et al.
Veröffentlicht: (2024)
von: Monajatipoor, Masoud, et al.
Veröffentlicht: (2024)
Does RAG Introduce Unfairness in LLMs? Evaluating Fairness in Retrieval-Augmented Generation Systems
von: Wu, Xuyang, et al.
Veröffentlicht: (2024)
von: Wu, Xuyang, et al.
Veröffentlicht: (2024)
SciRerankBench: Benchmarking Rerankers Towards Scientific Retrieval-Augmented Generated LLMs
von: Chen, Haotian, et al.
Veröffentlicht: (2025)
von: Chen, Haotian, et al.
Veröffentlicht: (2025)
When Do LLMs Need Retrieval Augmentation? Mitigating LLMs' Overconfidence Helps Retrieval Augmentation
von: Ni, Shiyu, et al.
Veröffentlicht: (2024)
von: Ni, Shiyu, et al.
Veröffentlicht: (2024)
RARE: Retrieval-Augmented Reasoning Modeling
von: Wang, Zhengren, et al.
Veröffentlicht: (2025)
von: Wang, Zhengren, et al.
Veröffentlicht: (2025)
Towards Automated Smart Contract Generation: Evaluation, Benchmarking, and Retrieval-Augmented Repair
von: Chen, Zaoyu, et al.
Veröffentlicht: (2025)
von: Chen, Zaoyu, et al.
Veröffentlicht: (2025)
Persona-Augmented Benchmarking: Evaluating LLMs Across Diverse Writing Styles
von: Truong, Kimberly Le, et al.
Veröffentlicht: (2025)
von: Truong, Kimberly Le, et al.
Veröffentlicht: (2025)
DAHL: Domain-specific Automated Hallucination Evaluation of Long-Form Text through a Benchmark Dataset in Biomedicine
von: Seo, Jean, et al.
Veröffentlicht: (2024)
von: Seo, Jean, et al.
Veröffentlicht: (2024)
Unanswerability Evaluation for Retrieval Augmented Generation
von: Peng, Xiangyu, et al.
Veröffentlicht: (2024)
von: Peng, Xiangyu, et al.
Veröffentlicht: (2024)
Benchmarking Retrieval-Augmented Generation for Chemistry
von: Zhong, Xianrui, et al.
Veröffentlicht: (2025)
von: Zhong, Xianrui, et al.
Veröffentlicht: (2025)
RARE: Retrieval-Aware Robustness Evaluation for Retrieval-Augmented Generation Systems
von: Zeng, Yixiao, et al.
Veröffentlicht: (2025)
von: Zeng, Yixiao, et al.
Veröffentlicht: (2025)
AgenticOCR: Parsing Only What You Need for Efficient Retrieval-Augmented Generation
von: Wang, Zhengren, et al.
Veröffentlicht: (2026)
von: Wang, Zhengren, et al.
Veröffentlicht: (2026)
MIRAGE: A Metric-Intensive Benchmark for Retrieval-Augmented Generation Evaluation
von: Park, Chanhee, et al.
Veröffentlicht: (2025)
von: Park, Chanhee, et al.
Veröffentlicht: (2025)
CFVBench: A Comprehensive Video Benchmark for Fine-grained Multimodal Retrieval-Augmented Generation
von: Wei, Kaiwen, et al.
Veröffentlicht: (2025)
von: Wei, Kaiwen, et al.
Veröffentlicht: (2025)
Med-R$^3$: Enhancing Medical Retrieval-Augmented Reasoning of LLMs via Progressive Reinforcement Learning
von: Lu, Keer, et al.
Veröffentlicht: (2025)
von: Lu, Keer, et al.
Veröffentlicht: (2025)
Learning to Shop Like Humans: A Review-driven Retrieval-Augmented Recommendation Framework with LLMs
von: Wei, Kaiwen, et al.
Veröffentlicht: (2025)
von: Wei, Kaiwen, et al.
Veröffentlicht: (2025)
Evaluating Cultural Knowledge Processing in Large Language Models: A Cognitive Benchmarking Framework Integrating Retrieval-Augmented Generation
von: Lee, Hung-Shin, et al.
Veröffentlicht: (2025)
von: Lee, Hung-Shin, et al.
Veröffentlicht: (2025)
A Survey for Large Language Models in Biomedicine
von: Wang, Chong, et al.
Veröffentlicht: (2024)
von: Wang, Chong, et al.
Veröffentlicht: (2024)
CRUD-RAG: A Comprehensive Chinese Benchmark for Retrieval-Augmented Generation of Large Language Models
von: Lyu, Yuanjie, et al.
Veröffentlicht: (2024)
von: Lyu, Yuanjie, et al.
Veröffentlicht: (2024)
DomainRAG: A Chinese Benchmark for Evaluating Domain-specific Retrieval-Augmented Generation
von: Wang, Shuting, et al.
Veröffentlicht: (2024)
von: Wang, Shuting, et al.
Veröffentlicht: (2024)
Lifelong Knowledge Editing for LLMs with Retrieval-Augmented Continuous Prompt Learning
von: Chen, Qizhou, et al.
Veröffentlicht: (2024)
von: Chen, Qizhou, et al.
Veröffentlicht: (2024)
PRGB Benchmark: A Robust Placeholder-Assisted Algorithm for Benchmarking Retrieval-Augmented Generation
von: Tan, Zhehao, et al.
Veröffentlicht: (2025)
von: Tan, Zhehao, et al.
Veröffentlicht: (2025)
REX-RAG: Reasoning Exploration with Policy Correction in Retrieval-Augmented Generation
von: Jiang, Wentao, et al.
Veröffentlicht: (2025)
von: Jiang, Wentao, et al.
Veröffentlicht: (2025)
Compound-QA: A Benchmark for Evaluating LLMs on Compound Questions
von: Hou, Yutao, et al.
Veröffentlicht: (2024)
von: Hou, Yutao, et al.
Veröffentlicht: (2024)
Predict the Retrieval! Test time adaptation for Retrieval Augmented Generation
von: Sun, Xin, et al.
Veröffentlicht: (2026)
von: Sun, Xin, et al.
Veröffentlicht: (2026)
Dialogue Benchmark Generation from Knowledge Graphs with Cost-Effective Retrieval-Augmented LLMs
von: Omar, Reham, et al.
Veröffentlicht: (2025)
von: Omar, Reham, et al.
Veröffentlicht: (2025)
BERGEN: A Benchmarking Library for Retrieval-Augmented Generation
von: Rau, David, et al.
Veröffentlicht: (2024)
von: Rau, David, et al.
Veröffentlicht: (2024)
AI for Biomedicine in the Era of Large Language Models
von: Bi, Zhenyu, et al.
Veröffentlicht: (2024)
von: Bi, Zhenyu, et al.
Veröffentlicht: (2024)
AgenticRAGTracer: A Hop-Aware Benchmark for Diagnosing Multi-Step Retrieval Reasoning in Agentic RAG
von: You, Qijie, et al.
Veröffentlicht: (2026)
von: You, Qijie, et al.
Veröffentlicht: (2026)
YpathRAG:A Retrieval-Augmented Generation Framework and Benchmark for Pathology
von: Yu, Deshui, et al.
Veröffentlicht: (2025)
von: Yu, Deshui, et al.
Veröffentlicht: (2025)
UltraMedical: Building Specialized Generalists in Biomedicine
von: Zhang, Kaiyan, et al.
Veröffentlicht: (2024)
von: Zhang, Kaiyan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Semantic Bridge: Universal Multi-Hop Question Generation via AMR-Driven Graph Synthesis
von: Chen, Linqing, et al.
Veröffentlicht: (2025) -
Hierarchical Multi-Label Generation with Probabilistic Level-Constraint
von: Chen, Linqing, et al.
Veröffentlicht: (2025) -
Large Language Model's Multi-Capability Alignment in Biomedical Domain
von: Wu, Wentao, et al.
Veröffentlicht: (2025) -
BiomedRAG: A Retrieval Augmented Large Language Model for Biomedicine
von: Li, Mingchen, et al.
Veröffentlicht: (2024) -
Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs
von: Son, Guijin, et al.
Veröffentlicht: (2026)