CRAB: A Benchmark for Evaluating Curation of Retrieval-Augmented LLMs in Biomedicine
Fuente:
arXiv
Saved in:
| Main Authors: | Zhong, Hanmeng, Chen, Linqing, Wu, Wentao, Wang, Weilei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Semantic Bridge: Universal Multi-Hop Question Generation via AMR-Driven Graph Synthesis
by: Chen, Linqing, et al.
Published: (2025)
by: Chen, Linqing, et al.
Published: (2025)
Hierarchical Multi-Label Generation with Probabilistic Level-Constraint
by: Chen, Linqing, et al.
Published: (2025)
by: Chen, Linqing, et al.
Published: (2025)
Large Language Model's Multi-Capability Alignment in Biomedical Domain
by: Wu, Wentao, et al.
Published: (2025)
by: Wu, Wentao, et al.
Published: (2025)
BiomedRAG: A Retrieval Augmented Large Language Model for Biomedicine
by: Li, Mingchen, et al.
Published: (2024)
by: Li, Mingchen, et al.
Published: (2024)
Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs
by: Son, Guijin, et al.
Published: (2026)
by: Son, Guijin, et al.
Published: (2026)
LLMs are Biased Evaluators But Not Biased for Retrieval Augmented Generation
by: Chen, Yen-Shan, et al.
Published: (2024)
by: Chen, Yen-Shan, et al.
Published: (2024)
On LLMs-Driven Synthetic Data Generation, Curation, and Evaluation: A Survey
by: Long, Lin, et al.
Published: (2024)
by: Long, Lin, et al.
Published: (2024)
Streamlining Biomedical Research with Specialized LLMs
by: Chen, Linqing, et al.
Published: (2025)
by: Chen, Linqing, et al.
Published: (2025)
RAGCap-Bench: Benchmarking Capabilities of LLMs in Agentic Retrieval Augmented Generation Systems
by: Lin, Jingru, et al.
Published: (2025)
by: Lin, Jingru, et al.
Published: (2025)
LLMs in Biomedicine: A study on clinical Named Entity Recognition
by: Monajatipoor, Masoud, et al.
Published: (2024)
by: Monajatipoor, Masoud, et al.
Published: (2024)
Does RAG Introduce Unfairness in LLMs? Evaluating Fairness in Retrieval-Augmented Generation Systems
by: Wu, Xuyang, et al.
Published: (2024)
by: Wu, Xuyang, et al.
Published: (2024)
SciRerankBench: Benchmarking Rerankers Towards Scientific Retrieval-Augmented Generated LLMs
by: Chen, Haotian, et al.
Published: (2025)
by: Chen, Haotian, et al.
Published: (2025)
When Do LLMs Need Retrieval Augmentation? Mitigating LLMs' Overconfidence Helps Retrieval Augmentation
by: Ni, Shiyu, et al.
Published: (2024)
by: Ni, Shiyu, et al.
Published: (2024)
RARE: Retrieval-Augmented Reasoning Modeling
by: Wang, Zhengren, et al.
Published: (2025)
by: Wang, Zhengren, et al.
Published: (2025)
Towards Automated Smart Contract Generation: Evaluation, Benchmarking, and Retrieval-Augmented Repair
by: Chen, Zaoyu, et al.
Published: (2025)
by: Chen, Zaoyu, et al.
Published: (2025)
Persona-Augmented Benchmarking: Evaluating LLMs Across Diverse Writing Styles
by: Truong, Kimberly Le, et al.
Published: (2025)
by: Truong, Kimberly Le, et al.
Published: (2025)
DAHL: Domain-specific Automated Hallucination Evaluation of Long-Form Text through a Benchmark Dataset in Biomedicine
by: Seo, Jean, et al.
Published: (2024)
by: Seo, Jean, et al.
Published: (2024)
Unanswerability Evaluation for Retrieval Augmented Generation
by: Peng, Xiangyu, et al.
Published: (2024)
by: Peng, Xiangyu, et al.
Published: (2024)
Benchmarking Retrieval-Augmented Generation for Chemistry
by: Zhong, Xianrui, et al.
Published: (2025)
by: Zhong, Xianrui, et al.
Published: (2025)
RARE: Retrieval-Aware Robustness Evaluation for Retrieval-Augmented Generation Systems
by: Zeng, Yixiao, et al.
Published: (2025)
by: Zeng, Yixiao, et al.
Published: (2025)
AgenticOCR: Parsing Only What You Need for Efficient Retrieval-Augmented Generation
by: Wang, Zhengren, et al.
Published: (2026)
by: Wang, Zhengren, et al.
Published: (2026)
MIRAGE: A Metric-Intensive Benchmark for Retrieval-Augmented Generation Evaluation
by: Park, Chanhee, et al.
Published: (2025)
by: Park, Chanhee, et al.
Published: (2025)
CFVBench: A Comprehensive Video Benchmark for Fine-grained Multimodal Retrieval-Augmented Generation
by: Wei, Kaiwen, et al.
Published: (2025)
by: Wei, Kaiwen, et al.
Published: (2025)
Med-R$^3$: Enhancing Medical Retrieval-Augmented Reasoning of LLMs via Progressive Reinforcement Learning
by: Lu, Keer, et al.
Published: (2025)
by: Lu, Keer, et al.
Published: (2025)
Learning to Shop Like Humans: A Review-driven Retrieval-Augmented Recommendation Framework with LLMs
by: Wei, Kaiwen, et al.
Published: (2025)
by: Wei, Kaiwen, et al.
Published: (2025)
Evaluating Cultural Knowledge Processing in Large Language Models: A Cognitive Benchmarking Framework Integrating Retrieval-Augmented Generation
by: Lee, Hung-Shin, et al.
Published: (2025)
by: Lee, Hung-Shin, et al.
Published: (2025)
A Survey for Large Language Models in Biomedicine
by: Wang, Chong, et al.
Published: (2024)
by: Wang, Chong, et al.
Published: (2024)
CRUD-RAG: A Comprehensive Chinese Benchmark for Retrieval-Augmented Generation of Large Language Models
by: Lyu, Yuanjie, et al.
Published: (2024)
by: Lyu, Yuanjie, et al.
Published: (2024)
DomainRAG: A Chinese Benchmark for Evaluating Domain-specific Retrieval-Augmented Generation
by: Wang, Shuting, et al.
Published: (2024)
by: Wang, Shuting, et al.
Published: (2024)
Lifelong Knowledge Editing for LLMs with Retrieval-Augmented Continuous Prompt Learning
by: Chen, Qizhou, et al.
Published: (2024)
by: Chen, Qizhou, et al.
Published: (2024)
PRGB Benchmark: A Robust Placeholder-Assisted Algorithm for Benchmarking Retrieval-Augmented Generation
by: Tan, Zhehao, et al.
Published: (2025)
by: Tan, Zhehao, et al.
Published: (2025)
REX-RAG: Reasoning Exploration with Policy Correction in Retrieval-Augmented Generation
by: Jiang, Wentao, et al.
Published: (2025)
by: Jiang, Wentao, et al.
Published: (2025)
Compound-QA: A Benchmark for Evaluating LLMs on Compound Questions
by: Hou, Yutao, et al.
Published: (2024)
by: Hou, Yutao, et al.
Published: (2024)
Predict the Retrieval! Test time adaptation for Retrieval Augmented Generation
by: Sun, Xin, et al.
Published: (2026)
by: Sun, Xin, et al.
Published: (2026)
Dialogue Benchmark Generation from Knowledge Graphs with Cost-Effective Retrieval-Augmented LLMs
by: Omar, Reham, et al.
Published: (2025)
by: Omar, Reham, et al.
Published: (2025)
BERGEN: A Benchmarking Library for Retrieval-Augmented Generation
by: Rau, David, et al.
Published: (2024)
by: Rau, David, et al.
Published: (2024)
AI for Biomedicine in the Era of Large Language Models
by: Bi, Zhenyu, et al.
Published: (2024)
by: Bi, Zhenyu, et al.
Published: (2024)
AgenticRAGTracer: A Hop-Aware Benchmark for Diagnosing Multi-Step Retrieval Reasoning in Agentic RAG
by: You, Qijie, et al.
Published: (2026)
by: You, Qijie, et al.
Published: (2026)
YpathRAG:A Retrieval-Augmented Generation Framework and Benchmark for Pathology
by: Yu, Deshui, et al.
Published: (2025)
by: Yu, Deshui, et al.
Published: (2025)
UltraMedical: Building Specialized Generalists in Biomedicine
by: Zhang, Kaiyan, et al.
Published: (2024)
by: Zhang, Kaiyan, et al.
Published: (2024)
Similar Items
-
Semantic Bridge: Universal Multi-Hop Question Generation via AMR-Driven Graph Synthesis
by: Chen, Linqing, et al.
Published: (2025) -
Hierarchical Multi-Label Generation with Probabilistic Level-Constraint
by: Chen, Linqing, et al.
Published: (2025) -
Large Language Model's Multi-Capability Alignment in Biomedical Domain
by: Wu, Wentao, et al.
Published: (2025) -
BiomedRAG: A Retrieval Augmented Large Language Model for Biomedicine
by: Li, Mingchen, et al.
Published: (2024) -
Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs
by: Son, Guijin, et al.
Published: (2026)