CRAB: A Benchmark for Evaluating Curation of Retrieval-Augmented LLMs in Biomedicine
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhong, Hanmeng, Chen, Linqing, Wu, Wentao, Wang, Weilei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Semantic Bridge: Universal Multi-Hop Question Generation via AMR-Driven Graph Synthesis
por: Chen, Linqing, et al.
Publicado: (2025)
por: Chen, Linqing, et al.
Publicado: (2025)
Hierarchical Multi-Label Generation with Probabilistic Level-Constraint
por: Chen, Linqing, et al.
Publicado: (2025)
por: Chen, Linqing, et al.
Publicado: (2025)
Large Language Model's Multi-Capability Alignment in Biomedical Domain
por: Wu, Wentao, et al.
Publicado: (2025)
por: Wu, Wentao, et al.
Publicado: (2025)
BiomedRAG: A Retrieval Augmented Large Language Model for Biomedicine
por: Li, Mingchen, et al.
Publicado: (2024)
por: Li, Mingchen, et al.
Publicado: (2024)
Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs
por: Son, Guijin, et al.
Publicado: (2026)
por: Son, Guijin, et al.
Publicado: (2026)
LLMs are Biased Evaluators But Not Biased for Retrieval Augmented Generation
por: Chen, Yen-Shan, et al.
Publicado: (2024)
por: Chen, Yen-Shan, et al.
Publicado: (2024)
On LLMs-Driven Synthetic Data Generation, Curation, and Evaluation: A Survey
por: Long, Lin, et al.
Publicado: (2024)
por: Long, Lin, et al.
Publicado: (2024)
Streamlining Biomedical Research with Specialized LLMs
por: Chen, Linqing, et al.
Publicado: (2025)
por: Chen, Linqing, et al.
Publicado: (2025)
RAGCap-Bench: Benchmarking Capabilities of LLMs in Agentic Retrieval Augmented Generation Systems
por: Lin, Jingru, et al.
Publicado: (2025)
por: Lin, Jingru, et al.
Publicado: (2025)
LLMs in Biomedicine: A study on clinical Named Entity Recognition
por: Monajatipoor, Masoud, et al.
Publicado: (2024)
por: Monajatipoor, Masoud, et al.
Publicado: (2024)
Does RAG Introduce Unfairness in LLMs? Evaluating Fairness in Retrieval-Augmented Generation Systems
por: Wu, Xuyang, et al.
Publicado: (2024)
por: Wu, Xuyang, et al.
Publicado: (2024)
SciRerankBench: Benchmarking Rerankers Towards Scientific Retrieval-Augmented Generated LLMs
por: Chen, Haotian, et al.
Publicado: (2025)
por: Chen, Haotian, et al.
Publicado: (2025)
When Do LLMs Need Retrieval Augmentation? Mitigating LLMs' Overconfidence Helps Retrieval Augmentation
por: Ni, Shiyu, et al.
Publicado: (2024)
por: Ni, Shiyu, et al.
Publicado: (2024)
RARE: Retrieval-Augmented Reasoning Modeling
por: Wang, Zhengren, et al.
Publicado: (2025)
por: Wang, Zhengren, et al.
Publicado: (2025)
Towards Automated Smart Contract Generation: Evaluation, Benchmarking, and Retrieval-Augmented Repair
por: Chen, Zaoyu, et al.
Publicado: (2025)
por: Chen, Zaoyu, et al.
Publicado: (2025)
Persona-Augmented Benchmarking: Evaluating LLMs Across Diverse Writing Styles
por: Truong, Kimberly Le, et al.
Publicado: (2025)
por: Truong, Kimberly Le, et al.
Publicado: (2025)
DAHL: Domain-specific Automated Hallucination Evaluation of Long-Form Text through a Benchmark Dataset in Biomedicine
por: Seo, Jean, et al.
Publicado: (2024)
por: Seo, Jean, et al.
Publicado: (2024)
Unanswerability Evaluation for Retrieval Augmented Generation
por: Peng, Xiangyu, et al.
Publicado: (2024)
por: Peng, Xiangyu, et al.
Publicado: (2024)
Benchmarking Retrieval-Augmented Generation for Chemistry
por: Zhong, Xianrui, et al.
Publicado: (2025)
por: Zhong, Xianrui, et al.
Publicado: (2025)
RARE: Retrieval-Aware Robustness Evaluation for Retrieval-Augmented Generation Systems
por: Zeng, Yixiao, et al.
Publicado: (2025)
por: Zeng, Yixiao, et al.
Publicado: (2025)
AgenticOCR: Parsing Only What You Need for Efficient Retrieval-Augmented Generation
por: Wang, Zhengren, et al.
Publicado: (2026)
por: Wang, Zhengren, et al.
Publicado: (2026)
MIRAGE: A Metric-Intensive Benchmark for Retrieval-Augmented Generation Evaluation
por: Park, Chanhee, et al.
Publicado: (2025)
por: Park, Chanhee, et al.
Publicado: (2025)
CFVBench: A Comprehensive Video Benchmark for Fine-grained Multimodal Retrieval-Augmented Generation
por: Wei, Kaiwen, et al.
Publicado: (2025)
por: Wei, Kaiwen, et al.
Publicado: (2025)
Med-R$^3$: Enhancing Medical Retrieval-Augmented Reasoning of LLMs via Progressive Reinforcement Learning
por: Lu, Keer, et al.
Publicado: (2025)
por: Lu, Keer, et al.
Publicado: (2025)
Learning to Shop Like Humans: A Review-driven Retrieval-Augmented Recommendation Framework with LLMs
por: Wei, Kaiwen, et al.
Publicado: (2025)
por: Wei, Kaiwen, et al.
Publicado: (2025)
Evaluating Cultural Knowledge Processing in Large Language Models: A Cognitive Benchmarking Framework Integrating Retrieval-Augmented Generation
por: Lee, Hung-Shin, et al.
Publicado: (2025)
por: Lee, Hung-Shin, et al.
Publicado: (2025)
A Survey for Large Language Models in Biomedicine
por: Wang, Chong, et al.
Publicado: (2024)
por: Wang, Chong, et al.
Publicado: (2024)
CRUD-RAG: A Comprehensive Chinese Benchmark for Retrieval-Augmented Generation of Large Language Models
por: Lyu, Yuanjie, et al.
Publicado: (2024)
por: Lyu, Yuanjie, et al.
Publicado: (2024)
DomainRAG: A Chinese Benchmark for Evaluating Domain-specific Retrieval-Augmented Generation
por: Wang, Shuting, et al.
Publicado: (2024)
por: Wang, Shuting, et al.
Publicado: (2024)
Lifelong Knowledge Editing for LLMs with Retrieval-Augmented Continuous Prompt Learning
por: Chen, Qizhou, et al.
Publicado: (2024)
por: Chen, Qizhou, et al.
Publicado: (2024)
PRGB Benchmark: A Robust Placeholder-Assisted Algorithm for Benchmarking Retrieval-Augmented Generation
por: Tan, Zhehao, et al.
Publicado: (2025)
por: Tan, Zhehao, et al.
Publicado: (2025)
REX-RAG: Reasoning Exploration with Policy Correction in Retrieval-Augmented Generation
por: Jiang, Wentao, et al.
Publicado: (2025)
por: Jiang, Wentao, et al.
Publicado: (2025)
Compound-QA: A Benchmark for Evaluating LLMs on Compound Questions
por: Hou, Yutao, et al.
Publicado: (2024)
por: Hou, Yutao, et al.
Publicado: (2024)
Predict the Retrieval! Test time adaptation for Retrieval Augmented Generation
por: Sun, Xin, et al.
Publicado: (2026)
por: Sun, Xin, et al.
Publicado: (2026)
Dialogue Benchmark Generation from Knowledge Graphs with Cost-Effective Retrieval-Augmented LLMs
por: Omar, Reham, et al.
Publicado: (2025)
por: Omar, Reham, et al.
Publicado: (2025)
BERGEN: A Benchmarking Library for Retrieval-Augmented Generation
por: Rau, David, et al.
Publicado: (2024)
por: Rau, David, et al.
Publicado: (2024)
AI for Biomedicine in the Era of Large Language Models
por: Bi, Zhenyu, et al.
Publicado: (2024)
por: Bi, Zhenyu, et al.
Publicado: (2024)
AgenticRAGTracer: A Hop-Aware Benchmark for Diagnosing Multi-Step Retrieval Reasoning in Agentic RAG
por: You, Qijie, et al.
Publicado: (2026)
por: You, Qijie, et al.
Publicado: (2026)
YpathRAG:A Retrieval-Augmented Generation Framework and Benchmark for Pathology
por: Yu, Deshui, et al.
Publicado: (2025)
por: Yu, Deshui, et al.
Publicado: (2025)
UltraMedical: Building Specialized Generalists in Biomedicine
por: Zhang, Kaiyan, et al.
Publicado: (2024)
por: Zhang, Kaiyan, et al.
Publicado: (2024)
Ejemplares similares
-
Semantic Bridge: Universal Multi-Hop Question Generation via AMR-Driven Graph Synthesis
por: Chen, Linqing, et al.
Publicado: (2025) -
Hierarchical Multi-Label Generation with Probabilistic Level-Constraint
por: Chen, Linqing, et al.
Publicado: (2025) -
Large Language Model's Multi-Capability Alignment in Biomedical Domain
por: Wu, Wentao, et al.
Publicado: (2025) -
BiomedRAG: A Retrieval Augmented Large Language Model for Biomedicine
por: Li, Mingchen, et al.
Publicado: (2024) -
Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs
por: Son, Guijin, et al.
Publicado: (2026)