Measuring Risk of Bias in Biomedical Reports: The RoBBR Benchmark

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Jianyou, Cao, Weili, Bao, Longtian, Zheng, Youze, Pasternak, Gil, Wang, Kaicheng, Wang, Xiaoyue, Paturi, Ramamohan, Bergen, Leon
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918144580255744
author Wang, Jianyou
Cao, Weili
Bao, Longtian
Zheng, Youze
Pasternak, Gil
Wang, Kaicheng
Wang, Xiaoyue
Paturi, Ramamohan
Bergen, Leon
author_facet Wang, Jianyou
Cao, Weili
Bao, Longtian
Zheng, Youze
Pasternak, Gil
Wang, Kaicheng
Wang, Xiaoyue
Paturi, Ramamohan
Bergen, Leon
contents Systems that answer questions by reviewing the scientific literature are becoming increasingly feasible. To draw reliable conclusions, these systems should take into account the quality of available evidence from different studies, placing more weight on studies that use a valid methodology. We present a benchmark for measuring the methodological strength of biomedical papers, drawing on the risk-of-bias framework used for systematic reviews. Derived from over 500 biomedical studies, the three benchmark tasks encompass expert reviewers' judgments of studies' research methodologies, including the assessments of risk of bias within these studies. The benchmark contains a human-validated annotation pipeline for fine-grained alignment of reviewers' judgments with research paper sentences. Our analyses show that large language models' reasoning and retrieval capabilities impact their effectiveness with risk-of-bias assessment. The dataset is available at https://github.com/RoBBR-Benchmark/RoBBR.
format Preprint
id arxiv_https___arxiv_org_abs_2411_18831
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Measuring Risk of Bias in Biomedical Reports: The RoBBR Benchmark
Wang, Jianyou
Cao, Weili
Bao, Longtian
Zheng, Youze
Pasternak, Gil
Wang, Kaicheng
Wang, Xiaoyue
Paturi, Ramamohan
Bergen, Leon
Computation and Language
Systems that answer questions by reviewing the scientific literature are becoming increasingly feasible. To draw reliable conclusions, these systems should take into account the quality of available evidence from different studies, placing more weight on studies that use a valid methodology. We present a benchmark for measuring the methodological strength of biomedical papers, drawing on the risk-of-bias framework used for systematic reviews. Derived from over 500 biomedical studies, the three benchmark tasks encompass expert reviewers' judgments of studies' research methodologies, including the assessments of risk of bias within these studies. The benchmark contains a human-validated annotation pipeline for fine-grained alignment of reviewers' judgments with research paper sentences. Our analyses show that large language models' reasoning and retrieval capabilities impact their effectiveness with risk-of-bias assessment. The dataset is available at https://github.com/RoBBR-Benchmark/RoBBR.
title Measuring Risk of Bias in Biomedical Reports: The RoBBR Benchmark
topic Computation and Language
url https://arxiv.org/abs/2411.18831