OpenScholar: Synthesizing Scientific Literature with Retrieval-augmented LMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Asai, Akari, He, Jacqueline, Shao, Rulin, Shi, Weijia, Singh, Amanpreet, Chang, Joseph Chee, Lo, Kyle, Soldaini, Luca, Feldman, Sergey, D'arcy, Mike, Wadden, David, Latzke, Matt, Tian, Minyang, Ji, Pan, Liu, Shengyan, Tong, Hao, Wu, Bohao, Xiong, Yanyu, Zettlemoyer, Luke, Neubig, Graham, Weld, Dan, Downey, Doug, Yih, Wen-tau, Koh, Pang Wei, Hajishirzi, Hannaneh
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909398521085952
author Asai, Akari
He, Jacqueline
Shao, Rulin
Shi, Weijia
Singh, Amanpreet
Chang, Joseph Chee
Lo, Kyle
Soldaini, Luca
Feldman, Sergey
D'arcy, Mike
Wadden, David
Latzke, Matt
Tian, Minyang
Ji, Pan
Liu, Shengyan
Tong, Hao
Wu, Bohao
Xiong, Yanyu
Zettlemoyer, Luke
Neubig, Graham
Weld, Dan
Downey, Doug
Yih, Wen-tau
Koh, Pang Wei
Hajishirzi, Hannaneh
author_facet Asai, Akari
He, Jacqueline
Shao, Rulin
Shi, Weijia
Singh, Amanpreet
Chang, Joseph Chee
Lo, Kyle
Soldaini, Luca
Feldman, Sergey
D'arcy, Mike
Wadden, David
Latzke, Matt
Tian, Minyang
Ji, Pan
Liu, Shengyan
Tong, Hao
Wu, Bohao
Xiong, Yanyu
Zettlemoyer, Luke
Neubig, Graham
Weld, Dan
Downey, Doug
Yih, Wen-tau
Koh, Pang Wei
Hajishirzi, Hannaneh
contents Scientific progress depends on researchers' ability to synthesize the growing body of literature. Can large language models (LMs) assist scientists in this task? We introduce OpenScholar, a specialized retrieval-augmented LM that answers scientific queries by identifying relevant passages from 45 million open-access papers and synthesizing citation-backed responses. To evaluate OpenScholar, we develop ScholarQABench, the first large-scale multi-domain benchmark for literature search, comprising 2,967 expert-written queries and 208 long-form answers across computer science, physics, neuroscience, and biomedicine. On ScholarQABench, OpenScholar-8B outperforms GPT-4o by 5% and PaperQA2 by 7% in correctness, despite being a smaller, open model. While GPT4o hallucinates citations 78 to 90% of the time, OpenScholar achieves citation accuracy on par with human experts. OpenScholar's datastore, retriever, and self-feedback inference loop also improves off-the-shelf LMs: for instance, OpenScholar-GPT4o improves GPT-4o's correctness by 12%. In human evaluations, experts preferred OpenScholar-8B and OpenScholar-GPT4o responses over expert-written ones 51% and 70% of the time, respectively, compared to GPT4o's 32%. We open-source all of our code, models, datastore, data and a public demo.
format Preprint
id arxiv_https___arxiv_org_abs_2411_14199
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle OpenScholar: Synthesizing Scientific Literature with Retrieval-augmented LMs
Asai, Akari
He, Jacqueline
Shao, Rulin
Shi, Weijia
Singh, Amanpreet
Chang, Joseph Chee
Lo, Kyle
Soldaini, Luca
Feldman, Sergey
D'arcy, Mike
Wadden, David
Latzke, Matt
Tian, Minyang
Ji, Pan
Liu, Shengyan
Tong, Hao
Wu, Bohao
Xiong, Yanyu
Zettlemoyer, Luke
Neubig, Graham
Weld, Dan
Downey, Doug
Yih, Wen-tau
Koh, Pang Wei
Hajishirzi, Hannaneh
Computation and Language
Artificial Intelligence
Digital Libraries
Information Retrieval
Machine Learning
Scientific progress depends on researchers' ability to synthesize the growing body of literature. Can large language models (LMs) assist scientists in this task? We introduce OpenScholar, a specialized retrieval-augmented LM that answers scientific queries by identifying relevant passages from 45 million open-access papers and synthesizing citation-backed responses. To evaluate OpenScholar, we develop ScholarQABench, the first large-scale multi-domain benchmark for literature search, comprising 2,967 expert-written queries and 208 long-form answers across computer science, physics, neuroscience, and biomedicine. On ScholarQABench, OpenScholar-8B outperforms GPT-4o by 5% and PaperQA2 by 7% in correctness, despite being a smaller, open model. While GPT4o hallucinates citations 78 to 90% of the time, OpenScholar achieves citation accuracy on par with human experts. OpenScholar's datastore, retriever, and self-feedback inference loop also improves off-the-shelf LMs: for instance, OpenScholar-GPT4o improves GPT-4o's correctness by 12%. In human evaluations, experts preferred OpenScholar-8B and OpenScholar-GPT4o responses over expert-written ones 51% and 70% of the time, respectively, compared to GPT4o's 32%. We open-source all of our code, models, datastore, data and a public demo.
title OpenScholar: Synthesizing Scientific Literature with Retrieval-augmented LMs
topic Computation and Language
Artificial Intelligence
Digital Libraries
Information Retrieval
Machine Learning
url https://arxiv.org/abs/2411.14199