OpenScholar: Synthesizing Scientific Literature with Retrieval-augmented LMs
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909398521085952 |
|---|---|
| author | Asai, Akari He, Jacqueline Shao, Rulin Shi, Weijia Singh, Amanpreet Chang, Joseph Chee Lo, Kyle Soldaini, Luca Feldman, Sergey D'arcy, Mike Wadden, David Latzke, Matt Tian, Minyang Ji, Pan Liu, Shengyan Tong, Hao Wu, Bohao Xiong, Yanyu Zettlemoyer, Luke Neubig, Graham Weld, Dan Downey, Doug Yih, Wen-tau Koh, Pang Wei Hajishirzi, Hannaneh |
| author_facet | Asai, Akari He, Jacqueline Shao, Rulin Shi, Weijia Singh, Amanpreet Chang, Joseph Chee Lo, Kyle Soldaini, Luca Feldman, Sergey D'arcy, Mike Wadden, David Latzke, Matt Tian, Minyang Ji, Pan Liu, Shengyan Tong, Hao Wu, Bohao Xiong, Yanyu Zettlemoyer, Luke Neubig, Graham Weld, Dan Downey, Doug Yih, Wen-tau Koh, Pang Wei Hajishirzi, Hannaneh |
| contents | Scientific progress depends on researchers' ability to synthesize the growing body of literature. Can large language models (LMs) assist scientists in this task? We introduce OpenScholar, a specialized retrieval-augmented LM that answers scientific queries by identifying relevant passages from 45 million open-access papers and synthesizing citation-backed responses. To evaluate OpenScholar, we develop ScholarQABench, the first large-scale multi-domain benchmark for literature search, comprising 2,967 expert-written queries and 208 long-form answers across computer science, physics, neuroscience, and biomedicine. On ScholarQABench, OpenScholar-8B outperforms GPT-4o by 5% and PaperQA2 by 7% in correctness, despite being a smaller, open model. While GPT4o hallucinates citations 78 to 90% of the time, OpenScholar achieves citation accuracy on par with human experts. OpenScholar's datastore, retriever, and self-feedback inference loop also improves off-the-shelf LMs: for instance, OpenScholar-GPT4o improves GPT-4o's correctness by 12%. In human evaluations, experts preferred OpenScholar-8B and OpenScholar-GPT4o responses over expert-written ones 51% and 70% of the time, respectively, compared to GPT4o's 32%. We open-source all of our code, models, datastore, data and a public demo. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_14199 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | OpenScholar: Synthesizing Scientific Literature with Retrieval-augmented LMs Asai, Akari He, Jacqueline Shao, Rulin Shi, Weijia Singh, Amanpreet Chang, Joseph Chee Lo, Kyle Soldaini, Luca Feldman, Sergey D'arcy, Mike Wadden, David Latzke, Matt Tian, Minyang Ji, Pan Liu, Shengyan Tong, Hao Wu, Bohao Xiong, Yanyu Zettlemoyer, Luke Neubig, Graham Weld, Dan Downey, Doug Yih, Wen-tau Koh, Pang Wei Hajishirzi, Hannaneh Computation and Language Artificial Intelligence Digital Libraries Information Retrieval Machine Learning Scientific progress depends on researchers' ability to synthesize the growing body of literature. Can large language models (LMs) assist scientists in this task? We introduce OpenScholar, a specialized retrieval-augmented LM that answers scientific queries by identifying relevant passages from 45 million open-access papers and synthesizing citation-backed responses. To evaluate OpenScholar, we develop ScholarQABench, the first large-scale multi-domain benchmark for literature search, comprising 2,967 expert-written queries and 208 long-form answers across computer science, physics, neuroscience, and biomedicine. On ScholarQABench, OpenScholar-8B outperforms GPT-4o by 5% and PaperQA2 by 7% in correctness, despite being a smaller, open model. While GPT4o hallucinates citations 78 to 90% of the time, OpenScholar achieves citation accuracy on par with human experts. OpenScholar's datastore, retriever, and self-feedback inference loop also improves off-the-shelf LMs: for instance, OpenScholar-GPT4o improves GPT-4o's correctness by 12%. In human evaluations, experts preferred OpenScholar-8B and OpenScholar-GPT4o responses over expert-written ones 51% and 70% of the time, respectively, compared to GPT4o's 32%. We open-source all of our code, models, datastore, data and a public demo. |
| title | OpenScholar: Synthesizing Scientific Literature with Retrieval-augmented LMs |
| topic | Computation and Language Artificial Intelligence Digital Libraries Information Retrieval Machine Learning |
| url | https://arxiv.org/abs/2411.14199 |