Language Models as Science Tutors

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chevalier, Alexis, Geng, Jiayi, Wettig, Alexander, Chen, Howard, Mizera, Sebastian, Annala, Toni, Aragon, Max Jameson, Fanlo, Arturo Rodríguez, Frieder, Simon, Machado, Simon, Prabhakar, Akshara, Thieu, Ellie, Wang, Jiachen T., Wang, Zirui, Wu, Xindi, Xia, Mengzhou, Xia, Wenhan, Yu, Jiatong, Zhu, Jun-Jie, Ren, Zhiyong Jason, Arora, Sanjeev, Chen, Danqi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914878690689024
author Chevalier, Alexis
Geng, Jiayi
Wettig, Alexander
Chen, Howard
Mizera, Sebastian
Annala, Toni
Aragon, Max Jameson
Fanlo, Arturo Rodríguez
Frieder, Simon
Machado, Simon
Prabhakar, Akshara
Thieu, Ellie
Wang, Jiachen T.
Wang, Zirui
Wu, Xindi
Xia, Mengzhou
Xia, Wenhan
Yu, Jiatong
Zhu, Jun-Jie
Ren, Zhiyong Jason
Arora, Sanjeev
Chen, Danqi
author_facet Chevalier, Alexis
Geng, Jiayi
Wettig, Alexander
Chen, Howard
Mizera, Sebastian
Annala, Toni
Aragon, Max Jameson
Fanlo, Arturo Rodríguez
Frieder, Simon
Machado, Simon
Prabhakar, Akshara
Thieu, Ellie
Wang, Jiachen T.
Wang, Zirui
Wu, Xindi
Xia, Mengzhou
Xia, Wenhan
Yu, Jiatong
Zhu, Jun-Jie
Ren, Zhiyong Jason
Arora, Sanjeev
Chen, Danqi
contents NLP has recently made exciting progress toward training language models (LMs) with strong scientific problem-solving skills. However, model development has not focused on real-life use-cases of LMs for science, including applications in education that require processing long scientific documents. To address this, we introduce TutorEval and TutorChat. TutorEval is a diverse question-answering benchmark consisting of questions about long chapters from STEM textbooks, written by experts. TutorEval helps measure real-life usability of LMs as scientific assistants, and it is the first benchmark combining long contexts, free-form generation, and multi-disciplinary scientific knowledge. Moreover, we show that fine-tuning base models with existing dialogue datasets leads to poor performance on TutorEval. Therefore, we create TutorChat, a dataset of 80,000 long synthetic dialogues about textbooks. We use TutorChat to fine-tune Llemma models with 7B and 34B parameters. These LM tutors specialized in math have a 32K-token context window, and they excel at TutorEval while performing strongly on GSM8K and MATH. Our datasets build on open-source materials, and we release our models, data, and evaluations.
format Preprint
id arxiv_https___arxiv_org_abs_2402_11111
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Language Models as Science Tutors
Chevalier, Alexis
Geng, Jiayi
Wettig, Alexander
Chen, Howard
Mizera, Sebastian
Annala, Toni
Aragon, Max Jameson
Fanlo, Arturo Rodríguez
Frieder, Simon
Machado, Simon
Prabhakar, Akshara
Thieu, Ellie
Wang, Jiachen T.
Wang, Zirui
Wu, Xindi
Xia, Mengzhou
Xia, Wenhan
Yu, Jiatong
Zhu, Jun-Jie
Ren, Zhiyong Jason
Arora, Sanjeev
Chen, Danqi
Computation and Language
NLP has recently made exciting progress toward training language models (LMs) with strong scientific problem-solving skills. However, model development has not focused on real-life use-cases of LMs for science, including applications in education that require processing long scientific documents. To address this, we introduce TutorEval and TutorChat. TutorEval is a diverse question-answering benchmark consisting of questions about long chapters from STEM textbooks, written by experts. TutorEval helps measure real-life usability of LMs as scientific assistants, and it is the first benchmark combining long contexts, free-form generation, and multi-disciplinary scientific knowledge. Moreover, we show that fine-tuning base models with existing dialogue datasets leads to poor performance on TutorEval. Therefore, we create TutorChat, a dataset of 80,000 long synthetic dialogues about textbooks. We use TutorChat to fine-tune Llemma models with 7B and 34B parameters. These LM tutors specialized in math have a 32K-token context window, and they excel at TutorEval while performing strongly on GSM8K and MATH. Our datasets build on open-source materials, and we release our models, data, and evaluations.
title Language Models as Science Tutors
topic Computation and Language
url https://arxiv.org/abs/2402.11111