Fine-Tuning Language Models on Multiple Datasets for Citation Intention Classification

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Shui, Zeren, Karypis, Petros, Karls, Daniel S., Wen, Mingjian, Manchanda, Saurav, Tadmor, Ellad B., Karypis, George
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916442758184960
author Shui, Zeren
Karypis, Petros
Karls, Daniel S.
Wen, Mingjian
Manchanda, Saurav
Tadmor, Ellad B.
Karypis, George
author_facet Shui, Zeren
Karypis, Petros
Karls, Daniel S.
Wen, Mingjian
Manchanda, Saurav
Tadmor, Ellad B.
Karypis, George
contents Citation intention Classification (CIC) tools classify citations by their intention (e.g., background, motivation) and assist readers in evaluating the contribution of scientific literature. Prior research has shown that pretrained language models (PLMs) such as SciBERT can achieve state-of-the-art performance on CIC benchmarks. PLMs are trained via self-supervision tasks on a large corpus of general text and can quickly adapt to CIC tasks via moderate fine-tuning on the corresponding dataset. Despite their advantages, PLMs can easily overfit small datasets during fine-tuning. In this paper, we propose a multi-task learning (MTL) framework that jointly fine-tunes PLMs on a dataset of primary interest together with multiple auxiliary CIC datasets to take advantage of additional supervision signals. We develop a data-driven task relation learning (TRL) method that controls the contribution of auxiliary datasets to avoid negative transfer and expensive hyper-parameter tuning. We conduct experiments on three CIC datasets and show that fine-tuning with additional datasets can improve the PLMs' generalization performance on the primary dataset. PLMs fine-tuned with our proposed framework outperform the current state-of-the-art models by 7% to 11% on small datasets while aligning with the best-performing model on a large dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2410_13332
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Fine-Tuning Language Models on Multiple Datasets for Citation Intention Classification
Shui, Zeren
Karypis, Petros
Karls, Daniel S.
Wen, Mingjian
Manchanda, Saurav
Tadmor, Ellad B.
Karypis, George
Computation and Language
Citation intention Classification (CIC) tools classify citations by their intention (e.g., background, motivation) and assist readers in evaluating the contribution of scientific literature. Prior research has shown that pretrained language models (PLMs) such as SciBERT can achieve state-of-the-art performance on CIC benchmarks. PLMs are trained via self-supervision tasks on a large corpus of general text and can quickly adapt to CIC tasks via moderate fine-tuning on the corresponding dataset. Despite their advantages, PLMs can easily overfit small datasets during fine-tuning. In this paper, we propose a multi-task learning (MTL) framework that jointly fine-tunes PLMs on a dataset of primary interest together with multiple auxiliary CIC datasets to take advantage of additional supervision signals. We develop a data-driven task relation learning (TRL) method that controls the contribution of auxiliary datasets to avoid negative transfer and expensive hyper-parameter tuning. We conduct experiments on three CIC datasets and show that fine-tuning with additional datasets can improve the PLMs' generalization performance on the primary dataset. PLMs fine-tuned with our proposed framework outperform the current state-of-the-art models by 7% to 11% on small datasets while aligning with the best-performing model on a large dataset.
title Fine-Tuning Language Models on Multiple Datasets for Citation Intention Classification
topic Computation and Language
url https://arxiv.org/abs/2410.13332