SC4ANM: Identifying Optimal Section Combinations for Automated Novelty Prediction in Academic Papers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Wenqing, Zhang, Chengzhi, Bao, Tong, Zhao, Yi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915297919762432
author Wu, Wenqing
Zhang, Chengzhi
Bao, Tong
Zhao, Yi
author_facet Wu, Wenqing
Zhang, Chengzhi
Bao, Tong
Zhao, Yi
contents Novelty is a core component of academic papers, and there are multiple perspectives on the assessment of novelty. Existing methods often focus on word or entity combinations, which provide limited insights. The content related to a paper's novelty is typically distributed across different core sections, e.g., Introduction, Methodology and Results. Therefore, exploring the optimal combination of sections for evaluating the novelty of a paper is important for advancing automated novelty assessment. In this paper, we utilize different combinations of sections from academic papers as inputs to drive language models to predict novelty scores. We then analyze the results to determine the optimal section combinations for novelty score prediction. We first employ natural language processing techniques to identify the sectional structure of academic papers, categorizing them into introduction, methods, results, and discussion (IMRaD). Subsequently, we used different combinations of these sections (e.g., introduction and methods) as inputs for pretrained language models (PLMs) and large language models (LLMs), employing novelty scores provided by human expert reviewers as ground truth labels to obtain prediction results. The results indicate that using introduction, results and discussion is most appropriate for assessing the novelty of a paper, while the use of the entire text does not yield significant results. Furthermore, based on the results of the PLMs and LLMs, the introduction and results appear to be the most important section for the task of novelty score prediction. The code and dataset for this paper can be accessed at https://github.com/njust-winchy/SC4ANM.
format Preprint
id arxiv_https___arxiv_org_abs_2505_16330
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SC4ANM: Identifying Optimal Section Combinations for Automated Novelty Prediction in Academic Papers
Wu, Wenqing
Zhang, Chengzhi
Bao, Tong
Zhao, Yi
Computation and Language
Artificial Intelligence
Digital Libraries
Novelty is a core component of academic papers, and there are multiple perspectives on the assessment of novelty. Existing methods often focus on word or entity combinations, which provide limited insights. The content related to a paper's novelty is typically distributed across different core sections, e.g., Introduction, Methodology and Results. Therefore, exploring the optimal combination of sections for evaluating the novelty of a paper is important for advancing automated novelty assessment. In this paper, we utilize different combinations of sections from academic papers as inputs to drive language models to predict novelty scores. We then analyze the results to determine the optimal section combinations for novelty score prediction. We first employ natural language processing techniques to identify the sectional structure of academic papers, categorizing them into introduction, methods, results, and discussion (IMRaD). Subsequently, we used different combinations of these sections (e.g., introduction and methods) as inputs for pretrained language models (PLMs) and large language models (LLMs), employing novelty scores provided by human expert reviewers as ground truth labels to obtain prediction results. The results indicate that using introduction, results and discussion is most appropriate for assessing the novelty of a paper, while the use of the entire text does not yield significant results. Furthermore, based on the results of the PLMs and LLMs, the introduction and results appear to be the most important section for the task of novelty score prediction. The code and dataset for this paper can be accessed at https://github.com/njust-winchy/SC4ANM.
title SC4ANM: Identifying Optimal Section Combinations for Automated Novelty Prediction in Academic Papers
topic Computation and Language
Artificial Intelligence
Digital Libraries
url https://arxiv.org/abs/2505.16330