Utilizing Pre-trained and Large Language Models for 10-K Items Segmentation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Lu, Hsin-Min, Chien, Yu-Tai, Yen, Huan-Hsun, Chen, Yen-Hsiu
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917388714246144
author Lu, Hsin-Min
Chien, Yu-Tai
Yen, Huan-Hsun
Chen, Yen-Hsiu
author_facet Lu, Hsin-Min
Chien, Yu-Tai
Yen, Huan-Hsun
Chen, Yen-Hsiu
contents Extracting specific items from 10-K reports is challenging due to variations in document formats and item presentation. To improve over traditional rule-based approaches, this study introduces and compares two advanced item segmentation methods: (1) GPT4ItemSeg, using a novel line-ID-based prompting mechanism to utilize a large language model, ChatGPT-4o, for item segmentation, and (2) BERT4ItemSeg, combining a pre-trained language model, BERT, with a Bi-LSTM model in a hierarchical structure to overcome context window constraints. Trained and evaluated on 3,737 annotated 10-K reports, BERT4ItemSeg achieves a macro-F1 of 0.9825, surpassing GPT4ItemSeg (0.9567), conditional random field (0.9818), and rule-based methods (0.9048) for core items (1, 1A, 3, and 7). These approaches enhance item segmentation performance, improving text analytics in accounting and finance. BERT4ItemSeg offers satisfactory item segmentation performance, while GPT4ItemSeg can easily adapt to regulatory changes. Together, they provide an extensible framework for 10-K item segmentation that supports reliable and reproducible results.
format Preprint
id arxiv_https___arxiv_org_abs_2502_08875
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Utilizing Pre-trained and Large Language Models for 10-K Items Segmentation
Lu, Hsin-Min
Chien, Yu-Tai
Yen, Huan-Hsun
Chen, Yen-Hsiu
General Finance
Extracting specific items from 10-K reports is challenging due to variations in document formats and item presentation. To improve over traditional rule-based approaches, this study introduces and compares two advanced item segmentation methods: (1) GPT4ItemSeg, using a novel line-ID-based prompting mechanism to utilize a large language model, ChatGPT-4o, for item segmentation, and (2) BERT4ItemSeg, combining a pre-trained language model, BERT, with a Bi-LSTM model in a hierarchical structure to overcome context window constraints. Trained and evaluated on 3,737 annotated 10-K reports, BERT4ItemSeg achieves a macro-F1 of 0.9825, surpassing GPT4ItemSeg (0.9567), conditional random field (0.9818), and rule-based methods (0.9048) for core items (1, 1A, 3, and 7). These approaches enhance item segmentation performance, improving text analytics in accounting and finance. BERT4ItemSeg offers satisfactory item segmentation performance, while GPT4ItemSeg can easily adapt to regulatory changes. Together, they provide an extensible framework for 10-K item segmentation that supports reliable and reproducible results.
title Utilizing Pre-trained and Large Language Models for 10-K Items Segmentation
topic General Finance
url https://arxiv.org/abs/2502.08875