Staff View: :: Library Catalog

Saved in:

Bibliographic Details
Main Authors:	D'Cruz, Célia, Bereder, Jean-Marc, Precioso, Frédéric, Riveill, Michel
Format:	Preprint
Published:	2024
Subjects:	Computation and Language
Online Access:	https://arxiv.org/abs/2408.13253
Tags:	Add Tag No Tags, Be the first to tag this record!

_version_	1866917756703604736
author	D'Cruz, Célia Bereder, Jean-Marc Precioso, Frédéric Riveill, Michel
author_facet	D'Cruz, Célia Bereder, Jean-Marc Precioso, Frédéric Riveill, Michel
contents	Large Language Models have undoubtedly revolutionized the Natural Language Processing field, the current trend being to promote one-model-for-all tasks (sentiment analysis, translation, etc.). However, the statistical mechanisms at work in the larger language models struggle to exploit the relevant information when it is very sparse, when it is a weak signal. This is the case, for example, for the classification of long domain-specific documents, when the relevance relies on a single relevant word or on very few relevant words from technical jargon. In the medical domain, it is essential to determine whether a given report contains critical information about a patient's condition. This critical information is often based on one or few specific isolated terms. In this paper, we propose a hierarchical model which exploits a short list of potential target terms to retrieve candidate sentences and represent them into the contextualized embedding of the target term(s) they contain. A pooling of the term(s) embedding(s) entails the document representation to be classified. We evaluate our model on one public medical document benchmark in English and on one private French medical dataset. We show that our narrower hierarchical model is better than larger language models for retrieving relevant long documents in a domain-specific context.
format	Preprint
id	arxiv_https___arxiv_org_abs_2408_13253
institution	arXiv
publishDate	2024
record_format	arxiv
spellingShingle	Domain-specific long text classification from sparse relevant information D'Cruz, Célia Bereder, Jean-Marc Precioso, Frédéric Riveill, Michel Computation and Language Large Language Models have undoubtedly revolutionized the Natural Language Processing field, the current trend being to promote one-model-for-all tasks (sentiment analysis, translation, etc.). However, the statistical mechanisms at work in the larger language models struggle to exploit the relevant information when it is very sparse, when it is a weak signal. This is the case, for example, for the classification of long domain-specific documents, when the relevance relies on a single relevant word or on very few relevant words from technical jargon. In the medical domain, it is essential to determine whether a given report contains critical information about a patient's condition. This critical information is often based on one or few specific isolated terms. In this paper, we propose a hierarchical model which exploits a short list of potential target terms to retrieve candidate sentences and represent them into the contextualized embedding of the target term(s) they contain. A pooling of the term(s) embedding(s) entails the document representation to be classified. We evaluate our model on one public medical document benchmark in English and on one private French medical dataset. We show that our narrower hierarchical model is better than larger language models for retrieving relevant long documents in a domain-specific context.
title	Domain-specific long text classification from sparse relevant information
topic	Computation and Language
url	https://arxiv.org/abs/2408.13253

Similar Items