Saved in:
Bibliographic Details
Main Authors: Azeraf, Elie, Monfrini, Emmanuel, Vignon, Emmanuel, Pieczynski, Wojciech
Format: Preprint
Published: 2021
Subjects:
Online Access:https://arxiv.org/abs/2102.11037
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909729635172352
author Azeraf, Elie
Monfrini, Emmanuel
Vignon, Emmanuel
Pieczynski, Wojciech
author_facet Azeraf, Elie
Monfrini, Emmanuel
Vignon, Emmanuel
Pieczynski, Wojciech
contents Natural Language Processing (NLP) models' current trend consists of using increasingly more extra-data to build the best models as possible. It implies more expensive computational costs and training time, difficulties for deployment, and worries about these models' carbon footprint reveal a critical problem in the future. Against this trend, our goal is to develop NLP models requiring no extra-data and minimizing training time. To do so, in this paper, we explore Markov chain models, Hidden Markov Chain (HMC) and Pairwise Markov Chain (PMC), for NLP segmentation tasks. We apply these models for three classic applications: POS Tagging, Named-Entity-Recognition, and Chunking. We develop an original method to adapt these models for text segmentation's specific challenges to obtain relevant performances with very short training and execution times. PMC achieves equivalent results to those obtained by Conditional Random Fields (CRF), one of the most applied models for these tasks when no extra-data are used. Moreover, PMC has training times 30 times shorter than the CRF ones, which validates this model given our objectives.
format Preprint
id arxiv_https___arxiv_org_abs_2102_11037
institution arXiv
publishDate 2021
record_format arxiv
spellingShingle Highly Fast Text Segmentation With Pairwise Markov Chains
Azeraf, Elie
Monfrini, Emmanuel
Vignon, Emmanuel
Pieczynski, Wojciech
Computation and Language
Machine Learning
Natural Language Processing (NLP) models' current trend consists of using increasingly more extra-data to build the best models as possible. It implies more expensive computational costs and training time, difficulties for deployment, and worries about these models' carbon footprint reveal a critical problem in the future. Against this trend, our goal is to develop NLP models requiring no extra-data and minimizing training time. To do so, in this paper, we explore Markov chain models, Hidden Markov Chain (HMC) and Pairwise Markov Chain (PMC), for NLP segmentation tasks. We apply these models for three classic applications: POS Tagging, Named-Entity-Recognition, and Chunking. We develop an original method to adapt these models for text segmentation's specific challenges to obtain relevant performances with very short training and execution times. PMC achieves equivalent results to those obtained by Conditional Random Fields (CRF), one of the most applied models for these tasks when no extra-data are used. Moreover, PMC has training times 30 times shorter than the CRF ones, which validates this model given our objectives.
title Highly Fast Text Segmentation With Pairwise Markov Chains
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2102.11037