CMULAB: An Open-Source Framework for Training and Deployment of Natural Language Processing Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sheikh, Zaid, Anastasopoulos, Antonios, Rijhwani, Shruti, Tjuatja, Lindia, Jimerson, Robbie, Neubig, Graham
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929301874540544
author Sheikh, Zaid
Anastasopoulos, Antonios
Rijhwani, Shruti
Tjuatja, Lindia
Jimerson, Robbie
Neubig, Graham
author_facet Sheikh, Zaid
Anastasopoulos, Antonios
Rijhwani, Shruti
Tjuatja, Lindia
Jimerson, Robbie
Neubig, Graham
contents Effectively using Natural Language Processing (NLP) tools in under-resourced languages requires a thorough understanding of the language itself, familiarity with the latest models and training methodologies, and technical expertise to deploy these models. This could present a significant obstacle for language community members and linguists to use NLP tools. This paper introduces the CMU Linguistic Annotation Backend, an open-source framework that simplifies model deployment and continuous human-in-the-loop fine-tuning of NLP models. CMULAB enables users to leverage the power of multilingual models to quickly adapt and extend existing tools for speech recognition, OCR, translation, and syntactic analysis to new languages, even with limited training data. We describe various tools and APIs that are currently available and how developers can easily add new models/functionality to the framework. Code is available at https://github.com/neulab/cmulab along with a live demo at https://cmulab.dev
format Preprint
id arxiv_https___arxiv_org_abs_2404_02408
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CMULAB: An Open-Source Framework for Training and Deployment of Natural Language Processing Models
Sheikh, Zaid
Anastasopoulos, Antonios
Rijhwani, Shruti
Tjuatja, Lindia
Jimerson, Robbie
Neubig, Graham
Computation and Language
Effectively using Natural Language Processing (NLP) tools in under-resourced languages requires a thorough understanding of the language itself, familiarity with the latest models and training methodologies, and technical expertise to deploy these models. This could present a significant obstacle for language community members and linguists to use NLP tools. This paper introduces the CMU Linguistic Annotation Backend, an open-source framework that simplifies model deployment and continuous human-in-the-loop fine-tuning of NLP models. CMULAB enables users to leverage the power of multilingual models to quickly adapt and extend existing tools for speech recognition, OCR, translation, and syntactic analysis to new languages, even with limited training data. We describe various tools and APIs that are currently available and how developers can easily add new models/functionality to the framework. Code is available at https://github.com/neulab/cmulab along with a live demo at https://cmulab.dev
title CMULAB: An Open-Source Framework for Training and Deployment of Natural Language Processing Models
topic Computation and Language
url https://arxiv.org/abs/2404.02408