CMULAB: An Open-Source Framework for Training and Deployment of Natural Language Processing Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866929301874540544 |
|---|---|
| author | Sheikh, Zaid Anastasopoulos, Antonios Rijhwani, Shruti Tjuatja, Lindia Jimerson, Robbie Neubig, Graham |
| author_facet | Sheikh, Zaid Anastasopoulos, Antonios Rijhwani, Shruti Tjuatja, Lindia Jimerson, Robbie Neubig, Graham |
| contents | Effectively using Natural Language Processing (NLP) tools in under-resourced languages requires a thorough understanding of the language itself, familiarity with the latest models and training methodologies, and technical expertise to deploy these models. This could present a significant obstacle for language community members and linguists to use NLP tools. This paper introduces the CMU Linguistic Annotation Backend, an open-source framework that simplifies model deployment and continuous human-in-the-loop fine-tuning of NLP models. CMULAB enables users to leverage the power of multilingual models to quickly adapt and extend existing tools for speech recognition, OCR, translation, and syntactic analysis to new languages, even with limited training data. We describe various tools and APIs that are currently available and how developers can easily add new models/functionality to the framework. Code is available at https://github.com/neulab/cmulab along with a live demo at https://cmulab.dev |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2404_02408 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | CMULAB: An Open-Source Framework for Training and Deployment of Natural Language Processing Models Sheikh, Zaid Anastasopoulos, Antonios Rijhwani, Shruti Tjuatja, Lindia Jimerson, Robbie Neubig, Graham Computation and Language Effectively using Natural Language Processing (NLP) tools in under-resourced languages requires a thorough understanding of the language itself, familiarity with the latest models and training methodologies, and technical expertise to deploy these models. This could present a significant obstacle for language community members and linguists to use NLP tools. This paper introduces the CMU Linguistic Annotation Backend, an open-source framework that simplifies model deployment and continuous human-in-the-loop fine-tuning of NLP models. CMULAB enables users to leverage the power of multilingual models to quickly adapt and extend existing tools for speech recognition, OCR, translation, and syntactic analysis to new languages, even with limited training data. We describe various tools and APIs that are currently available and how developers can easily add new models/functionality to the framework. Code is available at https://github.com/neulab/cmulab along with a live demo at https://cmulab.dev |
| title | CMULAB: An Open-Source Framework for Training and Deployment of Natural Language Processing Models |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2404.02408 |