Everyday Speech in the Indian Subcontinent
Fuente:
arXiv
Saved in:
| Main Author: | |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917932219498496 |
|---|---|
| author | P, Utkarsh |
| author_facet | P, Utkarsh |
| contents | India has 1369 languages of which 22 are official. About 13 different scripts are used to represent these languages. A Common Label Set (CLS) was developed based on phonetics to address the issue of large vocabulary of units required in the End-to-End (E2E) framework for multilingual synthesis. The Indian language text is first converted to CLS. This approach enables seamless code switching across 13 Indian languages and English in a given native speaker's voice, which corresponds to everyday speech in the Indian subcontinent, where the population is multilingual. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_10508 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Everyday Speech in the Indian Subcontinent P, Utkarsh Computation and Language Sound Audio and Speech Processing I.2.7 India has 1369 languages of which 22 are official. About 13 different scripts are used to represent these languages. A Common Label Set (CLS) was developed based on phonetics to address the issue of large vocabulary of units required in the End-to-End (E2E) framework for multilingual synthesis. The Indian language text is first converted to CLS. This approach enables seamless code switching across 13 Indian languages and English in a given native speaker's voice, which corresponds to everyday speech in the Indian subcontinent, where the population is multilingual. |
| title | Everyday Speech in the Indian Subcontinent |
| topic | Computation and Language Sound Audio and Speech Processing I.2.7 |
| url | https://arxiv.org/abs/2410.10508 |