Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Shankar, Lavanya, Perera, Leibny Paola Garcia
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2508.09430
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909735674970112
author Shankar, Lavanya
Perera, Leibny Paola Garcia
author_facet Shankar, Lavanya
Perera, Leibny Paola Garcia
contents Code-switching and language identification in child-directed scenarios present significant challenges, particularly in bilingual environments. This paper addresses this challenge by using Zipformer to handle the nuances of speech, which contains two imbalanced languages, Mandarin and English, in an utterance. This work demonstrates that the internal layers of the Zipformer effectively encode the language characteristics, which can be leveraged in language identification. We present the selection methodology of the inner layers to extract the embeddings and make a comparison with different back-ends. Our analysis shows that Zipformer is robust across these backends. Our approach effectively handles imbalanced data, achieving a Balanced Accuracy (BAC) of 81.89%, a 15.47% improvement over the language identification baseline. These findings highlight the potential of the transformer encoder architecture model in real scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2508_09430
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Leveraging Zipformer Model for Effective Language Identification in Code-Switched Child-Directed Speech
Shankar, Lavanya
Perera, Leibny Paola Garcia
Computation and Language
Sound
Code-switching and language identification in child-directed scenarios present significant challenges, particularly in bilingual environments. This paper addresses this challenge by using Zipformer to handle the nuances of speech, which contains two imbalanced languages, Mandarin and English, in an utterance. This work demonstrates that the internal layers of the Zipformer effectively encode the language characteristics, which can be leveraged in language identification. We present the selection methodology of the inner layers to extract the embeddings and make a comparison with different back-ends. Our analysis shows that Zipformer is robust across these backends. Our approach effectively handles imbalanced data, achieving a Balanced Accuracy (BAC) of 81.89%, a 15.47% improvement over the language identification baseline. These findings highlight the potential of the transformer encoder architecture model in real scenarios.
title Leveraging Zipformer Model for Effective Language Identification in Code-Switched Child-Directed Speech
topic Computation and Language
Sound
url https://arxiv.org/abs/2508.09430