SITA: Learning Speaker-Invariant and Tone-Aware Speech Representations for Low-Resource Tonal Languages
Fuente:
arXiv
Salvato in:
| Autori principali: | Xu, Tianyi, Ouyang, Xuan, Yao, Binwei, Xiong, Shoua, Misurelli, Sara, Lor, Maichou, Hu, Junjie |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Low-Resource Court Judgment Summarization for Common Law Systems
di: Liu, Shuaiqi, et al.
Pubblicazione: (2024)
di: Liu, Shuaiqi, et al.
Pubblicazione: (2024)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
di: Peters, Sydney, et al.
Pubblicazione: (2025)
di: Peters, Sydney, et al.
Pubblicazione: (2025)
Heidelberg-Boston @ SIGTYP 2024 Shared Task: Enhancing Low-Resource Language Analysis With Character-Aware Hierarchical Transformers
di: Riemenschneider, Frederick, et al.
Pubblicazione: (2024)
di: Riemenschneider, Frederick, et al.
Pubblicazione: (2024)
Spatially-Aware Speaker for Vision-and-Language Navigation Instruction Generation
di: Gopinathan, Muraleekrishna, et al.
Pubblicazione: (2024)
di: Gopinathan, Muraleekrishna, et al.
Pubblicazione: (2024)
Robust Long-Form Bangla Speech Processing: Automatic Speech Recognition and Speaker Diarization
di: Chowdhury, MD. Sagor, et al.
Pubblicazione: (2026)
di: Chowdhury, MD. Sagor, et al.
Pubblicazione: (2026)
PathBench: Speech Intelligibility Benchmark for Automatic Pathological Speech Assessment
di: Halpern, Bence Mark, et al.
Pubblicazione: (2026)
di: Halpern, Bence Mark, et al.
Pubblicazione: (2026)
Toward Architecture-Aware Evaluation Metrics for LLM Agents
di: Souza, Débora, et al.
Pubblicazione: (2026)
di: Souza, Débora, et al.
Pubblicazione: (2026)
FC-TTS: Style and Timbre Control in Zero-Shot Text-to-Speech with Disentangled Speech Representations
di: Lee, Yoonhyung, et al.
Pubblicazione: (2026)
di: Lee, Yoonhyung, et al.
Pubblicazione: (2026)
Forgotten Words: Benchmarking NeoBERT for Dementia Detection in Low-Resource Conversational Filipino and English Speech
di: Floresca, Rez Samantha Z., et al.
Pubblicazione: (2026)
di: Floresca, Rez Samantha Z., et al.
Pubblicazione: (2026)
Algorithm for Semantic Network Generation from Texts of Low Resource Languages Such as Kiswahili
di: Wanjawa, Barack Wamkaya, et al.
Pubblicazione: (2025)
di: Wanjawa, Barack Wamkaya, et al.
Pubblicazione: (2025)
A Multi-Task Benchmark for Abusive Language Detection in Low-Resource Settings
di: Gaim, Fitsum, et al.
Pubblicazione: (2025)
di: Gaim, Fitsum, et al.
Pubblicazione: (2025)
KinyaColBERT: A Lexically Grounded Retrieval Model for Low-Resource Retrieval-Augmented Generation
di: Nzeyimana, Antoine, et al.
Pubblicazione: (2025)
di: Nzeyimana, Antoine, et al.
Pubblicazione: (2025)
Reflective Translation: Improving Low-Resource Machine Translation via Structured Self-Reflection
di: Cheng, Nicholas
Pubblicazione: (2026)
di: Cheng, Nicholas
Pubblicazione: (2026)
MALT: Mechanistic Ablation of Lossy Translation in LLMs for a Low-Resource Language: Urdu
di: Bajwa, Taaha Saleem
Pubblicazione: (2025)
di: Bajwa, Taaha Saleem
Pubblicazione: (2025)
Dialect Matters: Cross-Lingual ASR Transfer for Low-Resource Indic Language Varieties
di: Dhasmana, Akriti, et al.
Pubblicazione: (2026)
di: Dhasmana, Akriti, et al.
Pubblicazione: (2026)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
di: Ashuach, Tomer, et al.
Pubblicazione: (2025)
di: Ashuach, Tomer, et al.
Pubblicazione: (2025)
Hard to Be Heard: Phoneme-Level ASR Analysis of Phonologically Complex, Low-Resource Endangered Languages
di: Akavarapu, V. S. D. S. Mahesh, et al.
Pubblicazione: (2026)
di: Akavarapu, V. S. D. S. Mahesh, et al.
Pubblicazione: (2026)
Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization
di: Wang, Hsuan-Yu, et al.
Pubblicazione: (2025)
di: Wang, Hsuan-Yu, et al.
Pubblicazione: (2025)
Is Textual Similarity Invariant under Machine Translation? Evidence Based on the Political Manifesto Corpus
di: Boratyn, Daria, et al.
Pubblicazione: (2026)
di: Boratyn, Daria, et al.
Pubblicazione: (2026)
WhisperAlign: Word-Boundary-Aware ASR and WhisperX-Anchored Pyannote Diarization for Long-Form Bengali Speech
di: Chowdhury, Aurchi, et al.
Pubblicazione: (2026)
di: Chowdhury, Aurchi, et al.
Pubblicazione: (2026)
Convex Low-resource Accent-Robust Language Detection in Speech Recognition
di: Feng, Miria, et al.
Pubblicazione: (2026)
di: Feng, Miria, et al.
Pubblicazione: (2026)
Towards Platonic Representation for Table Reasoning: A Foundation for Permutation-Invariant Retrieval
di: Tchuitcheu, Willy Carlos, et al.
Pubblicazione: (2026)
di: Tchuitcheu, Willy Carlos, et al.
Pubblicazione: (2026)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
di: Saji, Alan, et al.
Pubblicazione: (2025)
di: Saji, Alan, et al.
Pubblicazione: (2025)
Synthetic Voice Data for Automatic Speech Recognition in African Languages
di: DeRenzi, Brian, et al.
Pubblicazione: (2025)
di: DeRenzi, Brian, et al.
Pubblicazione: (2025)
Combining Data Generation and Active Learning for Low-Resource Question Answering
di: Kimmich, Maximilian, et al.
Pubblicazione: (2022)
di: Kimmich, Maximilian, et al.
Pubblicazione: (2022)
Blocks Architecture (BloArk): Efficient, Cost-Effective, and Incremental Dataset Architecture for Wikipedia Revision History
di: Li, Lingxi, et al.
Pubblicazione: (2024)
di: Li, Lingxi, et al.
Pubblicazione: (2024)
2M-BELEBELE: Highly Multilingual Speech and American Sign Language Comprehension Dataset
di: Costa-jussà, Marta R., et al.
Pubblicazione: (2024)
di: Costa-jussà, Marta R., et al.
Pubblicazione: (2024)
Named Entity Recognition for Address Extraction in Speech-to-Text Transcriptions Using Synthetic Data
di: Lajčinová, Bibiána, et al.
Pubblicazione: (2024)
di: Lajčinová, Bibiána, et al.
Pubblicazione: (2024)
The Curious Case of Visual Grounding: Different Effects for Speech- and Text-based Language Encoders
di: Sauter, Adrian, et al.
Pubblicazione: (2025)
di: Sauter, Adrian, et al.
Pubblicazione: (2025)
Low-resource neural machine translation with morphological modeling
di: Nzeyimana, Antoine
Pubblicazione: (2024)
di: Nzeyimana, Antoine
Pubblicazione: (2024)
Boosting Accuracy and Interpretability in Multilingual Hate Speech Detection Through Layer Freezing and Explainable AI
di: Bilehsavar, Meysam Shirdel, et al.
Pubblicazione: (2026)
di: Bilehsavar, Meysam Shirdel, et al.
Pubblicazione: (2026)
Sensitive Content Classification in Social Media: A Holistic Resource and Evaluation
di: Antypas, Dimosthenis, et al.
Pubblicazione: (2024)
di: Antypas, Dimosthenis, et al.
Pubblicazione: (2024)
Emotional Sequential Influence Modeling on False Information
di: Naskar, Debashis, et al.
Pubblicazione: (2024)
di: Naskar, Debashis, et al.
Pubblicazione: (2024)
What Language Models Know But Don't Say: Non-Generative Prior Extraction for Generalization
di: Rezaeimanesh, Sara, et al.
Pubblicazione: (2026)
di: Rezaeimanesh, Sara, et al.
Pubblicazione: (2026)
Culturally-Nuanced Story Generation for Reasoning in Low-Resource Languages: The Case of Javanese and Sundanese
di: Pranida, Salsabila Zahirah, et al.
Pubblicazione: (2025)
di: Pranida, Salsabila Zahirah, et al.
Pubblicazione: (2025)
Beyond Many-Shot Translation: Scaling In-Context Demonstrations For Low-Resource Machine Translation
di: Salim, Luis Frentzen, et al.
Pubblicazione: (2026)
di: Salim, Luis Frentzen, et al.
Pubblicazione: (2026)
Mean-Pooled Cosine Similarity is Not Length-Invariant: Theory and Cross-Domain Evidence for a Length-Invariant Alternative
di: Mitra, Sibayan, et al.
Pubblicazione: (2026)
di: Mitra, Sibayan, et al.
Pubblicazione: (2026)
Towards Fundamental Language Models: Does Linguistic Competence Scale with Model Size?
di: Collado-Montañez, Jaime, et al.
Pubblicazione: (2025)
di: Collado-Montañez, Jaime, et al.
Pubblicazione: (2025)
SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset
di: Smădu, Răzvan-Alexandru, et al.
Pubblicazione: (2025)
di: Smădu, Răzvan-Alexandru, et al.
Pubblicazione: (2025)
Automated Bug Triaging using Instruction-Tuned Large Language Models
di: Kiashemshaki, Kiana, et al.
Pubblicazione: (2025)
di: Kiashemshaki, Kiana, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Low-Resource Court Judgment Summarization for Common Law Systems
di: Liu, Shuaiqi, et al.
Pubblicazione: (2024) -
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
di: Peters, Sydney, et al.
Pubblicazione: (2025) -
Heidelberg-Boston @ SIGTYP 2024 Shared Task: Enhancing Low-Resource Language Analysis With Character-Aware Hierarchical Transformers
di: Riemenschneider, Frederick, et al.
Pubblicazione: (2024) -
Spatially-Aware Speaker for Vision-and-Language Navigation Instruction Generation
di: Gopinathan, Muraleekrishna, et al.
Pubblicazione: (2024) -
Robust Long-Form Bangla Speech Processing: Automatic Speech Recognition and Speaker Diarization
di: Chowdhury, MD. Sagor, et al.
Pubblicazione: (2026)