Lexical Tone is Hard to Quantize: Probing Discrete Speech Units in Mandarin and Yorùbá
Fuente:
arXiv
Saved in:
| Main Authors: | Osakuade, Opeyemi, King, Simon |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Do Discrete Self-Supervised Representations of Speech Capture Tone Distinctions?
by: Osakuade, Opeyemi, et al.
Published: (2024)
by: Osakuade, Opeyemi, et al.
Published: (2024)
Yoruba-G2P: A tone-aware grapheme-to-phoneme converter for Yorùbá
by: Osakuade, Opeyemi
Published: (2026)
by: Osakuade, Opeyemi
Published: (2026)
Yoruba-G2P: A tone-aware grapheme-to-phoneme converter for Yorùbá
by: Osakuade, Opeyemi
Published: (2026)
by: Osakuade, Opeyemi
Published: (2026)
Improving Direct Persian-English Speech-to-Speech Translation with Discrete Units and Synthetic Parallel Data
by: Rashidi, Sina, et al.
Published: (2025)
by: Rashidi, Sina, et al.
Published: (2025)
Residual Speech Embeddings for Tone Classification: Removing Linguistic Content to Enhance Paralinguistic Analysis
by: Ahbabi, Hamdan Al, et al.
Published: (2025)
by: Ahbabi, Hamdan Al, et al.
Published: (2025)
Transfer Learning via Lexical Relatedness: A Sarcasm and Hate Speech Case Study
by: Cabrera, Angelly, et al.
Published: (2025)
by: Cabrera, Angelly, et al.
Published: (2025)
Towards Homogeneous Lexical Tone Decoding from Heterogeneous Intracranial Recordings
by: Wu, Di, et al.
Published: (2024)
by: Wu, Di, et al.
Published: (2024)
Edeflip: Supervised Word Translation between English and Yoruba
by: Abioye, Ikeoluwa, et al.
Published: (2025)
by: Abioye, Ikeoluwa, et al.
Published: (2025)
On Lexical Invariance on Multisets and Graphs
by: Zhang, Muhan
Published: (2024)
by: Zhang, Muhan
Published: (2024)
Unsupervised Learning and Representation of Mandarin Tonal Categories by a Generative CNN
by: Schenck, Kai, et al.
Published: (2025)
by: Schenck, Kai, et al.
Published: (2025)
Lexical Hints of Accuracy in LLM Reasoning Chains
by: Vanhoyweghen, Arne, et al.
Published: (2025)
by: Vanhoyweghen, Arne, et al.
Published: (2025)
Luxical: High-Speed Lexical-Dense Text Embeddings
by: DatologyAI, et al.
Published: (2025)
by: DatologyAI, et al.
Published: (2025)
Analyzing Speech Unit Selection for Textless Speech-to-Speech Translation
by: Duret, Jarod, et al.
Published: (2024)
by: Duret, Jarod, et al.
Published: (2024)
PoeTone: A Framework for Constrained Generation of Structured Chinese Songci with LLMs
by: Qu, Zhan, et al.
Published: (2025)
by: Qu, Zhan, et al.
Published: (2025)
Neural Recovery of Historical Lexical Structure in Bantu Languages from Modern Data
by: Mutisya, Hillary, et al.
Published: (2026)
by: Mutisya, Hillary, et al.
Published: (2026)
Meta-Tuning LLMs to Leverage Lexical Knowledge for Generalizable Language Style Understanding
by: Guo, Ruohao, et al.
Published: (2023)
by: Guo, Ruohao, et al.
Published: (2023)
Moonshine v2: Ergodic Streaming Encoder ASR for Latency-Critical Speech Applications
by: Kudlur, Manjunath, et al.
Published: (2026)
by: Kudlur, Manjunath, et al.
Published: (2026)
Multilingual Lexical Feature Analysis of Spoken Language for Predicting Major Depression Symptom Severity
by: Tokareva, Anastasiia, et al.
Published: (2025)
by: Tokareva, Anastasiia, et al.
Published: (2025)
Model Internal Sleuthing: Finding Lexical Identity and Inflectional Features in Modern Language Models
by: Li, Michael, et al.
Published: (2025)
by: Li, Michael, et al.
Published: (2025)
UDDETTS: Unifying Discrete and Dimensional Emotions for Controllable Emotional Text-to-Speech
by: Liu, Jiaxuan, et al.
Published: (2025)
by: Liu, Jiaxuan, et al.
Published: (2025)
Codec-ASR: Training Performant Automatic Speech Recognition Systems with Discrete Speech Representations
by: Dhawan, Kunal, et al.
Published: (2024)
by: Dhawan, Kunal, et al.
Published: (2024)
Why is "Chicago" Predictive of Deceptive Reviews? Using LLMs to Discover Language Phenomena from Lexical Cues
by: Qu, Jiaming, et al.
Published: (2025)
by: Qu, Jiaming, et al.
Published: (2025)
Through a Compressed Lens: Investigating The Impact of Quantization on Factual Knowledge Recall
by: Wang, Qianli, et al.
Published: (2025)
by: Wang, Qianli, et al.
Published: (2025)
Shared Lexical Task Representations Explain Behavioral Variability In LLMs
by: Yang, Zhuonan, et al.
Published: (2026)
by: Yang, Zhuonan, et al.
Published: (2026)
Sensitivity of Generative VLMs to Semantically and Lexically Altered Prompts
by: Dumpala, Sri Harsha, et al.
Published: (2024)
by: Dumpala, Sri Harsha, et al.
Published: (2024)
Masked Gated Linear Unit
by: Tajima, Yukito, et al.
Published: (2025)
by: Tajima, Yukito, et al.
Published: (2025)
Do Lexical and Contextual Coreference Resolution Systems Degrade Differently under Mention Noise? An Empirical Study on Scientific Software Mentions
by: Alkan, Atilla Kaan, et al.
Published: (2026)
by: Alkan, Atilla Kaan, et al.
Published: (2026)
Zipf Distributions from Two-Stage Symbolic Processes: Stability Under Stochastic Lexical Filtering
by: Berman, Vladimir
Published: (2025)
by: Berman, Vladimir
Published: (2025)
Scaling Laws For Mixed Quantization
by: Cao, Zeyu, et al.
Published: (2024)
by: Cao, Zeyu, et al.
Published: (2024)
Supplementary Resources and Analysis for Automatic Speech Recognition Systems Trained on the Loquacious Dataset
by: Rossenbach, Nick, et al.
Published: (2025)
by: Rossenbach, Nick, et al.
Published: (2025)
Constrained Discrete Diffusion
by: Cardei, Michael, et al.
Published: (2025)
by: Cardei, Michael, et al.
Published: (2025)
Low-Bit Quantization Favors Undertrained LLMs: Scaling Laws for Quantized LLMs with 100T Training Tokens
by: Ouyang, Xu, et al.
Published: (2024)
by: Ouyang, Xu, et al.
Published: (2024)
EuroSpeech: A Multilingual Speech Corpus
by: Pfisterer, Samuel, et al.
Published: (2025)
by: Pfisterer, Samuel, et al.
Published: (2025)
Learning to Translate from Soft to Hard LLM Prompts
by: Kongsomjit, Pitipat, et al.
Published: (2026)
by: Kongsomjit, Pitipat, et al.
Published: (2026)
GPTVQ: The Blessing of Dimensionality for LLM Quantization
by: van Baalen, Mart, et al.
Published: (2024)
by: van Baalen, Mart, et al.
Published: (2024)
Scaling Law for Quantization-Aware Training
by: Chen, Mengzhao, et al.
Published: (2025)
by: Chen, Mengzhao, et al.
Published: (2025)
ZOQO: Zero-Order Quantized Optimization
by: Bar, Noga, et al.
Published: (2025)
by: Bar, Noga, et al.
Published: (2025)
SqueezeLLM: Dense-and-Sparse Quantization
by: Kim, Sehoon, et al.
Published: (2023)
by: Kim, Sehoon, et al.
Published: (2023)
Discrete Neural Algorithmic Reasoning
by: Rodionov, Gleb, et al.
Published: (2024)
by: Rodionov, Gleb, et al.
Published: (2024)
On the Importance of a Multi-Scale Calibration for Quantization
by: Son, Seungwoo, et al.
Published: (2026)
by: Son, Seungwoo, et al.
Published: (2026)
Similar Items
-
Do Discrete Self-Supervised Representations of Speech Capture Tone Distinctions?
by: Osakuade, Opeyemi, et al.
Published: (2024) -
Yoruba-G2P: A tone-aware grapheme-to-phoneme converter for Yorùbá
by: Osakuade, Opeyemi
Published: (2026) -
Yoruba-G2P: A tone-aware grapheme-to-phoneme converter for Yorùbá
by: Osakuade, Opeyemi
Published: (2026) -
Improving Direct Persian-English Speech-to-Speech Translation with Discrete Units and Synthetic Parallel Data
by: Rashidi, Sina, et al.
Published: (2025) -
Residual Speech Embeddings for Tone Classification: Removing Linguistic Content to Enhance Paralinguistic Analysis
by: Ahbabi, Hamdan Al, et al.
Published: (2025)