Saved in:
| Main Authors: | Chu, Wei, Dong, Yuanzhe, Tan, Ke, Han, Dong, Menendez-Pidal, Xavier, Fan, Ruchao, Miao, Chenfeng, Kim, Chanwoo, Raj, Bhiksha, Singh, Rita |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2509.04702 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
What Do Speech Foundation Models Not Learn About Speech?
by: Waheed, Abdul, et al.
Published: (2024)
by: Waheed, Abdul, et al.
Published: (2024)
DELULU: Discriminative Embedding Learning Using Latent Units for Speaker-Aware Self-Trained Speech Foundational Model
by: Baali, Massa, et al.
Published: (2025)
by: Baali, Massa, et al.
Published: (2025)
ZMM-TTS: Zero-shot Multilingual and Multispeaker Speech Synthesis Conditioned on Self-supervised Discrete Speech Representations
by: Gong, Cheng, et al.
Published: (2023)
by: Gong, Cheng, et al.
Published: (2023)
Speech vs. Transcript: Does It Matter for Human Annotators in Speech Summarization?
by: Sharma, Roshan, et al.
Published: (2024)
by: Sharma, Roshan, et al.
Published: (2024)
SELM: Enhancing Speech Emotion Recognition for Out-of-Domain Scenarios
by: Bukhari, Hazim, et al.
Published: (2024)
by: Bukhari, Hazim, et al.
Published: (2024)
Flor nueva de romances viejos / Ramón Menéndez Pidal
by: Menéndez Pidal, Ramón
Published: (1955)
by: Menéndez Pidal, Ramón
Published: (1955)
Los romances de América y otros estudios / Ramón Menéndez Pidal
by: Menéndez Pidal, Ramón
Published: (1972)
by: Menéndez Pidal, Ramón
Published: (1972)
Tres poetas primitivos : Elena y María, Roncesvalles, Historia troyana polimétrica / Ramón Menéndez Pidal
by: Menéndez Pidal, Ramón
Published: (1948)
by: Menéndez Pidal, Ramón
Published: (1948)
Poesía juglaresca y juglares : spectos de la historia literaria y cultural de España / Ramón Menéndez Pidal
by: Menéndez Pidal, Ramón
Published: (1983)
by: Menéndez Pidal, Ramón
Published: (1983)
La epopeya castellana a través de la literatura española / Ramón Menéndez Pidal
by: Menéndez Pidal, Ramón
Published: (1974)
by: Menéndez Pidal, Ramón
Published: (1974)
DART: Disentanglement of Accent and Speaker Representation in Multispeaker Text-to-Speech
by: Melechovsky, Jan, et al.
Published: (2024)
by: Melechovsky, Jan, et al.
Published: (2024)
Lost in Transcription, Found in Distribution Shift: Demystifying Hallucination in Speech Foundation Models
by: Atwany, Hanin, et al.
Published: (2025)
by: Atwany, Hanin, et al.
Published: (2025)
Large Language Model Guided Decoding for Self-Supervised Speech Recognition
by: Cohen, Eyal, et al.
Published: (2025)
by: Cohen, Eyal, et al.
Published: (2025)
La délimitation du champ littéraire dans les romans d’Ahmadou Kourouma
by: Laura Menéndez-Pidal Sendrail
Published: (2014)
by: Laura Menéndez-Pidal Sendrail
Published: (2014)
Retórica visual: una herramienta necesaria en la creación e interpretación de productos visuales
by: Silvia Nuere Menéndez-Pidal
Published: (2010)
by: Silvia Nuere Menéndez-Pidal
Published: (2010)
Human Voice is Unique
by: Singh, Rita, et al.
Published: (2025)
by: Singh, Rita, et al.
Published: (2025)
Speech Robust Bench: A Robustness Benchmark For Speech Recognition
by: Shah, Muhammad A., et al.
Published: (2024)
by: Shah, Muhammad A., et al.
Published: (2024)
Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
by: He, Haorui, et al.
Published: (2024)
by: He, Haorui, et al.
Published: (2024)
Black hole singularity resolution in unimodular gravity from unitarity
by: Gielen, Steffen, et al.
Published: (2024)
by: Gielen, Steffen, et al.
Published: (2024)
RelUNet: Relative Channel Fusion U-Net for Multichannel Speech Enhancement
by: Aldarmaki, Ibrahim, et al.
Published: (2024)
by: Aldarmaki, Ibrahim, et al.
Published: (2024)
Domain Adaptation for Contrastive Audio-Language Models
by: Deshmukh, Soham, et al.
Published: (2024)
by: Deshmukh, Soham, et al.
Published: (2024)
Benchmarking Children's ASR with Supervised and Self-supervised Speech Foundation Models
by: Fan, Ruchao, et al.
Published: (2024)
by: Fan, Ruchao, et al.
Published: (2024)
Speech LLMs are Contextual Reasoning Transcribers
by: Deng, Keqi, et al.
Published: (2026)
by: Deng, Keqi, et al.
Published: (2026)
UniEnc-CASSNAT: An Encoder-only Non-autoregressive ASR for Speech SSL Models
by: Fan, Ruchao, et al.
Published: (2024)
by: Fan, Ruchao, et al.
Published: (2024)
A Unified Speech LLM for Diarization and Speech Recognition in Multilingual Conversations
by: Saengthong, Phurich, et al.
Published: (2025)
by: Saengthong, Phurich, et al.
Published: (2025)
SOA: Reducing Domain Mismatch in SSL Pipeline by Speech Only Adaptation for Low Resource ASR
by: Shankar, Natarajan Balaji, et al.
Published: (2024)
by: Shankar, Natarajan Balaji, et al.
Published: (2024)
Speech-MASSIVE: A Multilingual Speech Dataset for SLU and Beyond
by: Lee, Beomseok, et al.
Published: (2024)
by: Lee, Beomseok, et al.
Published: (2024)
Advancing Topic Segmentation of Broadcasted Speech with Multilingual Semantic Embeddings
by: Shukla, Sakshi Deo, et al.
Published: (2024)
by: Shukla, Sakshi Deo, et al.
Published: (2024)
Psychoacoustic Challenges Of Speech Enhancement On VoIP Platforms
by: Konan, Joseph, et al.
Published: (2023)
by: Konan, Joseph, et al.
Published: (2023)
SVeritas: Benchmark for Robust Speaker Verification under Diverse Conditions
by: Baali, Massa, et al.
Published: (2025)
by: Baali, Massa, et al.
Published: (2025)
Emilia: A Large-Scale, Extensive, Multilingual, and Diverse Dataset for Speech Generation
by: He, Haorui, et al.
Published: (2025)
by: He, Haorui, et al.
Published: (2025)
NaturalL2S: End-to-End High-quality Multispeaker Lip-to-Speech Synthesis with Differential Digital Signal Processing
by: Liang, Yifan, et al.
Published: (2025)
by: Liang, Yifan, et al.
Published: (2025)
Summary on The Multilingual Conversational Speech Language Model Challenge: Datasets, Tasks, Baselines, and Methods
by: Mu, Bingshen, et al.
Published: (2025)
by: Mu, Bingshen, et al.
Published: (2025)
DisfluencySpeech -- Single-Speaker Conversational Speech Dataset with Paralanguage
by: Wang, Kyra, et al.
Published: (2024)
by: Wang, Kyra, et al.
Published: (2024)
Reasoning-Based Approach with Chain-of-Thought for Alzheimer's Detection Using Speech and Large Language Models
by: Park, Chanwoo, et al.
Published: (2025)
by: Park, Chanwoo, et al.
Published: (2025)
RLBR: Reinforcement Learning with Biasing Rewards for Contextual Speech Large Language Models
by: Ren, Bo, et al.
Published: (2026)
by: Ren, Bo, et al.
Published: (2026)
What and When to Learn: CURriculum Ranking Loss for Large-Scale Speaker Verification
by: Baali, Massa, et al.
Published: (2026)
by: Baali, Massa, et al.
Published: (2026)
ADIFF: Explaining audio difference using natural language
by: Deshmukh, Soham, et al.
Published: (2025)
by: Deshmukh, Soham, et al.
Published: (2025)
Deciphering GunType Hierarchy through Acoustic Analysis of Gunshot Recordings
by: Shah, Ankit, et al.
Published: (2025)
by: Shah, Ankit, et al.
Published: (2025)
ADIFF: Explaining audio difference using natural language
by: Deshmukh, Soham, et al.
Published: (2025)
by: Deshmukh, Soham, et al.
Published: (2025)
Similar Items
-
What Do Speech Foundation Models Not Learn About Speech?
by: Waheed, Abdul, et al.
Published: (2024) -
DELULU: Discriminative Embedding Learning Using Latent Units for Speaker-Aware Self-Trained Speech Foundational Model
by: Baali, Massa, et al.
Published: (2025) -
ZMM-TTS: Zero-shot Multilingual and Multispeaker Speech Synthesis Conditioned on Self-supervised Discrete Speech Representations
by: Gong, Cheng, et al.
Published: (2023) -
Speech vs. Transcript: Does It Matter for Human Annotators in Speech Summarization?
by: Sharma, Roshan, et al.
Published: (2024) -
SELM: Enhancing Speech Emotion Recognition for Out-of-Domain Scenarios
by: Bukhari, Hazim, et al.
Published: (2024)