Guardado en:
| Autores principales: | Xiao, Cihan, Liang, Ruixing, Zhang, Xiangyu, Tiryaki, Mehmet Emre, Bae, Veronica, Shankar, Lavanya, Yang, Rong, Poon, Ethan, Dupoux, Emmanuel, Khudanpur, Sanjeev, Perera, Leibny Paola Garcia |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2506.00267 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
HENT-SRT: Hierarchical Efficient Neural Transducer with Self-Distillation for Joint Speech Recognition and Translation
por: Hussein, Amir, et al.
Publicado: (2025)
por: Hussein, Amir, et al.
Publicado: (2025)
ShiftySpeech: A Large-Scale Synthetic Speech Dataset with Distribution Shifts
por: Garg, Ashi, et al.
Publicado: (2025)
por: Garg, Ashi, et al.
Publicado: (2025)
Universal Speech Content Factorization
por: Xinyuan, Henry Li, et al.
Publicado: (2026)
por: Xinyuan, Henry Li, et al.
Publicado: (2026)
On Speaker Attribution with SURT
por: Raj, Desh, et al.
Publicado: (2024)
por: Raj, Desh, et al.
Publicado: (2024)
Can LLMs Help Localize Fake Words in Partially Fake Speech?
por: Zhang, Lin, et al.
Publicado: (2026)
por: Zhang, Lin, et al.
Publicado: (2026)
Rapidly Adapting to New Voice Spoofing: Few-Shot Detection of Synthesized Speech Under Distribution Shifts
por: Garg, Ashi, et al.
Publicado: (2025)
por: Garg, Ashi, et al.
Publicado: (2025)
Scalable Controllable Accented TTS
por: Xinyuan, Henry Li, et al.
Publicado: (2025)
por: Xinyuan, Henry Li, et al.
Publicado: (2025)
Leveraging Zipformer Model for Effective Language Identification in Code-Switched Child-Directed Speech
por: Shankar, Lavanya, et al.
Publicado: (2025)
por: Shankar, Lavanya, et al.
Publicado: (2025)
GenVC: Self-Supervised Zero-Shot Voice Conversion
por: Cai, Zexin, et al.
Publicado: (2025)
por: Cai, Zexin, et al.
Publicado: (2025)
Privacy versus Emotion Preservation Trade-offs in Emotion-Preserving Speaker Anonymization
por: Cai, Zexin, et al.
Publicado: (2024)
por: Cai, Zexin, et al.
Publicado: (2024)
HLTCOE JHU Submission to the Voice Privacy Challenge 2024
por: Xinyuan, Henry Li, et al.
Publicado: (2024)
por: Xinyuan, Henry Li, et al.
Publicado: (2024)
Integrated Spoofing-Robust Automatic Speaker Verification via a Three-Class Formulation and LLR
por: Tan, Kai, et al.
Publicado: (2026)
por: Tan, Kai, et al.
Publicado: (2026)
Improving Neural Biasing for Contextual Speech Recognition by Early Context Injection and Text Perturbation
por: Huang, Ruizhe, et al.
Publicado: (2024)
por: Huang, Ruizhe, et al.
Publicado: (2024)
Aligning Speech to Languages to Enhance Code-switching Speech Recognition
por: Liu, Hexin, et al.
Publicado: (2024)
por: Liu, Hexin, et al.
Publicado: (2024)
Unsupervised Speech Enhancement using Data-defined Priors
por: Klement, Dominik, et al.
Publicado: (2025)
por: Klement, Dominik, et al.
Publicado: (2025)
Building Corpora for Single-Channel Speech Separation Across Multiple Domains
por: Maciejewski, Matthew, et al.
Publicado: (2018)
por: Maciejewski, Matthew, et al.
Publicado: (2018)
Acoustic modeling for Overlapping Speech Recognition: JHU Chime-5 Challenge System
por: Manohar, Vimal, et al.
Publicado: (2024)
por: Manohar, Vimal, et al.
Publicado: (2024)
Dimensionality-Aware Anomaly Detection in Learned Representations of Self-Supervised Speech Models
por: Arcos-Holzinger, Sandra, et al.
Publicado: (2026)
por: Arcos-Holzinger, Sandra, et al.
Publicado: (2026)
EmphAssess : a Prosodic Benchmark on Assessing Emphasis Transfer in Speech-to-Speech Models
por: de Seyssel, Maureen, et al.
Publicado: (2023)
por: de Seyssel, Maureen, et al.
Publicado: (2023)
Improving Spoken Language Modeling with Phoneme Classification: A Simple Fine-tuning Approach
por: Poli, Maxime, et al.
Publicado: (2024)
por: Poli, Maxime, et al.
Publicado: (2024)
fastabx: A library for efficient computation of ABX discriminability
por: Poli, Maxime, et al.
Publicado: (2025)
por: Poli, Maxime, et al.
Publicado: (2025)
Modeling Overlapped Speech with Shuffles
por: Wiesner, Matthew, et al.
Publicado: (2026)
por: Wiesner, Matthew, et al.
Publicado: (2026)
Adversarial Attacks and Defenses for Speech Recognition Systems
por: Żelasko, Piotr, et al.
Publicado: (2021)
por: Żelasko, Piotr, et al.
Publicado: (2021)
Speaking in Wavelet Domain: A Simple and Efficient Approach to Speed up Speech Diffusion Model
por: Zhang, Xiangyu, et al.
Publicado: (2024)
por: Zhang, Xiangyu, et al.
Publicado: (2024)
SpatialEmb: Extract and Encode Spatial Information for 1-Stage Multi-channel Multi-speaker ASR on Arbitrary Microphone Arrays
por: Shao, Yiwen, et al.
Publicado: (2026)
por: Shao, Yiwen, et al.
Publicado: (2026)
DiscoPhon: Benchmarking the Unsupervised Discovery of Phoneme Inventories With Discrete Speech Units
por: Poli, Maxime, et al.
Publicado: (2026)
por: Poli, Maxime, et al.
Publicado: (2026)
DiCoW: Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition
por: Polok, Alexander, et al.
Publicado: (2024)
por: Polok, Alexander, et al.
Publicado: (2024)
Target Speaker ASR with Whisper
por: Polok, Alexander, et al.
Publicado: (2024)
por: Polok, Alexander, et al.
Publicado: (2024)
Simulating Articulatory Trajectories with Phonological Feature Interpolation
por: Tandazo, Angelo Ortiz, et al.
Publicado: (2024)
por: Tandazo, Angelo Ortiz, et al.
Publicado: (2024)
Challenges in Automated Processing of Speech from Child Wearables: The Case of Voice Type Classifier
por: Kunze, Tarek, et al.
Publicado: (2025)
por: Kunze, Tarek, et al.
Publicado: (2025)
Multi-Channel Multi-Speaker ASR Using Target Speaker's Solo Segment
por: Shao, Yiwen, et al.
Publicado: (2024)
por: Shao, Yiwen, et al.
Publicado: (2024)
iMiGUE-Speech: A Spontaneous Speech Dataset for Affective Analysis
por: Kakouros, Sofoklis, et al.
Publicado: (2026)
por: Kakouros, Sofoklis, et al.
Publicado: (2026)
Auto-Landmark: Acoustic Landmark Dataset and Open-Source Toolkit for Landmark Extraction
por: Zhang, Xiangyu, et al.
Publicado: (2024)
por: Zhang, Xiangyu, et al.
Publicado: (2024)
SE-DiCoW: Self-Enrolled Diarization-Conditioned Whisper
por: Polok, Alexander, et al.
Publicado: (2026)
por: Polok, Alexander, et al.
Publicado: (2026)
An open-source voice type classifier for child-centered daylong recordings
por: Lavechin, Marvin, et al.
Publicado: (2020)
por: Lavechin, Marvin, et al.
Publicado: (2020)
MauBERT: Universal Phonetic Inductive Biases for Few-Shot Acoustic Units Discovery
por: Tandazo, Angelo Ortiz, et al.
Publicado: (2025)
por: Tandazo, Angelo Ortiz, et al.
Publicado: (2025)
Clean Label Attacks against SLU Systems
por: Xinyuan, Henry Li, et al.
Publicado: (2024)
por: Xinyuan, Henry Li, et al.
Publicado: (2024)
SAV-SE: Scene-aware Audio-Visual Speech Enhancement with Selective State Space Model
por: Qian, Xinyuan, et al.
Publicado: (2024)
por: Qian, Xinyuan, et al.
Publicado: (2024)
A Large Dataset of Spontaneous Speech with the Accent Spoken in São Paulo for Automatic Speech Recognition Evaluation
por: Lima, Rodrigo, et al.
Publicado: (2024)
por: Lima, Rodrigo, et al.
Publicado: (2024)
CHSER: A Dataset and Case Study on Generative Speech Error Correction for Child ASR
por: Shankar, Natarajan Balaji, et al.
Publicado: (2025)
por: Shankar, Natarajan Balaji, et al.
Publicado: (2025)
Ejemplares similares
-
HENT-SRT: Hierarchical Efficient Neural Transducer with Self-Distillation for Joint Speech Recognition and Translation
por: Hussein, Amir, et al.
Publicado: (2025) -
ShiftySpeech: A Large-Scale Synthetic Speech Dataset with Distribution Shifts
por: Garg, Ashi, et al.
Publicado: (2025) -
Universal Speech Content Factorization
por: Xinyuan, Henry Li, et al.
Publicado: (2026) -
On Speaker Attribution with SURT
por: Raj, Desh, et al.
Publicado: (2024) -
Can LLMs Help Localize Fake Words in Partially Fake Speech?
por: Zhang, Lin, et al.
Publicado: (2026)