Gespeichert in:
| Hauptverfasser: | Yeo, Yue Heng, Hu, Yuchen, Gopal, Shreyas, Peng, Yizhou, Liu, Hexin, Chng, Eng Siong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2601.00935 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Bi-directional Context-Enhanced Speech Large Language Models for Multilingual Conversational ASR
von: Peng, Yizhou, et al.
Veröffentlicht: (2025)
von: Peng, Yizhou, et al.
Veröffentlicht: (2025)
Code-switching Speech Recognition Under the Lens: Model- and Data-Centric Perspectives
von: Liu, Hexin, et al.
Veröffentlicht: (2025)
von: Liu, Hexin, et al.
Veröffentlicht: (2025)
Aligning Speech to Languages to Enhance Code-switching Speech Recognition
von: Liu, Hexin, et al.
Veröffentlicht: (2024)
von: Liu, Hexin, et al.
Veröffentlicht: (2024)
GenSE: Generative Speech Enhancement via Language Models using Hierarchical Modeling
von: Yao, Jixun, et al.
Veröffentlicht: (2025)
von: Yao, Jixun, et al.
Veröffentlicht: (2025)
Zero-shot Context Biasing with Trie-based Decoding using Synthetic Multi-Pronunciation
von: Liu, Changsong, et al.
Veröffentlicht: (2025)
von: Liu, Changsong, et al.
Veröffentlicht: (2025)
EASY: Emotion-aware Speaker Anonymization via Factorized Distillation
von: Yao, Jixun, et al.
Veröffentlicht: (2025)
von: Yao, Jixun, et al.
Veröffentlicht: (2025)
Bridging Speech and Text: Enhancing ASR with Pinyin-to-Character Pre-training in LLMs
von: Yuhang, Yang, et al.
Veröffentlicht: (2024)
von: Yuhang, Yang, et al.
Veröffentlicht: (2024)
Wav2code: Restore Clean Speech Representations via Codebook Lookup for Noise-Robust ASR
von: Hu, Yuchen, et al.
Veröffentlicht: (2023)
von: Hu, Yuchen, et al.
Veröffentlicht: (2023)
Noise-Aware Speech Separation with Contrastive Learning
von: Zhang, Zizheng, et al.
Veröffentlicht: (2023)
von: Zhang, Zizheng, et al.
Veröffentlicht: (2023)
Speechless: Speech Instruction Training Without Speech for Low Resource Languages
von: Dao, Alan, et al.
Veröffentlicht: (2025)
von: Dao, Alan, et al.
Veröffentlicht: (2025)
Noise-aware Speech Enhancement using Diffusion Probabilistic Model
von: Hu, Yuchen, et al.
Veröffentlicht: (2023)
von: Hu, Yuchen, et al.
Veröffentlicht: (2023)
Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
NTU Speechlab LLM-Based Multilingual ASR System for Interspeech MLC-SLM Challenge 2025
von: Peng, Yizhou, et al.
Veröffentlicht: (2025)
von: Peng, Yizhou, et al.
Veröffentlicht: (2025)
Speech Separation using Neural Audio Codecs with Embedding Loss
von: Yip, Jia Qi, et al.
Veröffentlicht: (2024)
von: Yip, Jia Qi, et al.
Veröffentlicht: (2024)
Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback
von: Chen, Chen, et al.
Veröffentlicht: (2024)
von: Chen, Chen, et al.
Veröffentlicht: (2024)
Speaking in Wavelet Domain: A Simple and Efficient Approach to Speed up Speech Diffusion Model
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2024)
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2024)
From KAN to GR-KAN: Advancing Speech Enhancement with KAN-Based Methodology
von: Li, Haoyang, et al.
Veröffentlicht: (2024)
von: Li, Haoyang, et al.
Veröffentlicht: (2024)
Proactive for Uncertainty: Cause-Aware Error Diagnosis and Interactive Clarification for Spoken Dialogue Systems
von: Peng, Yizhou, et al.
Veröffentlicht: (2026)
von: Peng, Yizhou, et al.
Veröffentlicht: (2026)
UniArray: Unified Spectral-Spatial Modeling for Array-Geometry-Agnostic Speech Separation
von: Chen, Weiguang, et al.
Veröffentlicht: (2025)
von: Chen, Weiguang, et al.
Veröffentlicht: (2025)
Hierarchical Self-Supervised Representation Learning for Depression Detection from Speech
von: Li, Yuxin, et al.
Veröffentlicht: (2025)
von: Li, Yuxin, et al.
Veröffentlicht: (2025)
Listen Again and Choose the Right Answer: A New Paradigm for Automatic Speech Recognition with Large Language Models
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action
von: Zhang, Haoyang, et al.
Veröffentlicht: (2026)
von: Zhang, Haoyang, et al.
Veröffentlicht: (2026)
GenTSE: Enhancing Target Speaker Extraction via a Coarse-to-Fine Generative Language Model
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
Summary on The Multilingual Conversational Speech Language Model Challenge: Datasets, Tasks, Baselines, and Methods
von: Mu, Bingshen, et al.
Veröffentlicht: (2025)
von: Mu, Bingshen, et al.
Veröffentlicht: (2025)
Robust Localization of Partially Fake Speech: Metrics and Out-of-Domain Evaluation
von: Luong, Hieu-Thi, et al.
Veröffentlicht: (2025)
von: Luong, Hieu-Thi, et al.
Veröffentlicht: (2025)
Dataset-Distillation Generative Model for Speech Emotion Recognition
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2024)
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2024)
LlamaPartialSpoof: An LLM-Driven Fake Speech Dataset Simulating Disinformation Generation
von: Luong, Hieu-Thi, et al.
Veröffentlicht: (2024)
von: Luong, Hieu-Thi, et al.
Veröffentlicht: (2024)
Analysis of Speaker Verification Performance Trade-offs with Neural Audio Codec Transmission
von: Thakur, Nirmalya Mallick, et al.
Veröffentlicht: (2025)
von: Thakur, Nirmalya Mallick, et al.
Veröffentlicht: (2025)
Data Augmentation for End-to-end Code-switching Speech Recognition
von: Du, Chenpeng, et al.
Veröffentlicht: (2020)
von: Du, Chenpeng, et al.
Veröffentlicht: (2020)
Multi-band Frequency Reconstruction for Neural Psychoacoustic Coding
von: Ng, Dianwen, et al.
Veröffentlicht: (2025)
von: Ng, Dianwen, et al.
Veröffentlicht: (2025)
StreamVoiceAnon+: Emotion-Preserving Streaming Speaker Anonymization via Frame-Level Acoustic Distillation
von: Kuzmin, Nikita, et al.
Veröffentlicht: (2026)
von: Kuzmin, Nikita, et al.
Veröffentlicht: (2026)
Towards Audio Codec-based Speech Separation
von: Yip, Jia Qi, et al.
Veröffentlicht: (2024)
von: Yip, Jia Qi, et al.
Veröffentlicht: (2024)
FD-Bench: A Full-Duplex Benchmarking Pipeline Designed for Full Duplex Spoken Dialogue Systems
von: Peng, Yizhou, et al.
Veröffentlicht: (2025)
von: Peng, Yizhou, et al.
Veröffentlicht: (2025)
Stream-Voice-Anon: Enhancing Utility of Real-Time Speaker Anonymization via Neural Audio Codec and Language Models
von: Kuzmin, Nikita, et al.
Veröffentlicht: (2026)
von: Kuzmin, Nikita, et al.
Veröffentlicht: (2026)
Audio Large Language Models Can Be Descriptive Speech Quality Evaluators
von: Chen, Chen, et al.
Veröffentlicht: (2025)
von: Chen, Chen, et al.
Veröffentlicht: (2025)
Speech Enhancement Using Continuous Embeddings of Neural Audio Codec
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
Continual Learning Optimizations for Auto-regressive Decoder of Multilingual ASR systems
von: Kwok, Chin Yuen, et al.
Veröffentlicht: (2024)
von: Kwok, Chin Yuen, et al.
Veröffentlicht: (2024)
Aligning Generative Speech Enhancement with Perceptual Feedback
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
Room Impulse Responses help attackers to evade Deep Fake Detection
von: Luong, Hieu-Thi, et al.
Veröffentlicht: (2024)
von: Luong, Hieu-Thi, et al.
Veröffentlicht: (2024)
SpeechT-RAG: Reliable Depression Detection in LLMs with Retrieval-Augmented Generation Using Speech Timing Information
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2025)
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Bi-directional Context-Enhanced Speech Large Language Models for Multilingual Conversational ASR
von: Peng, Yizhou, et al.
Veröffentlicht: (2025) -
Code-switching Speech Recognition Under the Lens: Model- and Data-Centric Perspectives
von: Liu, Hexin, et al.
Veröffentlicht: (2025) -
Aligning Speech to Languages to Enhance Code-switching Speech Recognition
von: Liu, Hexin, et al.
Veröffentlicht: (2024) -
GenSE: Generative Speech Enhancement via Language Models using Hierarchical Modeling
von: Yao, Jixun, et al.
Veröffentlicht: (2025) -
Zero-shot Context Biasing with Trie-based Decoding using Synthetic Multi-Pronunciation
von: Liu, Changsong, et al.
Veröffentlicht: (2025)