The Overview of Segmental Durations Modification Algorithms on Speech Signal Characteristics
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jang, Kyeomeun, Li, Jiaying, Wang, Yinuo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ERes2NetV2: Boosting Short-Duration Speaker Verification Performance with Computational Efficiency
von: Chen, Yafeng, et al.
Veröffentlicht: (2024)
von: Chen, Yafeng, et al.
Veröffentlicht: (2024)
Microphone Array Signal Processing and Deep Learning for Speech Enhancement
von: Haeb-Umbach, Reinhold, et al.
Veröffentlicht: (2025)
von: Haeb-Umbach, Reinhold, et al.
Veröffentlicht: (2025)
SpeechMLC: Speech Multi-label Classification
von: Kim, Miseul, et al.
Veröffentlicht: (2025)
von: Kim, Miseul, et al.
Veröffentlicht: (2025)
Rec-RIR: Monaural Blind Room Impulse Response Identification via DNN-based Reverberant Speech Reconstruction in STFT Domain
von: Wang, Pengyu, et al.
Veröffentlicht: (2025)
von: Wang, Pengyu, et al.
Veröffentlicht: (2025)
ParaS2S: Benchmarking and Aligning Spoken Language Models for Paralinguistic-aware Speech-to-Speech Interaction
von: Yang, Shu-wen, et al.
Veröffentlicht: (2025)
von: Yang, Shu-wen, et al.
Veröffentlicht: (2025)
Prompt-driven Target Speech Diarization
von: Jiang, Yidi, et al.
Veröffentlicht: (2023)
von: Jiang, Yidi, et al.
Veröffentlicht: (2023)
USDnet: Unsupervised Speech Dereverberation via Neural Forward Filtering
von: Wang, Zhong-Qiu
Veröffentlicht: (2024)
von: Wang, Zhong-Qiu
Veröffentlicht: (2024)
Mixture to Mixture: Leveraging Close-talk Mixtures as Weak-supervision for Speech Separation
von: Wang, Zhong-Qiu
Veröffentlicht: (2024)
von: Wang, Zhong-Qiu
Veröffentlicht: (2024)
A Speech Production Model for Radar: Connecting Speech Acoustics with Radar-Measured Vibrations
von: Lenz, Isabella, et al.
Veröffentlicht: (2025)
von: Lenz, Isabella, et al.
Veröffentlicht: (2025)
SELM: Speech Enhancement Using Discrete Tokens and Language Models
von: Wang, Ziqian, et al.
Veröffentlicht: (2023)
von: Wang, Ziqian, et al.
Veröffentlicht: (2023)
Comparison of Classification Algorithms for COVID19 Detection using Cough Acoustic Signals
von: Erdoğan, Yunus Emre, et al.
Veröffentlicht: (2022)
von: Erdoğan, Yunus Emre, et al.
Veröffentlicht: (2022)
Binaural Localization Model for Speech in Noise
von: Tokala, Vikas, et al.
Veröffentlicht: (2025)
von: Tokala, Vikas, et al.
Veröffentlicht: (2025)
Speech-Based Prioritization for Schizophrenia Intervention
von: Premananth, Gowtham, et al.
Veröffentlicht: (2025)
von: Premananth, Gowtham, et al.
Veröffentlicht: (2025)
SuperM2M: Supervised and Mixture-to-Mixture Co-Learning for Speech Enhancement and Noise-Robust ASR
von: Wang, Zhong-Qiu
Veröffentlicht: (2024)
von: Wang, Zhong-Qiu
Veröffentlicht: (2024)
Speech Enhancement based on cascaded two flows
von: Lee, Seonggyu, et al.
Veröffentlicht: (2025)
von: Lee, Seonggyu, et al.
Veröffentlicht: (2025)
Brain-Informed Speech Separation for Cochlear Implants
von: Gajecki, Tom, et al.
Veröffentlicht: (2026)
von: Gajecki, Tom, et al.
Veröffentlicht: (2026)
FlowSE: Efficient and High-Quality Speech Enhancement via Flow Matching
von: Wang, Ziqian, et al.
Veröffentlicht: (2025)
von: Wang, Ziqian, et al.
Veröffentlicht: (2025)
Independent Feature Enhanced Crossmodal Fusion for Match-Mismatch Classification of Speech Stimulus and EEG Response
von: Fan, Shitong, et al.
Veröffentlicht: (2024)
von: Fan, Shitong, et al.
Veröffentlicht: (2024)
FlowSE: Flow Matching-based Speech Enhancement
von: Lee, Seonggyu, et al.
Veröffentlicht: (2025)
von: Lee, Seonggyu, et al.
Veröffentlicht: (2025)
Bridging the Gap: Integrating Pre-trained Speech Enhancement and Recognition Models for Robust Speech Recognition
von: Wang, Kuan-Chen, et al.
Veröffentlicht: (2024)
von: Wang, Kuan-Chen, et al.
Veröffentlicht: (2024)
Harmonics to the Rescue: Why Voiced Speech is Not a Wss Process
von: Bologni, Giovanni, et al.
Veröffentlicht: (2025)
von: Bologni, Giovanni, et al.
Veröffentlicht: (2025)
Binaural Speech Enhancement Using Complex Convolutional Recurrent Networks
von: Tokala, Vikas, et al.
Veröffentlicht: (2025)
von: Tokala, Vikas, et al.
Veröffentlicht: (2025)
Analyzing the Impact of Accent on English Speech: Acoustic and Articulatory Perspectives
von: Premananth, Gowtham, et al.
Veröffentlicht: (2025)
von: Premananth, Gowtham, et al.
Veröffentlicht: (2025)
Unsupervised Face-Masked Speech Enhancement Using Generative Adversarial Networks With Human-in-the-Loop Assessment Metrics
von: Wang, Syu-Siang, et al.
Veröffentlicht: (2024)
von: Wang, Syu-Siang, et al.
Veröffentlicht: (2024)
TTSlow: Slow Down Text-to-Speech with Efficiency Robustness Evaluations
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2024)
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2024)
AffectSpeech: A Large-Scale Emotional Speech Dataset with Fine-Grained Textual Descriptions for Speech Emotion Captioning and Synthesis
von: Qi, Tianhua, et al.
Veröffentlicht: (2026)
von: Qi, Tianhua, et al.
Veröffentlicht: (2026)
Parameter-Efficient Fine-Tuning of Foundation Models for CLP Speech Classification
von: Bhattacharjee, Susmita, et al.
Veröffentlicht: (2025)
von: Bhattacharjee, Susmita, et al.
Veröffentlicht: (2025)
Exploring Disentangled Neural Speech Codecs from Self-Supervised Representations
von: Aihara, Ryo, et al.
Veröffentlicht: (2025)
von: Aihara, Ryo, et al.
Veröffentlicht: (2025)
Impact of Microphone Array Mismatches to Learning-based Replay Speech Detection
von: Neri, Michael, et al.
Veröffentlicht: (2025)
von: Neri, Michael, et al.
Veröffentlicht: (2025)
Multi-channel Replay Speech Detection using an Adaptive Learnable Beamformer
von: Neri, Michael, et al.
Veröffentlicht: (2025)
von: Neri, Michael, et al.
Veröffentlicht: (2025)
DeFTAN-II: Efficient Multichannel Speech Enhancement with Subgroup Processing
von: Lee, Dongheon, et al.
Veröffentlicht: (2023)
von: Lee, Dongheon, et al.
Veröffentlicht: (2023)
Entropy-Guided GRVQ for Ultra-Low Bitrate Neural Speech Codec
von: Ren, Yanzhou, et al.
Veröffentlicht: (2026)
von: Ren, Yanzhou, et al.
Veröffentlicht: (2026)
Semantic Communications for Speech Recognition
von: Weng, Zhenzi, et al.
Veröffentlicht: (2021)
von: Weng, Zhenzi, et al.
Veröffentlicht: (2021)
HyBeam: Hybrid Microphone-Beamforming Array-Agnostic Speech Enhancement for Wearables
von: Ilan, Yuval Bar, et al.
Veröffentlicht: (2025)
von: Ilan, Yuval Bar, et al.
Veröffentlicht: (2025)
BanglaNum -- A Public Dataset for Bengali Digit Recognition from Speech
von: Mohammad, Mir Sayeed, et al.
Veröffentlicht: (2024)
von: Mohammad, Mir Sayeed, et al.
Veröffentlicht: (2024)
Zero-Bit Transmission of Adaptive Pre- and De-emphasis Filters for Speech and Audio Coding
von: Piralideh, Niloofar Omidi, et al.
Veröffentlicht: (2024)
von: Piralideh, Niloofar Omidi, et al.
Veröffentlicht: (2024)
Speak in the Scene: Diffusion-based Acoustic Scene Transfer toward Immersive Speech Generation
von: Kim, Miseul, et al.
Veröffentlicht: (2024)
von: Kim, Miseul, et al.
Veröffentlicht: (2024)
Advanced Signal Analysis in Detecting Replay Attacks for Automatic Speaker Verification Systems
von: Kuang, Lee Shih
Veröffentlicht: (2024)
von: Kuang, Lee Shih
Veröffentlicht: (2024)
Generative AI in Signal Processing Education: An Audio Foundation Model Based Approach
von: Khan, Muhammad Salman, et al.
Veröffentlicht: (2026)
von: Khan, Muhammad Salman, et al.
Veröffentlicht: (2026)
Speaker and Style Disentanglement of Speech Based on Contrastive Predictive Coding Supported Factorized Variational Autoencoder
von: Xie, Yuying, et al.
Veröffentlicht: (2024)
von: Xie, Yuying, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ERes2NetV2: Boosting Short-Duration Speaker Verification Performance with Computational Efficiency
von: Chen, Yafeng, et al.
Veröffentlicht: (2024) -
Microphone Array Signal Processing and Deep Learning for Speech Enhancement
von: Haeb-Umbach, Reinhold, et al.
Veröffentlicht: (2025) -
SpeechMLC: Speech Multi-label Classification
von: Kim, Miseul, et al.
Veröffentlicht: (2025) -
Rec-RIR: Monaural Blind Room Impulse Response Identification via DNN-based Reverberant Speech Reconstruction in STFT Domain
von: Wang, Pengyu, et al.
Veröffentlicht: (2025) -
ParaS2S: Benchmarking and Aligning Spoken Language Models for Paralinguistic-aware Speech-to-Speech Interaction
von: Yang, Shu-wen, et al.
Veröffentlicht: (2025)