Zero-Bit Transmission of Adaptive Pre- and De-emphasis Filters for Speech and Audio Coding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Piralideh, Niloofar Omidi, Gournay, Philippe, Lefebvre, Roch |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
USDnet: Unsupervised Speech Dereverberation via Neural Forward Filtering
von: Wang, Zhong-Qiu
Veröffentlicht: (2024)
von: Wang, Zhong-Qiu
Veröffentlicht: (2024)
Robust Fixed-Filter Sound Zone Control with Audio-Based Position Tracking
von: Bhattacharjee, Sankha Subhra, et al.
Veröffentlicht: (2024)
von: Bhattacharjee, Sankha Subhra, et al.
Veröffentlicht: (2024)
Interpolation Filter Design for Sample Rate Independent Audio Effect RNNs
von: Carson, Alistair, et al.
Veröffentlicht: (2024)
von: Carson, Alistair, et al.
Veröffentlicht: (2024)
Multi-channel Replay Speech Detection using an Adaptive Learnable Beamformer
von: Neri, Michael, et al.
Veröffentlicht: (2025)
von: Neri, Michael, et al.
Veröffentlicht: (2025)
DeFTAN-II: Efficient Multichannel Speech Enhancement with Subgroup Processing
von: Lee, Dongheon, et al.
Veröffentlicht: (2023)
von: Lee, Dongheon, et al.
Veröffentlicht: (2023)
Significance of Chirp MFCC as a Feature in Speech and Audio Applications
von: Joysingh, S. Johanan, et al.
Veröffentlicht: (2024)
von: Joysingh, S. Johanan, et al.
Veröffentlicht: (2024)
Speaker and Style Disentanglement of Speech Based on Contrastive Predictive Coding Supported Factorized Variational Autoencoder
von: Xie, Yuying, et al.
Veröffentlicht: (2024)
von: Xie, Yuying, et al.
Veröffentlicht: (2024)
Bridging the Gap: Integrating Pre-trained Speech Enhancement and Recognition Models for Robust Speech Recognition
von: Wang, Kuan-Chen, et al.
Veröffentlicht: (2024)
von: Wang, Kuan-Chen, et al.
Veröffentlicht: (2024)
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding
von: Wang, Ziqian, et al.
Veröffentlicht: (2025)
von: Wang, Ziqian, et al.
Veröffentlicht: (2025)
Neural Speech and Audio Coding: Modern AI Technology Meets Traditional Codecs
von: Kim, Minje, et al.
Veröffentlicht: (2024)
von: Kim, Minje, et al.
Veröffentlicht: (2024)
SpeechMLC: Speech Multi-label Classification
von: Kim, Miseul, et al.
Veröffentlicht: (2025)
von: Kim, Miseul, et al.
Veröffentlicht: (2025)
A DNN Based Post-Filter to Enhance the Quality of Coded Speech in MDCT Domain
von: Gupta, Kishan, et al.
Veröffentlicht: (2022)
von: Gupta, Kishan, et al.
Veröffentlicht: (2022)
Cyclic Multichannel Wiener Filter for Acoustic Beamforming
von: Bologni, Giovanni, et al.
Veröffentlicht: (2025)
von: Bologni, Giovanni, et al.
Veröffentlicht: (2025)
Neural Spectral Band Generation for Audio Coding
von: Choi, Woongjib, et al.
Veröffentlicht: (2025)
von: Choi, Woongjib, et al.
Veröffentlicht: (2025)
Adaptive Per-Channel Energy Normalization Front-end for Robust Audio Signal Processing
von: Meng, Hanyu, et al.
Veröffentlicht: (2025)
von: Meng, Hanyu, et al.
Veröffentlicht: (2025)
A Speech Production Model for Radar: Connecting Speech Acoustics with Radar-Measured Vibrations
von: Lenz, Isabella, et al.
Veröffentlicht: (2025)
von: Lenz, Isabella, et al.
Veröffentlicht: (2025)
Detection of manatee vocalisations using the Audio Spectrogram Transformer
von: Schiappacasse, Stefano, et al.
Veröffentlicht: (2024)
von: Schiappacasse, Stefano, et al.
Veröffentlicht: (2024)
SUNAC: Source-aware Unified Neural Audio Codec
von: Aihara, Ryo, et al.
Veröffentlicht: (2025)
von: Aihara, Ryo, et al.
Veröffentlicht: (2025)
ParaS2S: Benchmarking and Aligning Spoken Language Models for Paralinguistic-aware Speech-to-Speech Interaction
von: Yang, Shu-wen, et al.
Veröffentlicht: (2025)
von: Yang, Shu-wen, et al.
Veröffentlicht: (2025)
Binaural Localization Model for Speech in Noise
von: Tokala, Vikas, et al.
Veröffentlicht: (2025)
von: Tokala, Vikas, et al.
Veröffentlicht: (2025)
Speech-Based Prioritization for Schizophrenia Intervention
von: Premananth, Gowtham, et al.
Veröffentlicht: (2025)
von: Premananth, Gowtham, et al.
Veröffentlicht: (2025)
Prompt-driven Target Speech Diarization
von: Jiang, Yidi, et al.
Veröffentlicht: (2023)
von: Jiang, Yidi, et al.
Veröffentlicht: (2023)
Perceptually-motivated Spatial Audio Codec for Higher-Order Ambisonics Compression
von: Hold, Christoph, et al.
Veröffentlicht: (2024)
von: Hold, Christoph, et al.
Veröffentlicht: (2024)
Benchmarking Audio Deepfake Detection Robustness in Real-world Communication Scenarios
von: Shi, Haohan, et al.
Veröffentlicht: (2025)
von: Shi, Haohan, et al.
Veröffentlicht: (2025)
Brain-Informed Speech Separation for Cochlear Implants
von: Gajecki, Tom, et al.
Veröffentlicht: (2026)
von: Gajecki, Tom, et al.
Veröffentlicht: (2026)
Speech Enhancement based on cascaded two flows
von: Lee, Seonggyu, et al.
Veröffentlicht: (2025)
von: Lee, Seonggyu, et al.
Veröffentlicht: (2025)
Generative AI in Signal Processing Education: An Audio Foundation Model Based Approach
von: Khan, Muhammad Salman, et al.
Veröffentlicht: (2026)
von: Khan, Muhammad Salman, et al.
Veröffentlicht: (2026)
FlowSE: Flow Matching-based Speech Enhancement
von: Lee, Seonggyu, et al.
Veröffentlicht: (2025)
von: Lee, Seonggyu, et al.
Veröffentlicht: (2025)
TTSlow: Slow Down Text-to-Speech with Efficiency Robustness Evaluations
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2024)
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2024)
Harmonics to the Rescue: Why Voiced Speech is Not a Wss Process
von: Bologni, Giovanni, et al.
Veröffentlicht: (2025)
von: Bologni, Giovanni, et al.
Veröffentlicht: (2025)
SELM: Speech Enhancement Using Discrete Tokens and Language Models
von: Wang, Ziqian, et al.
Veröffentlicht: (2023)
von: Wang, Ziqian, et al.
Veröffentlicht: (2023)
Binaural Speech Enhancement Using Complex Convolutional Recurrent Networks
von: Tokala, Vikas, et al.
Veröffentlicht: (2025)
von: Tokala, Vikas, et al.
Veröffentlicht: (2025)
Analyzing the Impact of Accent on English Speech: Acoustic and Articulatory Perspectives
von: Premananth, Gowtham, et al.
Veröffentlicht: (2025)
von: Premananth, Gowtham, et al.
Veröffentlicht: (2025)
The Overview of Segmental Durations Modification Algorithms on Speech Signal Characteristics
von: Jang, Kyeomeun, et al.
Veröffentlicht: (2025)
von: Jang, Kyeomeun, et al.
Veröffentlicht: (2025)
A Generalized Weighted Overlap-Add (WOLA) Filter Bank for Improved Subband System Identification
von: Sharma, Mohit, et al.
Veröffentlicht: (2025)
von: Sharma, Mohit, et al.
Veröffentlicht: (2025)
Listenable Maps for Zero-Shot Audio Classifiers
von: Paissan, Francesco, et al.
Veröffentlicht: (2024)
von: Paissan, Francesco, et al.
Veröffentlicht: (2024)
Exploring Audio-Visual Information Fusion for Sound Event Localization and Detection In Low-Resource Realistic Scenarios
von: Jiang, Ya, et al.
Veröffentlicht: (2024)
von: Jiang, Ya, et al.
Veröffentlicht: (2024)
Parameter-Efficient Fine-Tuning of Foundation Models for CLP Speech Classification
von: Bhattacharjee, Susmita, et al.
Veröffentlicht: (2025)
von: Bhattacharjee, Susmita, et al.
Veröffentlicht: (2025)
Exploring Disentangled Neural Speech Codecs from Self-Supervised Representations
von: Aihara, Ryo, et al.
Veröffentlicht: (2025)
von: Aihara, Ryo, et al.
Veröffentlicht: (2025)
Impact of Microphone Array Mismatches to Learning-based Replay Speech Detection
von: Neri, Michael, et al.
Veröffentlicht: (2025)
von: Neri, Michael, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
USDnet: Unsupervised Speech Dereverberation via Neural Forward Filtering
von: Wang, Zhong-Qiu
Veröffentlicht: (2024) -
Robust Fixed-Filter Sound Zone Control with Audio-Based Position Tracking
von: Bhattacharjee, Sankha Subhra, et al.
Veröffentlicht: (2024) -
Interpolation Filter Design for Sample Rate Independent Audio Effect RNNs
von: Carson, Alistair, et al.
Veröffentlicht: (2024) -
Multi-channel Replay Speech Detection using an Adaptive Learnable Beamformer
von: Neri, Michael, et al.
Veröffentlicht: (2025) -
DeFTAN-II: Efficient Multichannel Speech Enhancement with Subgroup Processing
von: Lee, Dongheon, et al.
Veröffentlicht: (2023)