Bayesian adaptive learning to latent variables via Variational Bayes and Maximum a Posteriori
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hu, Hu, Siniscalchi, Sabato Marco, Lee, Chin-Hui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Variational Bayesian Adaptive Learning of Deep Latent Variables for Acoustic Knowledge Transfer
von: Hu, Hu, et al.
Veröffentlicht: (2025)
von: Hu, Hu, et al.
Veröffentlicht: (2025)
An Explicit Consistency-Preserving Loss Function for Phase Reconstruction and Speech Enhancement
von: Ku, Pin-Jui, et al.
Veröffentlicht: (2024)
von: Ku, Pin-Jui, et al.
Veröffentlicht: (2024)
Language-Universal Speech Attributes Modeling for Zero-Shot Multilingual Spoken Keyword Recognition
von: Yen, Hao, et al.
Veröffentlicht: (2024)
von: Yen, Hao, et al.
Veröffentlicht: (2024)
voc2vec: A Foundation Model for Non-Verbal Vocalization
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
A Study on Zero-shot Non-intrusive Speech Assessment using Large Language Models
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2024)
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2024)
Feature Importance across Domains for Improving Non-Intrusive Speech Intelligibility Prediction in Hearing Aids
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2025)
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2025)
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition
von: Gao, Ming, et al.
Veröffentlicht: (2025)
von: Gao, Ming, et al.
Veröffentlicht: (2025)
Exploiting Consistency-Preserving Loss and Perceptual Contrast Stretching to Boost SSL-based Speech Enhancement
von: Khan, Muhammad Salman, et al.
Veröffentlicht: (2024)
von: Khan, Muhammad Salman, et al.
Veröffentlicht: (2024)
Towards Robust Assessment of Pathological Voices via Combined Low-Level Descriptors and Foundation Model Representations
von: Ariyanti, Whenty, et al.
Veröffentlicht: (2025)
von: Ariyanti, Whenty, et al.
Veröffentlicht: (2025)
Few-Shot and Pseudo-Label Guided Speech Quality Evaluation with Large Language Models
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2026)
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2026)
Hallucination Benchmark for Speech Foundation Models
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
Benchmarking Representations for Speech, Music, and Acoustic Events
von: La Quatra, Moreno, et al.
Veröffentlicht: (2024)
von: La Quatra, Moreno, et al.
Veröffentlicht: (2024)
A Knowledge-Driven Approach to Target Speech Extraction in the Presence of Background Sound Effects for Cinematic Audio Source Separation (CASS)
von: Ho, Chun-wei, et al.
Veröffentlicht: (2026)
von: Ho, Chun-wei, et al.
Veröffentlicht: (2026)
A Bottom-up Framework with Language-universal Speech Attribute Modeling for Syllable-based ASR
von: Yen, Hao, et al.
Veröffentlicht: (2025)
von: Yen, Hao, et al.
Veröffentlicht: (2025)
FlanEC: Exploring Flan-T5 for Post-ASR Error Correction
von: La Quatra, Moreno, et al.
Veröffentlicht: (2025)
von: La Quatra, Moreno, et al.
Veröffentlicht: (2025)
An Investigation of Incorporating Mamba for Speech Enhancement
von: Chao, Rong, et al.
Veröffentlicht: (2024)
von: Chao, Rong, et al.
Veröffentlicht: (2024)
Modeling strategies for speech enhancement in the latent space of a neural audio codec
von: Kammoun, Sofiene, et al.
Veröffentlicht: (2025)
von: Kammoun, Sofiene, et al.
Veröffentlicht: (2025)
It's Never Too Late: Fusing Acoustic Information into Large Language Models for Automatic Speech Recognition
von: Chen, Chen, et al.
Veröffentlicht: (2024)
von: Chen, Chen, et al.
Veröffentlicht: (2024)
Advancing Speech Quality Assessment Through Scientific Challenges and Open-source Activities
von: Huang, Wen-Chin
Veröffentlicht: (2025)
von: Huang, Wen-Chin
Veröffentlicht: (2025)
Improving fairness in speaker verification via Group-adapted Fusion Network
von: Shen, Hua, et al.
Veröffentlicht: (2022)
von: Shen, Hua, et al.
Veröffentlicht: (2022)
Conditioning and Sampling in Variational Diffusion Models for Speech Super-Resolution
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2022)
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2022)
Representational learning for an anomalous sound detection system with source separation model
von: Shin, Seunghyeon, et al.
Veröffentlicht: (2024)
von: Shin, Seunghyeon, et al.
Veröffentlicht: (2024)
Enhancing Spectrogram Realism in Singing Voice Synthesis via Explicit Bandwidth Extension Prior to Vocoder
von: Yang, Runxuan, et al.
Veröffentlicht: (2025)
von: Yang, Runxuan, et al.
Veröffentlicht: (2025)
Addressing Index Collapse of Large-Codebook Speech Tokenizer with Dual-Decoding Product-Quantized Variational Auto-Encoder
von: Guo, Haohan, et al.
Veröffentlicht: (2024)
von: Guo, Haohan, et al.
Veröffentlicht: (2024)
Differentiable Time-Varying Linear Prediction in the Context of End-to-End Analysis-by-Synthesis
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2024)
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2024)
Singing Voice Synthesis Using Differentiable LPC and Glottal-Flow-Inspired Wavetables
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2023)
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2023)
Period Singer: Integrating Periodic and Aperiodic Variational Autoencoders for Natural-Sounding End-to-End Singing Voice Synthesis
von: Kim, Taewoo, et al.
Veröffentlicht: (2024)
von: Kim, Taewoo, et al.
Veröffentlicht: (2024)
An Investigation on Combining Geometry and Consistency Constraints into Phase Estimation for Speech Enhancement
von: Ho, Chun-Wei, et al.
Veröffentlicht: (2025)
von: Ho, Chun-Wei, et al.
Veröffentlicht: (2025)
RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval
von: Sun, Haoqin, et al.
Veröffentlicht: (2025)
von: Sun, Haoqin, et al.
Veröffentlicht: (2025)
Sample adaptive data augmentation with progressive scheduling
von: Lu, Hongxuan, et al.
Veröffentlicht: (2024)
von: Lu, Hongxuan, et al.
Veröffentlicht: (2024)
STSR: High-Fidelity Speech Super-Resolution via Spectral-Transient Context Modeling
von: Yuan, Jiajun, et al.
Veröffentlicht: (2025)
von: Yuan, Jiajun, et al.
Veröffentlicht: (2025)
Diversifying and Expanding Frequency-Adaptive Convolution Kernels for Sound Event Detection
von: Nam, Hyeonuk, et al.
Veröffentlicht: (2024)
von: Nam, Hyeonuk, et al.
Veröffentlicht: (2024)
Monaural speech enhancement on drone via Adapter based transfer learning
von: Chen, Xingyu, et al.
Veröffentlicht: (2024)
von: Chen, Xingyu, et al.
Veröffentlicht: (2024)
MOS-Bench: Benchmarking Generalization Abilities of Subjective Speech Quality Assessment Models
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2024)
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2024)
SHEET: A Multi-purpose Open-source Speech Human Evaluation Estimation Toolkit
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2025)
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2025)
CodecMOS-Accent: A MOS Benchmark of Resynthesized and TTS Speech from Neural Codecs Across English Accents
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2026)
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2026)
Quality-Aware End-to-End Audio-Visual Neural Speaker Diarization
von: He, Mao-Kui, et al.
Veröffentlicht: (2024)
von: He, Mao-Kui, et al.
Veröffentlicht: (2024)
GenSE: Generative Speech Enhancement via Language Models using Hierarchical Modeling
von: Yao, Jixun, et al.
Veröffentlicht: (2025)
von: Yao, Jixun, et al.
Veröffentlicht: (2025)
Interactive singing melody extraction based on active adaptation
von: Saxena, Kavya Ranjan, et al.
Veröffentlicht: (2024)
von: Saxena, Kavya Ranjan, et al.
Veröffentlicht: (2024)
A Semantic Information-based Hierarchical Speech Enhancement Method Using Factorized Codec and Diffusion Model
von: Xiang, Yang, et al.
Veröffentlicht: (2025)
von: Xiang, Yang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Variational Bayesian Adaptive Learning of Deep Latent Variables for Acoustic Knowledge Transfer
von: Hu, Hu, et al.
Veröffentlicht: (2025) -
An Explicit Consistency-Preserving Loss Function for Phase Reconstruction and Speech Enhancement
von: Ku, Pin-Jui, et al.
Veröffentlicht: (2024) -
Language-Universal Speech Attributes Modeling for Zero-Shot Multilingual Spoken Keyword Recognition
von: Yen, Hao, et al.
Veröffentlicht: (2024) -
voc2vec: A Foundation Model for Non-Verbal Vocalization
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025) -
A Study on Zero-shot Non-intrusive Speech Assessment using Large Language Models
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2024)