On Calibration of Speech Classification Models: Insights from Energy-Based Model Investigations
Fuente:
arXiv
Salvato in:
| Autori principali: | Hao, Yaqian, Hu, Chenguang, Gao, Yingying, Zhang, Shilei, Feng, Junlan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
CEC: A Noisy Label Detection Method for Speaker Recognition
di: Shen, Yao, et al.
Pubblicazione: (2024)
di: Shen, Yao, et al.
Pubblicazione: (2024)
Exploring Energy-Based Models for Out-of-Distribution Detection in Dialect Identification
di: Hao, Yaqian, et al.
Pubblicazione: (2024)
di: Hao, Yaqian, et al.
Pubblicazione: (2024)
Plugin Speech Enhancement: A Universal Speech Enhancement Framework Inspired by Dynamic Neural Network
di: Chen, Yanan, et al.
Pubblicazione: (2024)
di: Chen, Yanan, et al.
Pubblicazione: (2024)
GenDistiller: Distilling Pre-trained Language Models based on an Autoregressive Generative Model
di: Gao, Yingying, et al.
Pubblicazione: (2024)
di: Gao, Yingying, et al.
Pubblicazione: (2024)
HarmoniFuse: A Component-Selective and Prompt-Adaptive Framework for Multi-Task Speech Language Modeling
di: Si, Yuke, et al.
Pubblicazione: (2025)
di: Si, Yuke, et al.
Pubblicazione: (2025)
PolySpeech: Exploring Unified Multitask Speech Models for Competitiveness with Single-task Models
di: Yang, Runyan, et al.
Pubblicazione: (2024)
di: Yang, Runyan, et al.
Pubblicazione: (2024)
MFSN: Multi-perspective Fusion Search Network For Pre-training Knowledge in Speech Emotion Recognition
di: Sun, Haiyang, et al.
Pubblicazione: (2023)
di: Sun, Haiyang, et al.
Pubblicazione: (2023)
Teaching Audio Models to Reason: A Unified Framework for Source- and Layer-wise Distillation
di: Yang, Runyan, et al.
Pubblicazione: (2025)
di: Yang, Runyan, et al.
Pubblicazione: (2025)
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion
di: Wang, Zhichao, et al.
Pubblicazione: (2026)
di: Wang, Zhichao, et al.
Pubblicazione: (2026)
Investigating Neural Audio Codecs for Speech Language Model-Based Speech Generation
di: Li, Jiaqi, et al.
Pubblicazione: (2024)
di: Li, Jiaqi, et al.
Pubblicazione: (2024)
DiffStyleTTS: Diffusion-based Hierarchical Prosody Modeling for Text-to-Speech with Diverse and Controllable Styles
di: Liu, Jiaxuan, et al.
Pubblicazione: (2024)
di: Liu, Jiaxuan, et al.
Pubblicazione: (2024)
B-GRPO: Unsupervised Speech Emotion Recognition based on Batched-Group Relative Policy Optimization
di: Gao, Yingying, et al.
Pubblicazione: (2026)
di: Gao, Yingying, et al.
Pubblicazione: (2026)
SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models
di: Zhang, Xin, et al.
Pubblicazione: (2023)
di: Zhang, Xin, et al.
Pubblicazione: (2023)
Investigating Polyglot Speech Foundation Models for Learning Collective Emotion from Crowds
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2025)
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2025)
Speech-Mamba: Long-Context Speech Recognition with Selective State Spaces Models
di: Gao, Xiaoxue, et al.
Pubblicazione: (2024)
di: Gao, Xiaoxue, et al.
Pubblicazione: (2024)
Vector Quantized Diffusion Model Based Speech Bandwidth Extension
di: Fang, Yuan, et al.
Pubblicazione: (2024)
di: Fang, Yuan, et al.
Pubblicazione: (2024)
Audio-Visual Representation Learning via Knowledge Distillation from Speech Foundation Models
di: Zhang, Jing-Xuan, et al.
Pubblicazione: (2025)
di: Zhang, Jing-Xuan, et al.
Pubblicazione: (2025)
Adaptive Speech Emotion Representation Learning Based On Dynamic Graph
di: Gao, Yingxue, et al.
Pubblicazione: (2024)
di: Gao, Yingxue, et al.
Pubblicazione: (2024)
FireRedASR: Open-Source Industrial-Grade Mandarin Speech Recognition Models from Encoder-Decoder to LLM Integration
di: Xu, Kai-Tuo, et al.
Pubblicazione: (2025)
di: Xu, Kai-Tuo, et al.
Pubblicazione: (2025)
Attention-Based Beamformer For Multi-Channel Speech Enhancement
di: Bai, Jinglin, et al.
Pubblicazione: (2024)
di: Bai, Jinglin, et al.
Pubblicazione: (2024)
Enabling Beam Search for Language Model-Based Text-to-Speech Synthesis
di: Tu, Zehai, et al.
Pubblicazione: (2024)
di: Tu, Zehai, et al.
Pubblicazione: (2024)
Leveraging Local and Global Knowledge Integration with Time-Frequency Calibrated Distillation for Speech Enhancement
di: Cheng, Jiaming, et al.
Pubblicazione: (2025)
di: Cheng, Jiaming, et al.
Pubblicazione: (2025)
Streaming Speech Recognition with Decoder-Only Large Language Models and Latency Optimization
di: Wan, Genshun, et al.
Pubblicazione: (2026)
di: Wan, Genshun, et al.
Pubblicazione: (2026)
Generative Speech Foundation Model Pretraining for High-Quality Speech Extraction and Restoration
di: Ku, Pin-Jui, et al.
Pubblicazione: (2024)
di: Ku, Pin-Jui, et al.
Pubblicazione: (2024)
PEFT-SER: On the Use of Parameter Efficient Transfer Learning Approaches For Speech Emotion Recognition Using Pre-trained Speech Models
di: Feng, Tiantian, et al.
Pubblicazione: (2023)
di: Feng, Tiantian, et al.
Pubblicazione: (2023)
Magnetoencephalography (MEG) Based Non-Invasive Chinese Speech Decoding
di: Jia, Zhihong, et al.
Pubblicazione: (2025)
di: Jia, Zhihong, et al.
Pubblicazione: (2025)
CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models
di: Chen, Junyang, et al.
Pubblicazione: (2026)
di: Chen, Junyang, et al.
Pubblicazione: (2026)
Speech Denoising with Auditory Models
di: Saddler, Mark R., et al.
Pubblicazione: (2020)
di: Saddler, Mark R., et al.
Pubblicazione: (2020)
SECodec: Structural Entropy-based Compressive Speech Representation Codec for Speech Language Models
di: Wang, Linqin, et al.
Pubblicazione: (2024)
di: Wang, Linqin, et al.
Pubblicazione: (2024)
A Semantic Information-based Hierarchical Speech Enhancement Method Using Factorized Codec and Diffusion Model
di: Xiang, Yang, et al.
Pubblicazione: (2025)
di: Xiang, Yang, et al.
Pubblicazione: (2025)
Distinguishing Neural Speech Synthesis Models Through Fingerprints in Speech Waveforms
di: Zhang, Chu Yuan, et al.
Pubblicazione: (2023)
di: Zhang, Chu Yuan, et al.
Pubblicazione: (2023)
Adaptive Convolution for CNN-based Speech Enhancement Models
di: Wang, Dahan, et al.
Pubblicazione: (2025)
di: Wang, Dahan, et al.
Pubblicazione: (2025)
Scale This, Not That: Investigating Key Dataset Attributes for Efficient Speech Enhancement Scaling
di: Zhang, Leying, et al.
Pubblicazione: (2024)
di: Zhang, Leying, et al.
Pubblicazione: (2024)
Speaker Disentanglement of Speech Pre-trained Model Based on Interpretability
di: Zhu, Xiaoxu, et al.
Pubblicazione: (2025)
di: Zhu, Xiaoxu, et al.
Pubblicazione: (2025)
Investigation of Deep Neural Network Acoustic Modelling Approaches for Low Resource Accented Mandarin Speech Recognition
di: Xie, Xurong, et al.
Pubblicazione: (2022)
di: Xie, Xurong, et al.
Pubblicazione: (2022)
Comprehend and Talk: Text to Speech Synthesis via Dual Language Modeling
di: Cao, Junjie, et al.
Pubblicazione: (2025)
di: Cao, Junjie, et al.
Pubblicazione: (2025)
HAM-TTS: Hierarchical Acoustic Modeling for Token-Based Zero-Shot Text-to-Speech with Model and Data Scaling
di: Wang, Chunhui, et al.
Pubblicazione: (2024)
di: Wang, Chunhui, et al.
Pubblicazione: (2024)
WenetSpeech-Wu: Datasets, Benchmarks, and Models for a Unified Chinese Wu Dialect Speech Processing Ecosystem
di: Wang, Chengyou, et al.
Pubblicazione: (2026)
di: Wang, Chengyou, et al.
Pubblicazione: (2026)
GenSE: Generative Speech Enhancement via Language Models using Hierarchical Modeling
di: Yao, Jixun, et al.
Pubblicazione: (2025)
di: Yao, Jixun, et al.
Pubblicazione: (2025)
Evaluating Self-Supervised Speech Models via Text-Based LLMS
di: Maekaku, Takashi, et al.
Pubblicazione: (2025)
di: Maekaku, Takashi, et al.
Pubblicazione: (2025)
Documenti analoghi
-
CEC: A Noisy Label Detection Method for Speaker Recognition
di: Shen, Yao, et al.
Pubblicazione: (2024) -
Exploring Energy-Based Models for Out-of-Distribution Detection in Dialect Identification
di: Hao, Yaqian, et al.
Pubblicazione: (2024) -
Plugin Speech Enhancement: A Universal Speech Enhancement Framework Inspired by Dynamic Neural Network
di: Chen, Yanan, et al.
Pubblicazione: (2024) -
GenDistiller: Distilling Pre-trained Language Models based on an Autoregressive Generative Model
di: Gao, Yingying, et al.
Pubblicazione: (2024) -
HarmoniFuse: A Component-Selective and Prompt-Adaptive Framework for Multi-Task Speech Language Modeling
di: Si, Yuke, et al.
Pubblicazione: (2025)