Probing the Information Encoded in Neural-based Acoustic Models of Automatic Speech Recognition Systems
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Raymondaud, Quentin, Rouvier, Mickael, Dufour, Richard |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MSP-Podcast SER Challenge 2024: L'antenne du Ventoux Multimodal Self-Supervised Learning for Speech Emotion Recognition
von: Duret, Jarod, et al.
Veröffentlicht: (2024)
von: Duret, Jarod, et al.
Veröffentlicht: (2024)
Serialized Speech Information Guidance with Overlapped Encoding Separation for Multi-Speaker Automatic Speech Recognition
von: Shi, Hao, et al.
Veröffentlicht: (2024)
von: Shi, Hao, et al.
Veröffentlicht: (2024)
An Empirical Analysis of Discrete Unit Representations in Speech Language Modeling Pre-training
von: Labrak, Yanis, et al.
Veröffentlicht: (2025)
von: Labrak, Yanis, et al.
Veröffentlicht: (2025)
Speech Recognition-based Feature Extraction for Enhanced Automatic Severity Classification in Dysarthric Speech
von: Choi, Yerin, et al.
Veröffentlicht: (2024)
von: Choi, Yerin, et al.
Veröffentlicht: (2024)
Tiny-Align: Bridging Automatic Speech Recognition and Large Language Model on the Edge
von: Qin, Ruiyang, et al.
Veröffentlicht: (2024)
von: Qin, Ruiyang, et al.
Veröffentlicht: (2024)
Automatic Speech Recognition in the Modern Era: Architectures, Training, and Evaluation
von: Nayeem, Md., et al.
Veröffentlicht: (2025)
von: Nayeem, Md., et al.
Veröffentlicht: (2025)
LLMs-Integrated Automatic Hate Speech Recognition Using Controllable Text Generation Models
von: Oshima, Ryutaro, et al.
Veröffentlicht: (2026)
von: Oshima, Ryutaro, et al.
Veröffentlicht: (2026)
It's Never Too Late: Fusing Acoustic Information into Large Language Models for Automatic Speech Recognition
von: Chen, Chen, et al.
Veröffentlicht: (2024)
von: Chen, Chen, et al.
Veröffentlicht: (2024)
Explaining Deep Learning Embeddings for Speech Emotion Recognition by Predicting Interpretable Acoustic Features
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
GEC-RAG: Improving Generative Error Correction via Retrieval-Augmented Generation for Automatic Speech Recognition Systems
von: Robatian, Amin, et al.
Veröffentlicht: (2025)
von: Robatian, Amin, et al.
Veröffentlicht: (2025)
Do we really need Self-Attention for Streaming Automatic Speech Recognition?
von: Dkhissi, Youness, et al.
Veröffentlicht: (2026)
von: Dkhissi, Youness, et al.
Veröffentlicht: (2026)
ACES: Accent Subspaces for Coupling, Explanations, and Stress-Testing in Automatic Speech Recognition
von: Parekh, Swapnil
Veröffentlicht: (2026)
von: Parekh, Swapnil
Veröffentlicht: (2026)
HuBERT-VIC: Improving Noise-Robust Automatic Speech Recognition of Speech Foundation Model via Variance-Invariance-Covariance Regularization
von: Ahn, Hyebin, et al.
Veröffentlicht: (2025)
von: Ahn, Hyebin, et al.
Veröffentlicht: (2025)
Cross-Corpus Validation of Speech Emotion Recognition in Urdu using Domain-Knowledge Acoustic Features
von: Talpur, Unzela, et al.
Veröffentlicht: (2025)
von: Talpur, Unzela, et al.
Veröffentlicht: (2025)
ICMC-ASR: The ICASSP 2024 In-Car Multi-Channel Automatic Speech Recognition Challenge
von: Wang, He, et al.
Veröffentlicht: (2024)
von: Wang, He, et al.
Veröffentlicht: (2024)
Interpreting Pretrained Speech Models for Automatic Speech Assessment of Voice Disorders
von: Lau, Hok-Shing, et al.
Veröffentlicht: (2024)
von: Lau, Hok-Shing, et al.
Veröffentlicht: (2024)
AMNet: An Acoustic Model Network for Enhanced Mandarin Speech Synthesis
von: Cao, Yubing, et al.
Veröffentlicht: (2025)
von: Cao, Yubing, et al.
Veröffentlicht: (2025)
Zero-Shot End-To-End Spoken Question Answering In Medical Domain
von: Labrak, Yanis, et al.
Veröffentlicht: (2024)
von: Labrak, Yanis, et al.
Veröffentlicht: (2024)
Improved Contextual Recognition In Automatic Speech Recognition Systems By Semantic Lattice Rescoring
von: Sudarshan, Ankitha, et al.
Veröffentlicht: (2023)
von: Sudarshan, Ankitha, et al.
Veröffentlicht: (2023)
Color-based Emotion Representation for Speech Emotion Recognition
von: Nagase, Ryotaro, et al.
Veröffentlicht: (2026)
von: Nagase, Ryotaro, et al.
Veröffentlicht: (2026)
A Multi-task Learning Balanced Attention Convolutional Neural Network Model for Few-shot Underwater Acoustic Target Recognition
von: Huang, Wei, et al.
Veröffentlicht: (2025)
von: Huang, Wei, et al.
Veröffentlicht: (2025)
SCRAPS: Speech Contrastive Representations of Acoustic and Phonetic Spaces
von: Vallés-Pérez, Ivan, et al.
Veröffentlicht: (2023)
von: Vallés-Pérez, Ivan, et al.
Veröffentlicht: (2023)
ABHINAYA -- A System for Speech Emotion Recognition In Naturalistic Conditions Challenge
von: Dutta, Soumya, et al.
Veröffentlicht: (2025)
von: Dutta, Soumya, et al.
Veröffentlicht: (2025)
TinyML for Speech Recognition
von: Barovic, Andrew, et al.
Veröffentlicht: (2025)
von: Barovic, Andrew, et al.
Veröffentlicht: (2025)
Automatic Speech Recognition using Advanced Deep Learning Approaches: A survey
von: Kheddar, Hamza, et al.
Veröffentlicht: (2024)
von: Kheddar, Hamza, et al.
Veröffentlicht: (2024)
ES4R: Speech Encoding Based on Prepositive Affective Modeling for Empathetic Response Generation
von: Gao, Zhuoyue, et al.
Veröffentlicht: (2026)
von: Gao, Zhuoyue, et al.
Veröffentlicht: (2026)
Clustering and Mining Accented Speech for Inclusive and Fair Speech Recognition
von: Kim, Jaeyoung, et al.
Veröffentlicht: (2024)
von: Kim, Jaeyoung, et al.
Veröffentlicht: (2024)
AAT: Adapting Audio Transformer for Various Acoustics Recognition Tasks
von: Liang, Yun, et al.
Veröffentlicht: (2024)
von: Liang, Yun, et al.
Veröffentlicht: (2024)
Cross-Speaker Encoding Network for Multi-Talker Speech Recognition
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
Self-supervised ASR Models and Features For Dysarthric and Elderly Speech Recognition
von: Hu, Shujie, et al.
Veröffentlicht: (2024)
von: Hu, Shujie, et al.
Veröffentlicht: (2024)
The NPU-ASLP-LiAuto System Description for Visual Speech Recognition in CNVSRC 2023
von: Wang, He, et al.
Veröffentlicht: (2024)
von: Wang, He, et al.
Veröffentlicht: (2024)
Asymmetric and trial-dependent modeling: the contribution of LIA to SdSV Challenge Task 2
von: Bousquet, Pierre-Michel, et al.
Veröffentlicht: (2024)
von: Bousquet, Pierre-Michel, et al.
Veröffentlicht: (2024)
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition
von: Gao, Ming, et al.
Veröffentlicht: (2025)
von: Gao, Ming, et al.
Veröffentlicht: (2025)
XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs
von: Gong, Yitian, et al.
Veröffentlicht: (2025)
von: Gong, Yitian, et al.
Veröffentlicht: (2025)
When LLMs Meets Acoustic Landmarks: An Efficient Approach to Integrate Speech into Large Language Models for Depression Detection
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2024)
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2024)
Self-Supervised Models for Phoneme Recognition: Applications in Children's Speech for Reading Learning
von: Medin, Lucas Block, et al.
Veröffentlicht: (2025)
von: Medin, Lucas Block, et al.
Veröffentlicht: (2025)
MLCA-AVSR: Multi-Layer Cross Attention Fusion based Audio-Visual Speech Recognition
von: Wang, He, et al.
Veröffentlicht: (2024)
von: Wang, He, et al.
Veröffentlicht: (2024)
The Computation of Generalized Embeddings for Underwater Acoustic Target Recognition using Contrastive Learning
von: Hummel, Hilde I., et al.
Veröffentlicht: (2025)
von: Hummel, Hilde I., et al.
Veröffentlicht: (2025)
SPEAR: Receiver-to-Receiver Acoustic Neural Warping Field
von: He, Yuhang, et al.
Veröffentlicht: (2024)
von: He, Yuhang, et al.
Veröffentlicht: (2024)
Structure-informed Positional Encoding for Music Generation
von: Agarwal, Manvi, et al.
Veröffentlicht: (2024)
von: Agarwal, Manvi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MSP-Podcast SER Challenge 2024: L'antenne du Ventoux Multimodal Self-Supervised Learning for Speech Emotion Recognition
von: Duret, Jarod, et al.
Veröffentlicht: (2024) -
Serialized Speech Information Guidance with Overlapped Encoding Separation for Multi-Speaker Automatic Speech Recognition
von: Shi, Hao, et al.
Veröffentlicht: (2024) -
An Empirical Analysis of Discrete Unit Representations in Speech Language Modeling Pre-training
von: Labrak, Yanis, et al.
Veröffentlicht: (2025) -
Speech Recognition-based Feature Extraction for Enhanced Automatic Severity Classification in Dysarthric Speech
von: Choi, Yerin, et al.
Veröffentlicht: (2024) -
Tiny-Align: Bridging Automatic Speech Recognition and Large Language Model on the Edge
von: Qin, Ruiyang, et al.
Veröffentlicht: (2024)