Enhancing Neural Spoken Language Recognition: An Exploration with Multilingual Datasets
Fuente:
arXiv
Guardado en:
| Autores principales: | Anidjar, Or Haim, Yozevitch, Roi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Whisper Turns Stronger: Augmenting Wav2Vec 2.0 for Superior ASR in Low-Resource Languages
por: Anidjar, Or Haim, et al.
Publicado: (2024)
por: Anidjar, Or Haim, et al.
Publicado: (2024)
Vocal Tract Length Warped Features for Spoken Keyword Spotting
por: Sarkar, Achintya kr., et al.
Publicado: (2025)
por: Sarkar, Achintya kr., et al.
Publicado: (2025)
Deep Neural Network for Musical Instrument Recognition using MFCCs
por: Mahanta, Saranga Kingkor, et al.
Publicado: (2021)
por: Mahanta, Saranga Kingkor, et al.
Publicado: (2021)
Spoken Language Intelligence of Large Language Models for Language Learning
por: Peng, Linkai, et al.
Publicado: (2023)
por: Peng, Linkai, et al.
Publicado: (2023)
Focal Loss based Residual Convolutional Neural Network for Speech Emotion Recognition
por: Tripathi, Suraj, et al.
Publicado: (2019)
por: Tripathi, Suraj, et al.
Publicado: (2019)
Enhancing Automatic Speech Recognition Through Integrated Noise Detection Architecture
por: Singh, Karamvir
Publicado: (2025)
por: Singh, Karamvir
Publicado: (2025)
WavReward: Spoken Dialogue Models With Generalist Reward Evaluators
por: Ji, Shengpeng, et al.
Publicado: (2025)
por: Ji, Shengpeng, et al.
Publicado: (2025)
Scaling Ambiguity: Augmenting Human Annotation in Speech Emotion Recognition with Audio-Language Models
por: Zhang, Wenda, et al.
Publicado: (2026)
por: Zhang, Wenda, et al.
Publicado: (2026)
Towards Speaker Identification with Minimal Dataset and Constrained Resources using 1D-Convolution Neural Network
por: Shahan, Irfan Nafiz, et al.
Publicado: (2024)
por: Shahan, Irfan Nafiz, et al.
Publicado: (2024)
Enhanced Speech Emotion Recognition with Efficient Channel Attention Guided Deep CNN-BiLSTM Framework
por: Kundu, Niloy Kumar, et al.
Publicado: (2024)
por: Kundu, Niloy Kumar, et al.
Publicado: (2024)
A Variational Framework for Improving Naturalness in Generative Spoken Language Models
por: Chen, Li-Wei, et al.
Publicado: (2025)
por: Chen, Li-Wei, et al.
Publicado: (2025)
SpokeN-100: A Cross-Lingual Benchmarking Dataset for The Classification of Spoken Numbers in Different Languages
por: Groh, René, et al.
Publicado: (2024)
por: Groh, René, et al.
Publicado: (2024)
HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling
por: Xue, Rongkun, et al.
Publicado: (2025)
por: Xue, Rongkun, et al.
Publicado: (2025)
Robust COVID-19 Detection from Cough Sounds using Deep Neural Decision Tree and Forest: A Comprehensive Cross-Datasets Evaluation
por: Islam, Rofiqul, et al.
Publicado: (2025)
por: Islam, Rofiqul, et al.
Publicado: (2025)
Analyzing Mitigation Strategies for Catastrophic Forgetting in End-to-End Training of Spoken Language Models
por: Hsiao, Chi-Yuan, et al.
Publicado: (2025)
por: Hsiao, Chi-Yuan, et al.
Publicado: (2025)
Bridging Modalities: Knowledge Distillation and Masked Training for Translating Multi-Modal Emotion Recognition to Uni-Modal, Speech-Only Emotion Recognition
por: Muaz, Muhammad, et al.
Publicado: (2024)
por: Muaz, Muhammad, et al.
Publicado: (2024)
Robust Cross-Etiology and Speaker-Independent Dysarthric Speech Recognition
por: Singh, Satwinder, et al.
Publicado: (2025)
por: Singh, Satwinder, et al.
Publicado: (2025)
LiteASR: Efficient Automatic Speech Recognition with Low-Rank Approximation
por: Kamahori, Keisuke, et al.
Publicado: (2025)
por: Kamahori, Keisuke, et al.
Publicado: (2025)
WhisperD: Dementia Speech Recognition and Filler Word Detection with Whisper
por: Akinrintoyo, Emmanuel, et al.
Publicado: (2025)
por: Akinrintoyo, Emmanuel, et al.
Publicado: (2025)
Epic-Sounds: A Large-scale Dataset of Actions That Sound
por: Huh, Jaesung, et al.
Publicado: (2023)
por: Huh, Jaesung, et al.
Publicado: (2023)
Speech Command Recognition Using LogNNet Reservoir Computing for Embedded Systems
por: Izotov, Yuriy, et al.
Publicado: (2025)
por: Izotov, Yuriy, et al.
Publicado: (2025)
$\texttt{AVROBUSTBENCH}$: Benchmarking the Robustness of Audio-Visual Recognition Models at Test-Time
por: Maharana, Sarthak Kumar, et al.
Publicado: (2025)
por: Maharana, Sarthak Kumar, et al.
Publicado: (2025)
Speech Emotion Recognition Using CNN and Its Use Case in Digital Healthcare
por: Nigar, Nishargo
Publicado: (2024)
por: Nigar, Nishargo
Publicado: (2024)
The Unreliability of Acoustic Systems in Alzheimer's Speech Datasets with Heterogeneous Recording Conditions
por: Gauder, Lara, et al.
Publicado: (2024)
por: Gauder, Lara, et al.
Publicado: (2024)
CloserMusicDB: A Modern Multipurpose Dataset of High Quality Music
por: Piekarzewicz, Aleksandra, et al.
Publicado: (2024)
por: Piekarzewicz, Aleksandra, et al.
Publicado: (2024)
Towards Unified Neural Decoding of Perceived, Spoken and Imagined Speech from EEG Signals
por: Lee, Jung-Sun, et al.
Publicado: (2024)
por: Lee, Jung-Sun, et al.
Publicado: (2024)
Designing Neural Synthesizers for Low-Latency Interaction
por: Caspe, Franco, et al.
Publicado: (2025)
por: Caspe, Franco, et al.
Publicado: (2025)
Learning Source Disentanglement in Neural Audio Codec
por: Bie, Xiaoyu, et al.
Publicado: (2024)
por: Bie, Xiaoyu, et al.
Publicado: (2024)
The NES Video-Music Database: A Dataset of Symbolic Video Game Music Paired with Gameplay Videos
por: Cardoso, Igor, et al.
Publicado: (2024)
por: Cardoso, Igor, et al.
Publicado: (2024)
Speaker Emotion Recognition: Leveraging Self-Supervised Models for Feature Extraction Using Wav2Vec2 and HuBERT
por: Jafarzadeh, Pourya, et al.
Publicado: (2024)
por: Jafarzadeh, Pourya, et al.
Publicado: (2024)
Speech Enhancement Using Continuous Embeddings of Neural Audio Codec
por: Li, Haoyang, et al.
Publicado: (2025)
por: Li, Haoyang, et al.
Publicado: (2025)
LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs
por: Mousavi, Pooneh, et al.
Publicado: (2025)
por: Mousavi, Pooneh, et al.
Publicado: (2025)
Run-Time Adaptation of Neural Beamforming for Robust Speech Dereverberation and Denoising
por: Fujita, Yoto, et al.
Publicado: (2024)
por: Fujita, Yoto, et al.
Publicado: (2024)
Lightweight Hopfield Neural Networks for Bioacoustic Detection and Call Monitoring of Captive Primates
por: Lomas, Wendy, et al.
Publicado: (2025)
por: Lomas, Wendy, et al.
Publicado: (2025)
PSLM: Parallel Generation of Text and Speech with LLMs for Low-Latency Spoken Dialogue Systems
por: Mitsui, Kentaro, et al.
Publicado: (2024)
por: Mitsui, Kentaro, et al.
Publicado: (2024)
SHAMaNS: Sound Localization with Hybrid Alpha-Stable Spatial Measure and Neural Steerer
por: Di Carlo, Diego, et al.
Publicado: (2025)
por: Di Carlo, Diego, et al.
Publicado: (2025)
JEPA as a Neural Tokenizer: Learning Robust Speech Representations with Density Adaptive Attention
por: Ioannides, Georgios, et al.
Publicado: (2025)
por: Ioannides, Georgios, et al.
Publicado: (2025)
GenTranslate: Large Language Models are Generative Multilingual Speech and Machine Translators
por: Hu, Yuchen, et al.
Publicado: (2024)
por: Hu, Yuchen, et al.
Publicado: (2024)
Acoustic Identification of Ae. aegypti Mosquitoes using Smartphone Apps and Residual Convolutional Neural Networks
por: Paim, Kayuã Oleques, et al.
Publicado: (2023)
por: Paim, Kayuã Oleques, et al.
Publicado: (2023)
SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes
por: Alex, Tony, et al.
Publicado: (2025)
por: Alex, Tony, et al.
Publicado: (2025)
Ejemplares similares
-
Whisper Turns Stronger: Augmenting Wav2Vec 2.0 for Superior ASR in Low-Resource Languages
por: Anidjar, Or Haim, et al.
Publicado: (2024) -
Vocal Tract Length Warped Features for Spoken Keyword Spotting
por: Sarkar, Achintya kr., et al.
Publicado: (2025) -
Deep Neural Network for Musical Instrument Recognition using MFCCs
por: Mahanta, Saranga Kingkor, et al.
Publicado: (2021) -
Spoken Language Intelligence of Large Language Models for Language Learning
por: Peng, Linkai, et al.
Publicado: (2023) -
Focal Loss based Residual Convolutional Neural Network for Speech Emotion Recognition
por: Tripathi, Suraj, et al.
Publicado: (2019)