Saved in:
| Main Authors: | Padhya, Dinanath, Maharjan, Sajen, Adhikari, Binita, Pokharel, Ishwor Raj |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.14736 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Design and Implementation of a Multi-Purpose Low-Cost Hall-Effect Sensor Glove for Sign Language Recognition
by: Padhya, Dinanath, et al.
Published: (2025)
by: Padhya, Dinanath, et al.
Published: (2025)
Audio-visual video-to-speech synthesis with synthesized input audio
by: Kefalas, Triantafyllos, et al.
Published: (2023)
by: Kefalas, Triantafyllos, et al.
Published: (2023)
Character-aware audio-visual subtitling in context
by: Huh, Jaesung, et al.
Published: (2024)
by: Huh, Jaesung, et al.
Published: (2024)
IsoNet: Causal Analysis of Multimodal Transformers for Neuromuscular Gesture Classification
by: Tyacke, Eion, et al.
Published: (2025)
by: Tyacke, Eion, et al.
Published: (2025)
Late fusion ensembles for speech recognition on diverse input audio representations
by: Jezidžić, Marin, et al.
Published: (2024)
by: Jezidžić, Marin, et al.
Published: (2024)
Enabling automatic transcription of child-centered audio recordings from real-world environments
by: Kocharov, Daniil, et al.
Published: (2025)
by: Kocharov, Daniil, et al.
Published: (2025)
Versatile audio-visual learning for emotion recognition
by: Goncalves, Lucas, et al.
Published: (2023)
by: Goncalves, Lucas, et al.
Published: (2023)
Throat and acoustic paired speech dataset for deep learning-based speech enhancement
by: Kim, Yunsik, et al.
Published: (2025)
by: Kim, Yunsik, et al.
Published: (2025)
Investigating self-supervised representations for audio-visual deepfake detection
by: Boldisor, Dragos-Alexandru, et al.
Published: (2025)
by: Boldisor, Dragos-Alexandru, et al.
Published: (2025)
Large-scale unsupervised audio pre-training for video-to-speech synthesis
by: Kefalas, Triantafyllos, et al.
Published: (2023)
by: Kefalas, Triantafyllos, et al.
Published: (2023)
Robustifying automatic speech recognition by extracting slowly varying features
by: Pizarro, Matías, et al.
Published: (2021)
by: Pizarro, Matías, et al.
Published: (2021)
Beyond saliency: enhancing explanation of speech emotion recognition with expert-referenced acoustic cues
by: Nasr, Seham, et al.
Published: (2025)
by: Nasr, Seham, et al.
Published: (2025)
Training chord recognition models on artificially generated audio
by: Majchrzak, Martyna, et al.
Published: (2025)
by: Majchrzak, Martyna, et al.
Published: (2025)
Adversarial multi-task underwater acoustic target recognition: towards robustness against various influential factors
by: Xie, Yuan, et al.
Published: (2024)
by: Xie, Yuan, et al.
Published: (2024)
Fusion approaches for emotion recognition from speech using acoustic and text-based features
by: Pepino, Leonardo, et al.
Published: (2024)
by: Pepino, Leonardo, et al.
Published: (2024)
Context-aware child-directed speech detection from long-form recordings
by: Charlot, Théo, et al.
Published: (2026)
by: Charlot, Théo, et al.
Published: (2026)
CNN-LSTM Hybrid Architecture for Over-the-Air Automatic Modulation Classification Using SDR
by: Padhya, Dinanath, et al.
Published: (2025)
by: Padhya, Dinanath, et al.
Published: (2025)
Lightweight and perceptually-guided voice conversion for electro-laryngeal speech
by: Mayrhofer, Benedikt, et al.
Published: (2026)
by: Mayrhofer, Benedikt, et al.
Published: (2026)
Modeling speech emotion with label variance and analyzing performance across speakers and unseen acoustic conditions
by: Mitra, Vikramjit, et al.
Published: (2025)
by: Mitra, Vikramjit, et al.
Published: (2025)
Better audio representations are more brain-like: linking model-brain alignment with performance in downstream auditory tasks
by: Pepino, Leonardo, et al.
Published: (2025)
by: Pepino, Leonardo, et al.
Published: (2025)
The silence of the weights: a structural pruning strategy for attention-based audio signal architectures with second order metrics
by: Diecidue, Andrea, et al.
Published: (2025)
by: Diecidue, Andrea, et al.
Published: (2025)
Transformation of audio embeddings into interpretable, concept-based representations
by: Zhang, Alice, et al.
Published: (2025)
by: Zhang, Alice, et al.
Published: (2025)
Towards generalizing deep-audio fake detection networks
by: Gasenzer, Konstantin, et al.
Published: (2023)
by: Gasenzer, Konstantin, et al.
Published: (2023)
Circumventing shortcuts in audio-visual deepfake detection datasets with unsupervised learning
by: Smeu, Stefan, et al.
Published: (2024)
by: Smeu, Stefan, et al.
Published: (2024)
RelUNet: Relative Channel Fusion U-Net for Multichannel Speech Enhancement
by: Aldarmaki, Ibrahim, et al.
Published: (2024)
by: Aldarmaki, Ibrahim, et al.
Published: (2024)
Testing chatbots on the creation of encoders for audio conditioned image generation
by: León, Jorge E., et al.
Published: (2025)
by: León, Jorge E., et al.
Published: (2025)
Unsupervised outlier detection to improve bird audio dataset labels
by: Collins, Bruce
Published: (2025)
by: Collins, Bruce
Published: (2025)
Combining audio control and style transfer using latent diffusion
by: Demerlé, Nils, et al.
Published: (2024)
by: Demerlé, Nils, et al.
Published: (2024)
U-Mamba-Net: A highly efficient Mamba-based U-net style network for noisy and reverberant speech separation
by: Dang, Shaoxiang, et al.
Published: (2024)
by: Dang, Shaoxiang, et al.
Published: (2024)
MPIPN: A Multi Physics-Informed PointNet for solving parametric acoustic-structure systems
by: Wang, Chu, et al.
Published: (2024)
by: Wang, Chu, et al.
Published: (2024)
Deepfake audio as a data augmentation technique for training automatic speech to text transcription models
by: Ferreira, Alexandre R., et al.
Published: (2023)
by: Ferreira, Alexandre R., et al.
Published: (2023)
Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks
by: Maiti, Soumi, et al.
Published: (2023)
by: Maiti, Soumi, et al.
Published: (2023)
LRS-VoxMM: A benchmark for in-the-wild audio-visual speech recognition
by: Kwak, Doyeop, et al.
Published: (2026)
by: Kwak, Doyeop, et al.
Published: (2026)
Selfsupervised learning for pathological speech detection
by: Sheikh, Shakeel Ahmad
Published: (2024)
by: Sheikh, Shakeel Ahmad
Published: (2024)
Towards the Synthesis of Non-speech Vocalizations
by: Hoq, Enjamamul, et al.
Published: (2024)
by: Hoq, Enjamamul, et al.
Published: (2024)
EnCodecMAE: Leveraging neural codecs for universal audio representation learning
by: Pepino, Leonardo, et al.
Published: (2023)
by: Pepino, Leonardo, et al.
Published: (2023)
Virtual boundary integral neural network for three-dimensional exterior acoustic problems
by: Li, Jiahao, et al.
Published: (2026)
by: Li, Jiahao, et al.
Published: (2026)
GRAM: Spatial general-purpose audio representations for real-world environments
by: Yuksel, Goksenin, et al.
Published: (2026)
by: Yuksel, Goksenin, et al.
Published: (2026)
Recomposer: Event-roll-guided generative audio editing
by: Ellis, Daniel P. W., et al.
Published: (2025)
by: Ellis, Daniel P. W., et al.
Published: (2025)
Introduction to speech recognition
by: Dauphin, Gabriel
Published: (2024)
by: Dauphin, Gabriel
Published: (2024)
Similar Items
-
Design and Implementation of a Multi-Purpose Low-Cost Hall-Effect Sensor Glove for Sign Language Recognition
by: Padhya, Dinanath, et al.
Published: (2025) -
Audio-visual video-to-speech synthesis with synthesized input audio
by: Kefalas, Triantafyllos, et al.
Published: (2023) -
Character-aware audio-visual subtitling in context
by: Huh, Jaesung, et al.
Published: (2024) -
IsoNet: Causal Analysis of Multimodal Transformers for Neuromuscular Gesture Classification
by: Tyacke, Eion, et al.
Published: (2025) -
Late fusion ensembles for speech recognition on diverse input audio representations
by: Jezidžić, Marin, et al.
Published: (2024)