Deep Learning Models in Speech Recognition: Measuring GPU Energy Consumption, Impact of Noise and Model Quantization for Edge Deployment
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Chakravarty, Aditya |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Adapting Whisper for Lightweight and Efficient Automatic Speech Recognition of Children for On-device Edge Applications
von: Dutta, Satwik, et al.
Veröffentlicht: (2025)
von: Dutta, Satwik, et al.
Veröffentlicht: (2025)
A Joint Cross-Attention Model for Audio-Visual Fusion in Dimensional Emotion Recognition
von: Praveen, R. Gnana, et al.
Veröffentlicht: (2022)
von: Praveen, R. Gnana, et al.
Veröffentlicht: (2022)
Directional Source Separation for Robust Speech Recognition on Smart Glasses
von: Feng, Tiantian, et al.
Veröffentlicht: (2023)
von: Feng, Tiantian, et al.
Veröffentlicht: (2023)
Recreating Neural Activity During Speech Production with Language and Speech Model Embeddings
von: Khanday, Owais Mujtaba, et al.
Veröffentlicht: (2025)
von: Khanday, Owais Mujtaba, et al.
Veröffentlicht: (2025)
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition
von: Chen, Youjun, et al.
Veröffentlicht: (2025)
von: Chen, Youjun, et al.
Veröffentlicht: (2025)
Using Confidence Scores to Improve Eyes-free Detection of Speech Recognition Errors
von: Nowrin, Sadia, et al.
Veröffentlicht: (2024)
von: Nowrin, Sadia, et al.
Veröffentlicht: (2024)
AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals
von: Zhou, Dongliang, et al.
Veröffentlicht: (2025)
von: Zhou, Dongliang, et al.
Veröffentlicht: (2025)
Real-Time Word-Level Temporal Segmentation in Streaming Speech Recognition
von: Nishida, Naoto, et al.
Veröffentlicht: (2025)
von: Nishida, Naoto, et al.
Veröffentlicht: (2025)
HiCMAE: Hierarchical Contrastive Masked Autoencoder for Self-Supervised Audio-Visual Emotion Recognition
von: Sun, Licai, et al.
Veröffentlicht: (2024)
von: Sun, Licai, et al.
Veröffentlicht: (2024)
Personalized Speech Emotion Recognition in Human-Robot Interaction using Vision Transformers
von: Mishra, Ruchik, et al.
Veröffentlicht: (2024)
von: Mishra, Ruchik, et al.
Veröffentlicht: (2024)
A Near-Real-Time Processing Ego Speech Filtering Pipeline Designed for Speech Interruption During Human-Robot Interaction
von: Li, Yue, et al.
Veröffentlicht: (2024)
von: Li, Yue, et al.
Veröffentlicht: (2024)
Towards Temporally Explainable Dysarthric Speech Clarity Assessment
von: Park, Seohyun, et al.
Veröffentlicht: (2025)
von: Park, Seohyun, et al.
Veröffentlicht: (2025)
SACM: SEEG-Audio Contrastive Matching for Chinese Speech Decoding
von: Wang, Hongbin, et al.
Veröffentlicht: (2025)
von: Wang, Hongbin, et al.
Veröffentlicht: (2025)
VoiceX: A Text-To-Speech Framework for Custom Voices
von: Mertes, Silvan, et al.
Veröffentlicht: (2024)
von: Mertes, Silvan, et al.
Veröffentlicht: (2024)
Psychophysiology-aided Perceptually Fluent Speech Analysis of Children Who Stutter
von: Xiao, Yi, et al.
Veröffentlicht: (2022)
von: Xiao, Yi, et al.
Veröffentlicht: (2022)
NeuroIncept Decoder for High-Fidelity Speech Reconstruction from Neural Activity
von: Khanday, Owais Mujtaba, et al.
Veröffentlicht: (2025)
von: Khanday, Owais Mujtaba, et al.
Veröffentlicht: (2025)
Improving Multimodal Emotion Recognition by Leveraging Acoustic Adaptation and Visual Alignment
von: Zhao, Zhixian, et al.
Veröffentlicht: (2024)
von: Zhao, Zhixian, et al.
Veröffentlicht: (2024)
USpeech: Ultrasound-Enhanced Speech with Minimal Human Effort via Cross-Modal Synthesis
von: Yu, Luca Jiang-Tao, et al.
Veröffentlicht: (2024)
von: Yu, Luca Jiang-Tao, et al.
Veröffentlicht: (2024)
How Private is Low-Frequency Speech Audio in the Wild? An Analysis of Verbal Intelligibility by Humans and Machines
von: Liu, Ailin, et al.
Veröffentlicht: (2024)
von: Liu, Ailin, et al.
Veröffentlicht: (2024)
Human Feedback Driven Dynamic Speech Emotion Recognition
von: Fedorov, Ilya, et al.
Veröffentlicht: (2025)
von: Fedorov, Ilya, et al.
Veröffentlicht: (2025)
Investigating the Effects of Large-Scale Pseudo-Stereo Data and Different Speech Foundation Model on Dialogue Generative Spoken Language Model
von: Fu, Yu-Kuan, et al.
Veröffentlicht: (2024)
von: Fu, Yu-Kuan, et al.
Veröffentlicht: (2024)
Interactive Sonification for Health and Energy using ChucK and Unity
von: Zhao, Yichun, et al.
Veröffentlicht: (2024)
von: Zhao, Yichun, et al.
Veröffentlicht: (2024)
Optimizing Multilingual Text-To-Speech with Accents & Emotions
von: Pawar, Pranav, et al.
Veröffentlicht: (2025)
von: Pawar, Pranav, et al.
Veröffentlicht: (2025)
SingVisio: Visual Analytics of Diffusion Model for Singing Voice Conversion
von: Xue, Liumeng, et al.
Veröffentlicht: (2024)
von: Xue, Liumeng, et al.
Veröffentlicht: (2024)
A Cross-Modal Approach to Silent Speech with LLM-Enhanced Recognition
von: Benster, Tyler, et al.
Veröffentlicht: (2024)
von: Benster, Tyler, et al.
Veröffentlicht: (2024)
Cross-Lingual Speech Emotion Recognition: Humans vs. Self-Supervised Models
von: Han, Zhichen, et al.
Veröffentlicht: (2024)
von: Han, Zhichen, et al.
Veröffentlicht: (2024)
Speech vs. Transcript: Does It Matter for Human Annotators in Speech Summarization?
von: Sharma, Roshan, et al.
Veröffentlicht: (2024)
von: Sharma, Roshan, et al.
Veröffentlicht: (2024)
STAA-Net: A Sparse and Transferable Adversarial Attack for Speech Emotion Recognition
von: Chang, Yi, et al.
Veröffentlicht: (2024)
von: Chang, Yi, et al.
Veröffentlicht: (2024)
AIx Speed: Playback Speed Optimization Using Listening Comprehension of Speech Recognition Models
von: Kawamura, Kazuki, et al.
Veröffentlicht: (2024)
von: Kawamura, Kazuki, et al.
Veröffentlicht: (2024)
Cervical Auscultation Machine Learning for Dysphagia Assessment
von: Chia, An An, et al.
Veröffentlicht: (2024)
von: Chia, An An, et al.
Veröffentlicht: (2024)
Sound2Hap: Learning Audio-to-Vibrotactile Haptic Generation from Human Ratings
von: Li, Yinan, et al.
Veröffentlicht: (2026)
von: Li, Yinan, et al.
Veröffentlicht: (2026)
Robust Dual-Modal Speech Keyword Spotting for XR Headsets
von: Cai, Zhuojiang, et al.
Veröffentlicht: (2024)
von: Cai, Zhuojiang, et al.
Veröffentlicht: (2024)
Integrating Representational Gestures into Automatically Generated Embodied Explanations and its Effects on Understanding and Interaction Quality
von: Robrecht, Amelie Sophie, et al.
Veröffentlicht: (2024)
von: Robrecht, Amelie Sophie, et al.
Veröffentlicht: (2024)
Soundify: Matching Sound Effects to Video
von: Lin, David Chuan-En, et al.
Veröffentlicht: (2021)
von: Lin, David Chuan-En, et al.
Veröffentlicht: (2021)
Emotion-Disentangled Embedding Alignment for Noise-Robust and Cross-Corpus Speech Emotion Recognition
von: Tiwari, Upasana, et al.
Veröffentlicht: (2025)
von: Tiwari, Upasana, et al.
Veröffentlicht: (2025)
Detecting COPD Through Speech Analysis: A Dataset of Danish Speech and Machine Learning Approach
von: Sankey-Olsen, Cuno, et al.
Veröffentlicht: (2025)
von: Sankey-Olsen, Cuno, et al.
Veröffentlicht: (2025)
Sound-Based Recognition of Touch Gestures and Emotions for Enhanced Human-Robot Interaction
von: Hou, Yuanbo, et al.
Veröffentlicht: (2024)
von: Hou, Yuanbo, et al.
Veröffentlicht: (2024)
InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training
von: Wang, Dingdong, et al.
Veröffentlicht: (2025)
von: Wang, Dingdong, et al.
Veröffentlicht: (2025)
Subject Disentanglement Neural Network for Speech Envelope Reconstruction from EEG
von: Zhang, Li, et al.
Veröffentlicht: (2025)
von: Zhang, Li, et al.
Veröffentlicht: (2025)
Speech-driven Personalized Gesture Synthetics: Harnessing Automatic Fuzzy Feature Inference
von: Zhang, Fan, et al.
Veröffentlicht: (2024)
von: Zhang, Fan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Adapting Whisper for Lightweight and Efficient Automatic Speech Recognition of Children for On-device Edge Applications
von: Dutta, Satwik, et al.
Veröffentlicht: (2025) -
A Joint Cross-Attention Model for Audio-Visual Fusion in Dimensional Emotion Recognition
von: Praveen, R. Gnana, et al.
Veröffentlicht: (2022) -
Directional Source Separation for Robust Speech Recognition on Smart Glasses
von: Feng, Tiantian, et al.
Veröffentlicht: (2023) -
Recreating Neural Activity During Speech Production with Language and Speech Model Embeddings
von: Khanday, Owais Mujtaba, et al.
Veröffentlicht: (2025) -
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition
von: Chen, Youjun, et al.
Veröffentlicht: (2025)