Predicting Heart Activity from Speech using Data-driven and Knowledge-based features
Fuente:
arXiv
Guardado en:
| Autores principales: | Elbanna, Gasser, Mostaani, Zohreh, -Doss, Mathew Magimai. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Evaluating Speaker Identity Coding in Self-supervised Models and Humans
por: Elbanna, Gasser
Publicado: (2024)
por: Elbanna, Gasser
Publicado: (2024)
Assessment of Personality Dimensions Across Situations Using Conversational Speech
por: Zhang, Alice, et al.
Publicado: (2025)
por: Zhang, Alice, et al.
Publicado: (2025)
Incorporating Talker Identity Aids With Improving Speech Recognition in Adversarial Environments
por: Alavilli, Sagarika, et al.
Publicado: (2024)
por: Alavilli, Sagarika, et al.
Publicado: (2024)
Explaining Deep Learning Embeddings for Speech Emotion Recognition by Predicting Interpretable Acoustic Features
por: Dixit, Satvik, et al.
Publicado: (2024)
por: Dixit, Satvik, et al.
Publicado: (2024)
Unsupervised Rhythm and Voice Conversion to Improve ASR on Dysarthric Speech
por: Hajal, Karl El, et al.
Publicado: (2025)
por: Hajal, Karl El, et al.
Publicado: (2025)
Unsupervised Rhythm and Voice Conversion of Dysarthric to Healthy Speech for ASR
por: Hajal, Karl El, et al.
Publicado: (2025)
por: Hajal, Karl El, et al.
Publicado: (2025)
On the Utility of Speech and Audio Foundation Models for Marmoset Call Analysis
por: Sarkar, Eklavya, et al.
Publicado: (2024)
por: Sarkar, Eklavya, et al.
Publicado: (2024)
kNN Retrieval for Simple and Effective Zero-Shot Multi-speaker Text-to-Speech
por: Hajal, Karl El, et al.
Publicado: (2024)
por: Hajal, Karl El, et al.
Publicado: (2024)
Unveiling Audio Deepfake Origins: A Deep Metric learning And Conformer Network Approach With Ensemble Fusion
por: Kulkarni, Ajinkya, et al.
Publicado: (2025)
por: Kulkarni, Ajinkya, et al.
Publicado: (2025)
Joint Semantic Knowledge Distillation and Masked Acoustic Modeling for Full-band Speech Restoration with Improved Intelligibility
por: Liu, Xiaoyu, et al.
Publicado: (2024)
por: Liu, Xiaoyu, et al.
Publicado: (2024)
Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions
por: Gao, Xiaoxue, et al.
Publicado: (2025)
por: Gao, Xiaoxue, et al.
Publicado: (2025)
Automatic Speech Recognition using Advanced Deep Learning Approaches: A survey
por: Kheddar, Hamza, et al.
Publicado: (2024)
por: Kheddar, Hamza, et al.
Publicado: (2024)
Speech Boosting: Low-Latency Live Speech Enhancement for TWS Earbuds
por: Bae, Hanbin, et al.
Publicado: (2024)
por: Bae, Hanbin, et al.
Publicado: (2024)
CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation
por: Kim, Ji-Hoon, et al.
Publicado: (2024)
por: Kim, Ji-Hoon, et al.
Publicado: (2024)
Classification of Heart Sounds Using Multi-Branch Deep Convolutional Network and LSTM-CNN
por: Latifi, Seyed Amir, et al.
Publicado: (2024)
por: Latifi, Seyed Amir, et al.
Publicado: (2024)
Speech Enhancement Based on Drifting Models
por: Xu, Liang, et al.
Publicado: (2026)
por: Xu, Liang, et al.
Publicado: (2026)
Toward Fully-End-to-End Listened Speech Decoding from EEG Signals
por: Lee, Jihwan, et al.
Publicado: (2024)
por: Lee, Jihwan, et al.
Publicado: (2024)
Construction and Evaluation of Mandarin Multimodal Emotional Speech Database
por: Ting, Zhu, et al.
Publicado: (2024)
por: Ting, Zhu, et al.
Publicado: (2024)
A Hybrid Model for Weakly-Supervised Speech Dereverberation
por: Bahrman, Louis, et al.
Publicado: (2025)
por: Bahrman, Louis, et al.
Publicado: (2025)
Tool Wear Prediction in CNC Turning Operations using Ultrasonic Microphone Arrays and CNNs
por: Steckel, Jan, et al.
Publicado: (2024)
por: Steckel, Jan, et al.
Publicado: (2024)
JenGAN: Stacked Shifted Filters in GAN-Based Speech Synthesis
por: Cho, Hyunjae, et al.
Publicado: (2024)
por: Cho, Hyunjae, et al.
Publicado: (2024)
Neural Speech and Audio Coding: Modern AI Technology Meets Traditional Codecs
por: Kim, Minje, et al.
Publicado: (2024)
por: Kim, Minje, et al.
Publicado: (2024)
Discriminating real and synthetic super-resolved audio samples using embedding-based classifiers
por: Silaev, Mikhail, et al.
Publicado: (2026)
por: Silaev, Mikhail, et al.
Publicado: (2026)
EEG-Based Speech Decoding: A Novel Approach Using Multi-Kernel Ensemble Diffusion Models
por: Kim, Soowon, et al.
Publicado: (2024)
por: Kim, Soowon, et al.
Publicado: (2024)
Single-stage TTS with Masked Audio Token Modeling and Semantic Knowledge Distillation
por: Gállego, Gerard I., et al.
Publicado: (2024)
por: Gállego, Gerard I., et al.
Publicado: (2024)
A Domain-Knowledge-Inspired Music Embedding Space and a Novel Attention Mechanism for Symbolic Music Modeling
por: Guo, Z., et al.
Publicado: (2022)
por: Guo, Z., et al.
Publicado: (2022)
Recent Advances in Discrete Speech Tokens: A Review
por: Guo, Yiwei, et al.
Publicado: (2025)
por: Guo, Yiwei, et al.
Publicado: (2025)
IS${}^3$ : Generic Impulsive--Stationary Sound Separation in Acoustic Scenes using Deep Filtering
por: Berger, Clémentine, et al.
Publicado: (2025)
por: Berger, Clémentine, et al.
Publicado: (2025)
Contrastive Knowledge Distillation for Embedding Refinement in Personalized Speech Enhancement
por: Serre, Thomas, et al.
Publicado: (2026)
por: Serre, Thomas, et al.
Publicado: (2026)
EMOCONV-DIFF: Diffusion-based Speech Emotion Conversion for Non-parallel and In-the-wild Data
por: Prabhu, Navin Raj, et al.
Publicado: (2023)
por: Prabhu, Navin Raj, et al.
Publicado: (2023)
Robust Generative Audio Quality Assessment: Disentangling Quality from Spurious Correlations
por: Huang, Kuan-Tang, et al.
Publicado: (2026)
por: Huang, Kuan-Tang, et al.
Publicado: (2026)
Compressing Quaternion Convolutional Neural Networks for Audio Classification
por: Singh, Arshdeep, et al.
Publicado: (2025)
por: Singh, Arshdeep, et al.
Publicado: (2025)
AI-Generated Music Detection in Broadcast Monitoring
por: López-Ayala, David, et al.
Publicado: (2026)
por: López-Ayala, David, et al.
Publicado: (2026)
Differentiable Modal Synthesis for Physical Modeling of Planar String Sound and Motion Simulation
por: Lee, Jin Woo, et al.
Publicado: (2024)
por: Lee, Jin Woo, et al.
Publicado: (2024)
Towards Low-Latency Tracking of Multiple Speakers With Short-Context Speaker Embeddings
por: Iatariene, Taous, et al.
Publicado: (2025)
por: Iatariene, Taous, et al.
Publicado: (2025)
U-DREAM: Unsupervised Dereverberation guided by a Reverberation Model
por: Bahrman, Louis, et al.
Publicado: (2025)
por: Bahrman, Louis, et al.
Publicado: (2025)
Perceptual Noise-Masking with Music through Deep Spectral Envelope Shaping
por: Berger, Clémentine, et al.
Publicado: (2025)
por: Berger, Clémentine, et al.
Publicado: (2025)
SWIM: Short-Window CNN Integrated with Mamba for EEG-Based Auditory Spatial Attention Decoding
por: Zhang, Ziyang, et al.
Publicado: (2024)
por: Zhang, Ziyang, et al.
Publicado: (2024)
Wavetable Synthesis Using CVAE for Timbre Control Based on Semantic Label
por: Yutani, Tsugumasa, et al.
Publicado: (2024)
por: Yutani, Tsugumasa, et al.
Publicado: (2024)
CSL-L2M: Controllable Song-Level Lyric-to-Melody Generation Based on Conditional Transformer with Fine-Grained Lyric and Musical Controls
por: Chai, Li, et al.
Publicado: (2024)
por: Chai, Li, et al.
Publicado: (2024)
Ejemplares similares
-
Evaluating Speaker Identity Coding in Self-supervised Models and Humans
por: Elbanna, Gasser
Publicado: (2024) -
Assessment of Personality Dimensions Across Situations Using Conversational Speech
por: Zhang, Alice, et al.
Publicado: (2025) -
Incorporating Talker Identity Aids With Improving Speech Recognition in Adversarial Environments
por: Alavilli, Sagarika, et al.
Publicado: (2024) -
Explaining Deep Learning Embeddings for Speech Emotion Recognition by Predicting Interpretable Acoustic Features
por: Dixit, Satvik, et al.
Publicado: (2024) -
Unsupervised Rhythm and Voice Conversion to Improve ASR on Dysarthric Speech
por: Hajal, Karl El, et al.
Publicado: (2025)