Foundation Model Hidden Representations for Heart Rate Estimation from Auscultation
Fuente:
arXiv
Guardado en:
| Autores principales: | Nie, Jingping, Tran, Dung T., Thakkar, Karan, Kowtha, Vasudha, Huang, Jon, Avendano, Carlos, Azemi, Erdrin, Mitra, Vikramjit |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Model-driven Heart Rate Estimation and Heart Murmur Detection based on Phonocardiogram
por: Nie, Jingping, et al.
Publicado: (2024)
por: Nie, Jingping, et al.
Publicado: (2024)
Modeling speech emotion with label variance and analyzing performance across speakers and unseen acoustic conditions
por: Mitra, Vikramjit, et al.
Publicado: (2025)
por: Mitra, Vikramjit, et al.
Publicado: (2025)
Voice Quality Dimensions as Interpretable Primitives for Speaking Style for Atypical Speech and Affect
por: Narain, Jaya, et al.
Publicado: (2025)
por: Narain, Jaya, et al.
Publicado: (2025)
Pre-Trained Foundation Model representations to uncover Breathing patterns in Speech
por: Mitra, Vikramjit, et al.
Publicado: (2024)
por: Mitra, Vikramjit, et al.
Publicado: (2024)
A Generalist Audio Foundation Model for Comprehensive Body Sound Auscultation
por: Wang, Pingjie, et al.
Publicado: (2024)
por: Wang, Pingjie, et al.
Publicado: (2024)
Switchboard-Affect: Emotion Perception Labels from Conversational Speech
por: Romana, Amrit, et al.
Publicado: (2025)
por: Romana, Amrit, et al.
Publicado: (2025)
Intelligent Cardiac Auscultation for Murmur Detection via Parallel-Attentive Models with Uncertainty Estimation
por: Zhang, Zixing, et al.
Publicado: (2024)
por: Zhang, Zixing, et al.
Publicado: (2024)
SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion Transformer
por: Wang, Helin, et al.
Publicado: (2024)
por: Wang, Helin, et al.
Publicado: (2024)
LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes
por: Dang, Trung, et al.
Publicado: (2024)
por: Dang, Trung, et al.
Publicado: (2024)
Cervical Auscultation Machine Learning for Dysphagia Assessment
por: Chia, An An, et al.
Publicado: (2024)
por: Chia, An An, et al.
Publicado: (2024)
Patient-Level Multimodal Question Answering from Multi-Site Auscultation Recordings
por: Wu, Fan, et al.
Publicado: (2026)
por: Wu, Fan, et al.
Publicado: (2026)
SegAug: CTC-Aligned Segmented Augmentation For Robust RNN-Transducer Based Speech Recognition
por: Le, Khanh, et al.
Publicado: (2025)
por: Le, Khanh, et al.
Publicado: (2025)
Zero-Shot Text-to-Speech from Continuous Text Streams
por: Dang, Trung, et al.
Publicado: (2024)
por: Dang, Trung, et al.
Publicado: (2024)
ChunkFormer: Masked Chunking Conformer For Long-Form Speech Transcription
por: Le, Khanh, et al.
Publicado: (2025)
por: Le, Khanh, et al.
Publicado: (2025)
Assessing the Alignment of Audio Representations with Timbre Similarity Ratings
por: Tian, Haokun, et al.
Publicado: (2025)
por: Tian, Haokun, et al.
Publicado: (2025)
MAGE: A Coarse-to-Fine Speech Enhancer with Masked Generative Model
por: Pham, The Hieu, et al.
Publicado: (2025)
por: Pham, The Hieu, et al.
Publicado: (2025)
Using Speech Foundational Models in Loss Functions for Hearing Aid Speech Enhancement
por: Sutherland, Robert, et al.
Publicado: (2024)
por: Sutherland, Robert, et al.
Publicado: (2024)
BowelRCNN: Region-based Convolutional Neural Network System for Bowel Sound Auscultation
por: Matynia, Igor, et al.
Publicado: (2025)
por: Matynia, Igor, et al.
Publicado: (2025)
Auditory Representation Effective for Estimating Vocal Tract Information
por: Irino, Toshio, et al.
Publicado: (2023)
por: Irino, Toshio, et al.
Publicado: (2023)
Single-channel speech enhancement using learnable loss mixup
por: Chang, Oscar, et al.
Publicado: (2023)
por: Chang, Oscar, et al.
Publicado: (2023)
Audio-Visual Representation Learning via Knowledge Distillation from Speech Foundation Models
por: Zhang, Jing-Xuan, et al.
Publicado: (2025)
por: Zhang, Jing-Xuan, et al.
Publicado: (2025)
Generative Deep Learning and Signal Processing for Data Augmentation of Cardiac Auscultation Signals: Improving Model Robustness Using Synthetic Audio
por: Abbott, Leigh, et al.
Publicado: (2024)
por: Abbott, Leigh, et al.
Publicado: (2024)
Towards Fusion of Neural Audio Codec-based Representations with Spectral for Heart Murmur Classification via Bandit-based Cross-Attention Mechanism
por: Phukan, Orchid Chetia, et al.
Publicado: (2025)
por: Phukan, Orchid Chetia, et al.
Publicado: (2025)
Fast Word Error Rate Estimation Using Self-Supervised Representations for Speech and Text
por: Park, Chanho, et al.
Publicado: (2023)
por: Park, Chanho, et al.
Publicado: (2023)
Rene: A Pre-trained Multi-modal Architecture for Auscultation of Respiratory Diseases
por: Zhang, Pengfei, et al.
Publicado: (2024)
por: Zhang, Pengfei, et al.
Publicado: (2024)
HSDreport: Heart Sound Diagnosis with Echocardiography Reports
por: Zhao, Zihan, et al.
Publicado: (2024)
por: Zhao, Zihan, et al.
Publicado: (2024)
Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition
por: Wang, Shih-heng, et al.
Publicado: (2024)
por: Wang, Shih-heng, et al.
Publicado: (2024)
Noro: Noise-Robust One-shot Voice Conversion with Hidden Speaker Representation Learning
por: He, Haorui, et al.
Publicado: (2024)
por: He, Haorui, et al.
Publicado: (2024)
Breathing and Semantic Pause Detection and Exertion-Level Classification in Post-Exercise Speech
por: Wang, Yuyu, et al.
Publicado: (2025)
por: Wang, Yuyu, et al.
Publicado: (2025)
BUET Multi-disease Heart Sound Dataset: A Comprehensive Auscultation Dataset for Developing Computer-Aided Diagnostic Systems
por: Ali, Shams Nafisa, et al.
Publicado: (2024)
por: Ali, Shams Nafisa, et al.
Publicado: (2024)
uaMix-MAE: Efficient Tuning of Pretrained Audio Transformers with Unsupervised Audio Mixtures
por: Tabassum, Afrina, et al.
Publicado: (2024)
por: Tabassum, Afrina, et al.
Publicado: (2024)
MambaRate: Speech Quality Assessment Across Different Sampling Rates
por: Kakoulidis, Panos, et al.
Publicado: (2025)
por: Kakoulidis, Panos, et al.
Publicado: (2025)
Point to the Hidden: Exposing Speech Audio Splicing via Signal Pointer Nets
por: Moussa, Denise, et al.
Publicado: (2023)
por: Moussa, Denise, et al.
Publicado: (2023)
Resource-Efficient Adaptation of Speech Foundation Models for Multi-Speaker ASR
por: Wang, Weiqing, et al.
Publicado: (2024)
por: Wang, Weiqing, et al.
Publicado: (2024)
Revealing the Hidden Temporal Structure of HubertSoft Embeddings based on the Russian Phonetic Corpus
por: Ananeva, Anastasia, et al.
Publicado: (2025)
por: Ananeva, Anastasia, et al.
Publicado: (2025)
Generative Speech Foundation Model Pretraining for High-Quality Speech Extraction and Restoration
por: Ku, Pin-Jui, et al.
Publicado: (2024)
por: Ku, Pin-Jui, et al.
Publicado: (2024)
Production and Manufacturing of 3D Printed Acoustic Guitars
por: Tran, Timothy, et al.
Publicado: (2025)
por: Tran, Timothy, et al.
Publicado: (2025)
UniAudio: An Audio Foundation Model Toward Universal Audio Generation
por: Yang, Dongchao, et al.
Publicado: (2023)
por: Yang, Dongchao, et al.
Publicado: (2023)
Exploring Pre-trained General-purpose Audio Representations for Heart Murmur Detection
por: Niizumi, Daisuke, et al.
Publicado: (2024)
por: Niizumi, Daisuke, et al.
Publicado: (2024)
Music Similarity Representation Learning Focusing on Individual Instruments with Source Separation and Human Preference
por: Imamura, Takehiro, et al.
Publicado: (2025)
por: Imamura, Takehiro, et al.
Publicado: (2025)
Ejemplares similares
-
Model-driven Heart Rate Estimation and Heart Murmur Detection based on Phonocardiogram
por: Nie, Jingping, et al.
Publicado: (2024) -
Modeling speech emotion with label variance and analyzing performance across speakers and unseen acoustic conditions
por: Mitra, Vikramjit, et al.
Publicado: (2025) -
Voice Quality Dimensions as Interpretable Primitives for Speaking Style for Atypical Speech and Affect
por: Narain, Jaya, et al.
Publicado: (2025) -
Pre-Trained Foundation Model representations to uncover Breathing patterns in Speech
por: Mitra, Vikramjit, et al.
Publicado: (2024) -
A Generalist Audio Foundation Model for Comprehensive Body Sound Auscultation
por: Wang, Pingjie, et al.
Publicado: (2024)