voice2mode: Phonation Mode Classification in Singing using Self-Supervised Speech Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Justus, Aju Ani, Agrawal, Ruchit, Kadiri, Sudarsana Reddy, Narayanan, Shrikanth |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Layer-Wise Analysis of Self-Supervised Representations for Age and Gender Classification in Children's Speech
di: Sinha, Abhijit, et al.
Pubblicazione: (2025)
di: Sinha, Abhijit, et al.
Pubblicazione: (2025)
Can a Machine Distinguish High and Low Amount of Social Creak in Speech?
di: Laukkanen, Anne-Maria, et al.
Pubblicazione: (2024)
di: Laukkanen, Anne-Maria, et al.
Pubblicazione: (2024)
Can Layer-wise SSL Features Improve Zero-Shot ASR Performance for Children's Speech?
di: Sinha, Abhijit, et al.
Pubblicazione: (2025)
di: Sinha, Abhijit, et al.
Pubblicazione: (2025)
Evaluation of Speech Foundation Models for ASR on Child-Adult Conversations in Autism Diagnostic Sessions
di: Ashvin, Aditya, et al.
Pubblicazione: (2024)
di: Ashvin, Aditya, et al.
Pubblicazione: (2024)
MMSD-Net: Towards Multi-modal Stuttering Detection
di: Nie, Liangyu, et al.
Pubblicazione: (2024)
di: Nie, Liangyu, et al.
Pubblicazione: (2024)
Music Generation using Human-In-The-Loop Reinforcement Learning
di: Justus, Aju Ani
Pubblicazione: (2025)
di: Justus, Aju Ani
Pubblicazione: (2025)
Speech Emotion Recognition with Phonation Excitation Information and Articulatory Kinematics
di: Zhang, Ziqian, et al.
Pubblicazione: (2025)
di: Zhang, Ziqian, et al.
Pubblicazione: (2025)
Can Synthetic Audio From Generative Foundation Models Assist Audio Recognition and Speech Modeling?
di: Feng, Tiantian, et al.
Pubblicazione: (2024)
di: Feng, Tiantian, et al.
Pubblicazione: (2024)
Toward Fully-End-to-End Listened Speech Decoding from EEG Signals
di: Lee, Jihwan, et al.
Pubblicazione: (2024)
di: Lee, Jihwan, et al.
Pubblicazione: (2024)
Low-Resource Cross-Domain Singing Voice Synthesis via Reduced Self-Supervised Speech Representations
di: Kakoulidis, Panos, et al.
Pubblicazione: (2024)
di: Kakoulidis, Panos, et al.
Pubblicazione: (2024)
Affect Decoding in Phonated and Silent Speech Production from Surface EMG
di: Pistrosch, Simon, et al.
Pubblicazione: (2026)
di: Pistrosch, Simon, et al.
Pubblicazione: (2026)
Towards disentangling the contributions of articulation and acoustics in multimodal phoneme recognition
di: Foley, Sean, et al.
Pubblicazione: (2025)
di: Foley, Sean, et al.
Pubblicazione: (2025)
A Unified Model For Voice and Accent Conversion In Speech and Singing using Self-Supervised Learning and Feature Extraction
di: Cheripally, Sowmya
Pubblicazione: (2024)
di: Cheripally, Sowmya
Pubblicazione: (2024)
Singing Voice Conversion with Accompaniment Using Self-Supervised Representation-Based Melody Features
di: Chen, Wei, et al.
Pubblicazione: (2025)
di: Chen, Wei, et al.
Pubblicazione: (2025)
PEFT-SER: On the Use of Parameter Efficient Transfer Learning Approaches For Speech Emotion Recognition Using Pre-trained Speech Models
di: Feng, Tiantian, et al.
Pubblicazione: (2023)
di: Feng, Tiantian, et al.
Pubblicazione: (2023)
Who Said What WSW 2.0? Enhanced Automated Analysis of Preschool Classroom Speech
di: Sun, Anchen, et al.
Pubblicazione: (2025)
di: Sun, Anchen, et al.
Pubblicazione: (2025)
LLMs as Policy-Agnostic Teammates: A Case Study in Human Proxy Design for Heterogeneous Agent Teams
di: Justus, Aju Ani, et al.
Pubblicazione: (2025)
di: Justus, Aju Ani, et al.
Pubblicazione: (2025)
A Semi-Supervised Framework for Speech Confidence Detection using Whisper
di: Wynn, Adam, et al.
Pubblicazione: (2026)
di: Wynn, Adam, et al.
Pubblicazione: (2026)
Zero-Shot KWS for Children's Speech using Layer-Wise Features from SSL Models
di: Kutum, Subham, et al.
Pubblicazione: (2025)
di: Kutum, Subham, et al.
Pubblicazione: (2025)
Examining Test-Time Adaptation for Personalized Child Speech Recognition
di: Shi, Zhonghao, et al.
Pubblicazione: (2024)
di: Shi, Zhonghao, et al.
Pubblicazione: (2024)
A Dataset for Automatic Vocal Mode Classification
di: Hinrichs, Reemt, et al.
Pubblicazione: (2026)
di: Hinrichs, Reemt, et al.
Pubblicazione: (2026)
Generative Multi-modal Feedback for Singing Voice Synthesis Evaluation
di: Li, Xueyan, et al.
Pubblicazione: (2025)
di: Li, Xueyan, et al.
Pubblicazione: (2025)
Towards Early Prediction of Self-Supervised Speech Model Performance
di: Whetten, Ryan, et al.
Pubblicazione: (2025)
di: Whetten, Ryan, et al.
Pubblicazione: (2025)
The Effect of Batch Size on Contrastive Self-Supervised Speech Representation Learning
di: Vaessen, Nik, et al.
Pubblicazione: (2024)
di: Vaessen, Nik, et al.
Pubblicazione: (2024)
Self-Supervised Singing Voice Pre-Training towards Speech-to-Singing Conversion
di: Li, Ruiqi, et al.
Pubblicazione: (2024)
di: Li, Ruiqi, et al.
Pubblicazione: (2024)
Lightweight and perceptually-guided voice conversion for electro-laryngeal speech
di: Mayrhofer, Benedikt, et al.
Pubblicazione: (2026)
di: Mayrhofer, Benedikt, et al.
Pubblicazione: (2026)
Self-Supervised Learning for Few-Shot Bird Sound Classification
di: Moummad, Ilyass, et al.
Pubblicazione: (2023)
di: Moummad, Ilyass, et al.
Pubblicazione: (2023)
Property Neurons in Self-Supervised Speech Transformers
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2024)
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2024)
Impact of Speech Mode in Automatic Pathological Speech Detection
di: Sheikh, Shakeel A., et al.
Pubblicazione: (2024)
di: Sheikh, Shakeel A., et al.
Pubblicazione: (2024)
A Novel Fusion Architecture for PD Detection Using Semi-Supervised Speech Embeddings
di: Adnan, Tariq, et al.
Pubblicazione: (2024)
di: Adnan, Tariq, et al.
Pubblicazione: (2024)
Enhancing Listened Speech Decoding from EEG via Parallel Phoneme Sequence Prediction
di: Lee, Jihwan, et al.
Pubblicazione: (2025)
di: Lee, Jihwan, et al.
Pubblicazione: (2025)
Losses Can Be Blessings: Routing Self-Supervised Speech Representations Towards Efficient Multilingual and Multitask Speech Processing
di: Fu, Yonggan, et al.
Pubblicazione: (2022)
di: Fu, Yonggan, et al.
Pubblicazione: (2022)
Understanding Self-Supervised Learning of Speech Representation via Invariance and Redundancy Reduction
di: Brima, Yusuf, et al.
Pubblicazione: (2023)
di: Brima, Yusuf, et al.
Pubblicazione: (2023)
On the Transferability of Large-Scale Self-Supervision to Few-Shot Audio Classification
di: Heggan, Calum, et al.
Pubblicazione: (2024)
di: Heggan, Calum, et al.
Pubblicazione: (2024)
Windowed SummaryMixing: An Efficient Fine-Tuning of Self-Supervised Learning Models for Low-resource Speech Recognition
di: Menon, Aditya Srinivas, et al.
Pubblicazione: (2026)
di: Menon, Aditya Srinivas, et al.
Pubblicazione: (2026)
Self-Supervised Speech Quality Estimation and Enhancement Using Only Clean Speech
di: Fu, Szu-Wei, et al.
Pubblicazione: (2024)
di: Fu, Szu-Wei, et al.
Pubblicazione: (2024)
Efficient Training of Self-Supervised Speech Foundation Models on a Compute Budget
di: Liu, Andy T., et al.
Pubblicazione: (2024)
di: Liu, Andy T., et al.
Pubblicazione: (2024)
DAISY: Data Adaptive Self-Supervised Early Exit for Speech Representation Models
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2024)
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2024)
Can DeepFake Speech be Reliably Detected?
di: Liu, Hongbin, et al.
Pubblicazione: (2024)
di: Liu, Hongbin, et al.
Pubblicazione: (2024)
Huntington Disease Automatic Speech Recognition with Biomarker Supervision
di: Wang, Charles L., et al.
Pubblicazione: (2026)
di: Wang, Charles L., et al.
Pubblicazione: (2026)
Documenti analoghi
-
Layer-Wise Analysis of Self-Supervised Representations for Age and Gender Classification in Children's Speech
di: Sinha, Abhijit, et al.
Pubblicazione: (2025) -
Can a Machine Distinguish High and Low Amount of Social Creak in Speech?
di: Laukkanen, Anne-Maria, et al.
Pubblicazione: (2024) -
Can Layer-wise SSL Features Improve Zero-Shot ASR Performance for Children's Speech?
di: Sinha, Abhijit, et al.
Pubblicazione: (2025) -
Evaluation of Speech Foundation Models for ASR on Child-Adult Conversations in Autism Diagnostic Sessions
di: Ashvin, Aditya, et al.
Pubblicazione: (2024) -
MMSD-Net: Towards Multi-modal Stuttering Detection
di: Nie, Liangyu, et al.
Pubblicazione: (2024)