Weakly Supervised Tabla Stroke Transcription via TI-SDRM: A Rhythm-Aware Lattice Rescoring Framework
Fuente:
arXiv
Salvato in:
| Autori principali: | Kodag, Rahul Bapusaheb, Arora, Vipul |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Meta-learning-based percussion transcription and $t\bar{a}la$ identification from low-resource audio
di: Kodag, Rahul Bapusaheb, et al.
Pubblicazione: (2025)
di: Kodag, Rahul Bapusaheb, et al.
Pubblicazione: (2025)
$T\bar{a}laGen:$ A System for Automatic $T\bar{a}la$ Identification and Generation
di: Kodag, Rahul Bapusaheb, et al.
Pubblicazione: (2024)
di: Kodag, Rahul Bapusaheb, et al.
Pubblicazione: (2024)
AudioNet: Supervised Deep Hashing for Retrieval of Similar Audio Events
di: Dutta, Sagar, et al.
Pubblicazione: (2025)
di: Dutta, Sagar, et al.
Pubblicazione: (2025)
Uncertainty Quantification in Melody Estimation using Histogram Representation
di: Saxena, Kavya Ranjan, et al.
Pubblicazione: (2025)
di: Saxena, Kavya Ranjan, et al.
Pubblicazione: (2025)
SyncNet: correlating objective for time delay estimation in audio signals
di: Raina, Akshay, et al.
Pubblicazione: (2022)
di: Raina, Akshay, et al.
Pubblicazione: (2022)
BEST-STD2.0: Balanced and Efficient Speech Tokenizer for Spoken Term Detection
di: Singh, Anup, et al.
Pubblicazione: (2025)
di: Singh, Anup, et al.
Pubblicazione: (2025)
Improving Active Learning for Melody Estimation by Disentangling Uncertainties
di: Jaiswal, Aayush, et al.
Pubblicazione: (2025)
di: Jaiswal, Aayush, et al.
Pubblicazione: (2025)
Interactive singing melody extraction based on active adaptation
di: Saxena, Kavya Ranjan, et al.
Pubblicazione: (2024)
di: Saxena, Kavya Ranjan, et al.
Pubblicazione: (2024)
Explainable Deep Learning Analysis for Raga Identification in Indian Art Music
di: Singh, Parampreet, et al.
Pubblicazione: (2024)
di: Singh, Parampreet, et al.
Pubblicazione: (2024)
H-QuEST: Accelerating Query-by-Example Spoken Term Detection with Hierarchical Indexing
di: Singh, Akanksha, et al.
Pubblicazione: (2025)
di: Singh, Akanksha, et al.
Pubblicazione: (2025)
Identification and Clustering of Unseen Ragas in Indian Art Music
di: Singh, Parampreet, et al.
Pubblicazione: (2024)
di: Singh, Parampreet, et al.
Pubblicazione: (2024)
Attention-Based Audio Embeddings for Query-by-Example
di: Singh, Anup, et al.
Pubblicazione: (2022)
di: Singh, Anup, et al.
Pubblicazione: (2022)
Learning to Discover: A Generalized Framework for Raga Identification without Forgetting
di: Singh, Parampreet, et al.
Pubblicazione: (2026)
di: Singh, Parampreet, et al.
Pubblicazione: (2026)
Rhythm Features for Speaker Identification
di: Mehlman, Nick, et al.
Pubblicazione: (2025)
di: Mehlman, Nick, et al.
Pubblicazione: (2025)
Initial Decoding with Minimally Augmented Language Model for Improved Lattice Rescoring in Low Resource ASR
di: Murthy, Savitha, et al.
Pubblicazione: (2024)
di: Murthy, Savitha, et al.
Pubblicazione: (2024)
Improved Contextual Recognition In Automatic Speech Recognition Systems By Semantic Lattice Rescoring
di: Sudarshan, Ankitha, et al.
Pubblicazione: (2023)
di: Sudarshan, Ankitha, et al.
Pubblicazione: (2023)
TeLeS: Temporal Lexeme Similarity Score to Estimate Confidence in End-to-End ASR
di: Ravi, Nagarathna, et al.
Pubblicazione: (2024)
di: Ravi, Nagarathna, et al.
Pubblicazione: (2024)
Automatic Detection and Analysis of Singing Mistakes for Music Pedagogy
di: Kumar, Sumit, et al.
Pubblicazione: (2026)
di: Kumar, Sumit, et al.
Pubblicazione: (2026)
Weakly Supervised Phonological Features for Pathological Speech Analysis
di: Thienpondt, Jenthe, et al.
Pubblicazione: (2025)
di: Thienpondt, Jenthe, et al.
Pubblicazione: (2025)
ProGRes: Prompted Generative Rescoring on ASR n-Best
di: Tur, Ada Defne, et al.
Pubblicazione: (2024)
di: Tur, Ada Defne, et al.
Pubblicazione: (2024)
Speech Recognition Rescoring with Large Speech-Text Foundation Models
di: Shivakumar, Prashanth Gurunath, et al.
Pubblicazione: (2024)
di: Shivakumar, Prashanth Gurunath, et al.
Pubblicazione: (2024)
DiffRhythm 2: Efficient and High Fidelity Song Generation via Block Flow Matching
di: Jiang, Yuepeng, et al.
Pubblicazione: (2025)
di: Jiang, Yuepeng, et al.
Pubblicazione: (2025)
Learning from Limited Labels: Transductive Graph Label Propagation for Indian Music Analysis
di: Singh, Parampreet, et al.
Pubblicazione: (2026)
di: Singh, Parampreet, et al.
Pubblicazione: (2026)
SoulX-Transcriber: A Robust End-to-End Framework for Multi-Speaker Speech Transcription
di: Dai, Yuhang, et al.
Pubblicazione: (2026)
di: Dai, Yuhang, et al.
Pubblicazione: (2026)
BEST-STD: Bidirectional Mamba-Enhanced Speech Tokenization for Spoken Term Detection
di: Singh, Anup, et al.
Pubblicazione: (2024)
di: Singh, Anup, et al.
Pubblicazione: (2024)
Towards Weakly Supervised Text-to-Audio Grounding
di: Xu, Xuenan, et al.
Pubblicazione: (2024)
di: Xu, Xuenan, et al.
Pubblicazione: (2024)
Exploiting Noise Inseparability for Weakly-Supervised Discriminative Speech Denoising Using Noisy Targets
di: Maciejewski, Matthew, et al.
Pubblicazione: (2026)
di: Maciejewski, Matthew, et al.
Pubblicazione: (2026)
DeePAQ: A Perceptual Audio Quality Metric Based On Foundational Models and Weakly Supervised Learning
di: Jiang, Guanxin, et al.
Pubblicazione: (2025)
di: Jiang, Guanxin, et al.
Pubblicazione: (2025)
Generating Rhythm Game Music with Jukebox
di: Yan, Nicholas
Pubblicazione: (2023)
di: Yan, Nicholas
Pubblicazione: (2023)
Pseudo Strong Labels from Frame-Level Predictions for Weakly Supervised Sound Event Detection
di: Zhang, Yuliang, et al.
Pubblicazione: (2025)
di: Zhang, Yuliang, et al.
Pubblicazione: (2025)
Mixture to Beamformed Mixture: Leveraging Beamformed Mixture as Weak-Supervision for Speech Enhancement and Noise-Robust ASR
di: Wang, Zhong-Qiu, et al.
Pubblicazione: (2025)
di: Wang, Zhong-Qiu, et al.
Pubblicazione: (2025)
Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching
di: Zuo, Jialong, et al.
Pubblicazione: (2025)
di: Zuo, Jialong, et al.
Pubblicazione: (2025)
Applying LLMs for Rescoring N-best ASR Hypotheses of Casual Conversations: Effects of Domain Adaptation and Context Carry-over
di: Ogawa, Atsunori, et al.
Pubblicazione: (2024)
di: Ogawa, Atsunori, et al.
Pubblicazione: (2024)
Learning When to Trust Which Teacher for Weakly Supervised ASR
di: Agrawal, Aakriti, et al.
Pubblicazione: (2023)
di: Agrawal, Aakriti, et al.
Pubblicazione: (2023)
STARS: A Unified Framework for Singing Transcription, Alignment, and Refined Style Annotation
di: Guo, Wenxiang, et al.
Pubblicazione: (2025)
di: Guo, Wenxiang, et al.
Pubblicazione: (2025)
Recognizing Ornaments in Vocal Indian Art Music with Active Annotation
di: Kumar, Sumit, et al.
Pubblicazione: (2025)
di: Kumar, Sumit, et al.
Pubblicazione: (2025)
Spatially Aware Self-Supervised Models for Multi-Channel Neural Speaker Diarization
di: Han, Jiangyu, et al.
Pubblicazione: (2025)
di: Han, Jiangyu, et al.
Pubblicazione: (2025)
Leveraging AM and FM Rhythm Spectrograms for Dementia Classification and Assessment
di: Gogoi, Parismita, et al.
Pubblicazione: (2025)
di: Gogoi, Parismita, et al.
Pubblicazione: (2025)
DiffRhythm: Blazingly Fast and Embarrassingly Simple End-to-End Full-Length Song Generation with Latent Diffusion
di: Ning, Ziqian, et al.
Pubblicazione: (2025)
di: Ning, Ziqian, et al.
Pubblicazione: (2025)
Speaker Embeddings With Weakly Supervised Voice Activity Detection For Efficient Speaker Diarization
di: Thienpondt, Jenthe, et al.
Pubblicazione: (2024)
di: Thienpondt, Jenthe, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Meta-learning-based percussion transcription and $t\bar{a}la$ identification from low-resource audio
di: Kodag, Rahul Bapusaheb, et al.
Pubblicazione: (2025) -
$T\bar{a}laGen:$ A System for Automatic $T\bar{a}la$ Identification and Generation
di: Kodag, Rahul Bapusaheb, et al.
Pubblicazione: (2024) -
AudioNet: Supervised Deep Hashing for Retrieval of Similar Audio Events
di: Dutta, Sagar, et al.
Pubblicazione: (2025) -
Uncertainty Quantification in Melody Estimation using Histogram Representation
di: Saxena, Kavya Ranjan, et al.
Pubblicazione: (2025) -
SyncNet: correlating objective for time delay estimation in audio signals
di: Raina, Akshay, et al.
Pubblicazione: (2022)