Dhvani: A Weakly-supervised Phonemic Error Detection and Personalized Feedback System for Hindi
Fuente:
arXiv
Guardado en:
| Autores principales: | Rustagi, Arnav, Bajpai, Satvik, Kaur, Nimrat, Siddharth, Siddharth |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
RepCNN: Micro-sized, Mighty Models for Wakeword Detection
por: Kundu, Arnav, et al.
Publicado: (2024)
por: Kundu, Arnav, et al.
Publicado: (2024)
Aligning Generative Speech Enhancement with Perceptual Feedback
por: Li, Haoyang, et al.
Publicado: (2025)
por: Li, Haoyang, et al.
Publicado: (2025)
Defense Against Synthetic Speech: Real-Time Detection of RVC Voice Conversion Attacks
por: Chinchmalatpure, Prajwal, et al.
Publicado: (2025)
por: Chinchmalatpure, Prajwal, et al.
Publicado: (2025)
Weakly Supervised Detection and Temporal Localization of Whale Calls in Long-Duration Bioacoustic Data
por: Nihal, Ragib Amin, et al.
Publicado: (2025)
por: Nihal, Ragib Amin, et al.
Publicado: (2025)
HyWA: Hypernetwork Weight Adapting Personalized Voice Activity Detection
por: Nejad, Mahsa Ghazvini, et al.
Publicado: (2025)
por: Nejad, Mahsa Ghazvini, et al.
Publicado: (2025)
Speaker- and Text-Independent Estimation of Articulatory Movements and Phoneme Alignments from Speech
por: Weise, Tobias, et al.
Publicado: (2024)
por: Weise, Tobias, et al.
Publicado: (2024)
End to end Hindi to English speech conversion using Bark, mBART and a finetuned XLSR Wav2Vec2
por: Tathe, Aniket, et al.
Publicado: (2024)
por: Tathe, Aniket, et al.
Publicado: (2024)
Exploring bat song syllable representations in self-supervised audio encoders
por: Kloots, Marianne de Heer, et al.
Publicado: (2024)
por: Kloots, Marianne de Heer, et al.
Publicado: (2024)
AURA Score: A Metric For Holistic Audio Question Answering Evaluation
por: Dixit, Satvik, et al.
Publicado: (2025)
por: Dixit, Satvik, et al.
Publicado: (2025)
PESTO: Real-Time Pitch Estimation with Self-supervised Transposition-equivariant Objective
por: Riou, Alain, et al.
Publicado: (2025)
por: Riou, Alain, et al.
Publicado: (2025)
Phoneme-Level Feature Discrepancies: A Key to Detecting Sophisticated Speech Deepfakes
por: Zhang, Kuiyuan, et al.
Publicado: (2024)
por: Zhang, Kuiyuan, et al.
Publicado: (2024)
MEBM-Phoneme: Multi-scale Enhanced BrainMagic for End-to-End MEG Phoneme Classification
por: Jinghua, Liang, et al.
Publicado: (2026)
por: Jinghua, Liang, et al.
Publicado: (2026)
Continuous Autoregressive Models with Noise Augmentation Avoid Error Accumulation
por: Pasini, Marco, et al.
Publicado: (2024)
por: Pasini, Marco, et al.
Publicado: (2024)
PPINtonus: Early Detection of Parkinson's Disease Using Deep-Learning Tonal Analysis
por: Reddy, Varun
Publicado: (2024)
por: Reddy, Varun
Publicado: (2024)
Personalized Speech Enhancement Without a Separate Speaker Embedding Model
por: Pärnamaa, Tanel, et al.
Publicado: (2024)
por: Pärnamaa, Tanel, et al.
Publicado: (2024)
Optimizing Estonian TV Subtitles with Semi-supervised Learning and LLMs
por: Fedorchenko, Artem, et al.
Publicado: (2025)
por: Fedorchenko, Artem, et al.
Publicado: (2025)
Data-Efficient ASR Personalization for Non-Normative Speech Using an Uncertainty-Based Phoneme Difficulty Score for Guided Sampling
por: Pokel, Niclas, et al.
Publicado: (2025)
por: Pokel, Niclas, et al.
Publicado: (2025)
The 2025 PNPL Competition: Speech Detection and Phoneme Classification in the LibriBrain Dataset
por: Landau, Gilad, et al.
Publicado: (2025)
por: Landau, Gilad, et al.
Publicado: (2025)
A SUPERB-Style Benchmark of Self-Supervised Speech Models for Audio Deepfake Detection
por: Ali, Hashim, et al.
Publicado: (2026)
por: Ali, Hashim, et al.
Publicado: (2026)
Phoneme-Level Deepfake Detection Across Emotional Conditions Using Self-Supervised Embeddings
por: Nallaguntla, Vamshi, et al.
Publicado: (2026)
por: Nallaguntla, Vamshi, et al.
Publicado: (2026)
Towards Human-in-the-Loop Onset Detection: A Transfer Learning Approach for Maracatu
por: Pinto, António Sá
Publicado: (2025)
por: Pinto, António Sá
Publicado: (2025)
Hindi audio-video-Deepfake (HAV-DF): A Hindi language-based Audio-video Deepfake Dataset
por: Kaur, Sukhandeep, et al.
Publicado: (2024)
por: Kaur, Sukhandeep, et al.
Publicado: (2024)
A Novel Hybrid Deep Learning Technique for Speech Emotion Detection using Feature Engineering
por: Chowdhury, Shahana Yasmin, et al.
Publicado: (2025)
por: Chowdhury, Shahana Yasmin, et al.
Publicado: (2025)
AG-LSEC: Audio Grounded Lexical Speaker Error Correction
por: Paturi, Rohit, et al.
Publicado: (2024)
por: Paturi, Rohit, et al.
Publicado: (2024)
Weakly-supervised Audio Separation via Bi-modal Semantic Similarity
por: Mahmud, Tanvir, et al.
Publicado: (2024)
por: Mahmud, Tanvir, et al.
Publicado: (2024)
Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning
por: Abdelfattah, Abdullah, et al.
Publicado: (2025)
por: Abdelfattah, Abdullah, et al.
Publicado: (2025)
A Data-Driven Diffusion-based Approach for Audio Deepfake Explanations
por: Grinberg, Petr, et al.
Publicado: (2025)
por: Grinberg, Petr, et al.
Publicado: (2025)
Audio-Based Pedestrian Detection in the Presence of Vehicular Noise
por: Kim, Yonghyun, et al.
Publicado: (2025)
por: Kim, Yonghyun, et al.
Publicado: (2025)
A Closer Look at Wav2Vec2 Embeddings for On-Device Single-Channel Speech Enhancement
por: Shankar, Ravi, et al.
Publicado: (2024)
por: Shankar, Ravi, et al.
Publicado: (2024)
Investigating the Effectiveness of Explainability Methods in Parkinson's Detection from Speech
por: Mancini, Eleonora, et al.
Publicado: (2024)
por: Mancini, Eleonora, et al.
Publicado: (2024)
SwiftF0: Fast and Accurate Monophonic Pitch Detection
por: Nieradzik, Lars
Publicado: (2025)
por: Nieradzik, Lars
Publicado: (2025)
Evaluating Fake Music Detection Performance Under Audio Augmentations
por: Sroka, Tomasz, et al.
Publicado: (2025)
por: Sroka, Tomasz, et al.
Publicado: (2025)
Advancing Marine Bioacoustics with Deep Generative Models: A Hybrid Augmentation Strategy for Southern Resident Killer Whale Detection
por: Padovese, Bruno, et al.
Publicado: (2025)
por: Padovese, Bruno, et al.
Publicado: (2025)
Enhancing Automatic Speech Recognition Through Integrated Noise Detection Architecture
por: Singh, Karamvir
Publicado: (2025)
por: Singh, Karamvir
Publicado: (2025)
WhisperD: Dementia Speech Recognition and Filler Word Detection with Whisper
por: Akinrintoyo, Emmanuel, et al.
Publicado: (2025)
por: Akinrintoyo, Emmanuel, et al.
Publicado: (2025)
Hybrid Disagreement-Diversity Active Learning for Bioacoustic Sound Event Detection
por: Zhang, Shiqi, et al.
Publicado: (2025)
por: Zhang, Shiqi, et al.
Publicado: (2025)
Music Plagiarism Detection: Problem Formulation and a Segment-based Solution
por: Go, Seonghyeon, et al.
Publicado: (2026)
por: Go, Seonghyeon, et al.
Publicado: (2026)
Controllable Singing Voice Synthesis using Phoneme-Level Energy Sequence
por: Ryu, Yerin, et al.
Publicado: (2025)
por: Ryu, Yerin, et al.
Publicado: (2025)
ProKWS: Personalized Keyword Spotting via Collaborative Learning of Phonemes and Prosody
por: Pan, Jianan, et al.
Publicado: (2026)
por: Pan, Jianan, et al.
Publicado: (2026)
Lightweight Hopfield Neural Networks for Bioacoustic Detection and Call Monitoring of Captive Primates
por: Lomas, Wendy, et al.
Publicado: (2025)
por: Lomas, Wendy, et al.
Publicado: (2025)
Ejemplares similares
-
RepCNN: Micro-sized, Mighty Models for Wakeword Detection
por: Kundu, Arnav, et al.
Publicado: (2024) -
Aligning Generative Speech Enhancement with Perceptual Feedback
por: Li, Haoyang, et al.
Publicado: (2025) -
Defense Against Synthetic Speech: Real-Time Detection of RVC Voice Conversion Attacks
por: Chinchmalatpure, Prajwal, et al.
Publicado: (2025) -
Weakly Supervised Detection and Temporal Localization of Whale Calls in Long-Duration Bioacoustic Data
por: Nihal, Ragib Amin, et al.
Publicado: (2025) -
HyWA: Hypernetwork Weight Adapting Personalized Voice Activity Detection
por: Nejad, Mahsa Ghazvini, et al.
Publicado: (2025)