H-QuEST: Accelerating Query-by-Example Spoken Term Detection with Hierarchical Indexing
Fuente:
arXiv
Salvato in:
| Autori principali: | Singh, Akanksha, Chen, Yi-Ping Phoebe, Arora, Vipul |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
BEST-STD2.0: Balanced and Efficient Speech Tokenizer for Spoken Term Detection
di: Singh, Anup, et al.
Pubblicazione: (2025)
di: Singh, Anup, et al.
Pubblicazione: (2025)
Attention-Based Audio Embeddings for Query-by-Example
di: Singh, Anup, et al.
Pubblicazione: (2022)
di: Singh, Anup, et al.
Pubblicazione: (2022)
BEST-STD: Bidirectional Mamba-Enhanced Speech Tokenization for Spoken Term Detection
di: Singh, Anup, et al.
Pubblicazione: (2024)
di: Singh, Anup, et al.
Pubblicazione: (2024)
Cross-Lingual Query-by-Example Spoken Term Detection: A Transformer-Based Approach
di: Fatemeh, Allahdadi, et al.
Pubblicazione: (2024)
di: Fatemeh, Allahdadi, et al.
Pubblicazione: (2024)
Improving Active Learning for Melody Estimation by Disentangling Uncertainties
di: Jaiswal, Aayush, et al.
Pubblicazione: (2025)
di: Jaiswal, Aayush, et al.
Pubblicazione: (2025)
Explainable Deep Learning Analysis for Raga Identification in Indian Art Music
di: Singh, Parampreet, et al.
Pubblicazione: (2024)
di: Singh, Parampreet, et al.
Pubblicazione: (2024)
AudioNet: Supervised Deep Hashing for Retrieval of Similar Audio Events
di: Dutta, Sagar, et al.
Pubblicazione: (2025)
di: Dutta, Sagar, et al.
Pubblicazione: (2025)
Identification and Clustering of Unseen Ragas in Indian Art Music
di: Singh, Parampreet, et al.
Pubblicazione: (2024)
di: Singh, Parampreet, et al.
Pubblicazione: (2024)
Uncertainty Quantification in Melody Estimation using Histogram Representation
di: Saxena, Kavya Ranjan, et al.
Pubblicazione: (2025)
di: Saxena, Kavya Ranjan, et al.
Pubblicazione: (2025)
Meta-learning-based percussion transcription and $t\bar{a}la$ identification from low-resource audio
di: Kodag, Rahul Bapusaheb, et al.
Pubblicazione: (2025)
di: Kodag, Rahul Bapusaheb, et al.
Pubblicazione: (2025)
Weakly Supervised Tabla Stroke Transcription via TI-SDRM: A Rhythm-Aware Lattice Rescoring Framework
di: Kodag, Rahul Bapusaheb, et al.
Pubblicazione: (2026)
di: Kodag, Rahul Bapusaheb, et al.
Pubblicazione: (2026)
SyncNet: correlating objective for time delay estimation in audio signals
di: Raina, Akshay, et al.
Pubblicazione: (2022)
di: Raina, Akshay, et al.
Pubblicazione: (2022)
Automatic Detection and Analysis of Singing Mistakes for Music Pedagogy
di: Kumar, Sumit, et al.
Pubblicazione: (2026)
di: Kumar, Sumit, et al.
Pubblicazione: (2026)
Written Term Detection Improves Spoken Term Detection
di: Yusuf, Bolaji, et al.
Pubblicazione: (2024)
di: Yusuf, Bolaji, et al.
Pubblicazione: (2024)
Interactive singing melody extraction based on active adaptation
di: Saxena, Kavya Ranjan, et al.
Pubblicazione: (2024)
di: Saxena, Kavya Ranjan, et al.
Pubblicazione: (2024)
$T\bar{a}laGen:$ A System for Automatic $T\bar{a}la$ Identification and Generation
di: Kodag, Rahul Bapusaheb, et al.
Pubblicazione: (2024)
di: Kodag, Rahul Bapusaheb, et al.
Pubblicazione: (2024)
Spoken-Term Discovery using Discrete Speech Units
di: van Niekerk, Benjamin, et al.
Pubblicazione: (2024)
di: van Niekerk, Benjamin, et al.
Pubblicazione: (2024)
HiPPO: Exploring A Novel Hierarchical Pronunciation Assessment Approach for Spoken Languages
di: Yan, Bi-Cheng, et al.
Pubblicazione: (2025)
di: Yan, Bi-Cheng, et al.
Pubblicazione: (2025)
Learning from Limited Labels: Transductive Graph Label Propagation for Indian Music Analysis
di: Singh, Parampreet, et al.
Pubblicazione: (2026)
di: Singh, Parampreet, et al.
Pubblicazione: (2026)
Learning to Discover: A Generalized Framework for Raga Identification without Forgetting
di: Singh, Parampreet, et al.
Pubblicazione: (2026)
di: Singh, Parampreet, et al.
Pubblicazione: (2026)
Towards Hierarchical Spoken Language Dysfluency Modeling
di: Lian, Jiachen, et al.
Pubblicazione: (2024)
di: Lian, Jiachen, et al.
Pubblicazione: (2024)
Recognizing Ornaments in Vocal Indian Art Music with Active Annotation
di: Kumar, Sumit, et al.
Pubblicazione: (2025)
di: Kumar, Sumit, et al.
Pubblicazione: (2025)
TeLeS: Temporal Lexeme Similarity Score to Estimate Confidence in End-to-End ASR
di: Ravi, Nagarathna, et al.
Pubblicazione: (2024)
di: Ravi, Nagarathna, et al.
Pubblicazione: (2024)
Finding Task-specific Subnetworks in Multi-task Spoken Language Understanding Model
di: Futami, Hayato, et al.
Pubblicazione: (2024)
di: Futami, Hayato, et al.
Pubblicazione: (2024)
Acoustic and Semantic Modeling of Emotion in Spoken Language
di: Dutta, Soumya
Pubblicazione: (2026)
di: Dutta, Soumya
Pubblicazione: (2026)
Semantic-Aware Interruption Detection in Spoken Dialogue Systems: Benchmark, Metric, and Model
di: Xia, Kangxiang, et al.
Pubblicazione: (2026)
di: Xia, Kangxiang, et al.
Pubblicazione: (2026)
On The Landscape of Spoken Language Models: A Comprehensive Survey
di: Arora, Siddhant, et al.
Pubblicazione: (2025)
di: Arora, Siddhant, et al.
Pubblicazione: (2025)
Spoken Language Corpora Augmentation with Domain-Specific Voice-Cloned Speech
di: Czyżnikiewicz, Mateusz, et al.
Pubblicazione: (2024)
di: Czyżnikiewicz, Mateusz, et al.
Pubblicazione: (2024)
DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action
di: Zhang, Haoyang, et al.
Pubblicazione: (2026)
di: Zhang, Haoyang, et al.
Pubblicazione: (2026)
Prompting Whisper for QA-driven Zero-shot End-to-end Spoken Language Understanding
di: Li, Mohan, et al.
Pubblicazione: (2024)
di: Li, Mohan, et al.
Pubblicazione: (2024)
E-chat: Emotion-sensitive Spoken Dialogue System with Large Language Models
di: Xue, Hongfei, et al.
Pubblicazione: (2023)
di: Xue, Hongfei, et al.
Pubblicazione: (2023)
Spoken Stereoset: On Evaluating Social Bias Toward Speaker in Speech Large Language Models
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2024)
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2024)
Exploring Text-Queried Sound Event Detection with Audio Source Separation
di: Yin, Han, et al.
Pubblicazione: (2024)
di: Yin, Han, et al.
Pubblicazione: (2024)
On the Evaluation of Speech Foundation Models for Spoken Language Understanding
di: Arora, Siddhant, et al.
Pubblicazione: (2024)
di: Arora, Siddhant, et al.
Pubblicazione: (2024)
ESPnet-SDS: Unified Toolkit and Demo for Spoken Dialogue Systems
di: Arora, Siddhant, et al.
Pubblicazione: (2025)
di: Arora, Siddhant, et al.
Pubblicazione: (2025)
Evaluating Hallucinations in Audio-Visual Multimodal LLMs with Spoken Queries under Diverse Acoustic Conditions
di: Park, Hansol, et al.
Pubblicazione: (2025)
di: Park, Hansol, et al.
Pubblicazione: (2025)
On the use of Performer and Agent Attention for Spoken Language Identification
di: dhiman, Jitendra Kumar, et al.
Pubblicazione: (2025)
di: dhiman, Jitendra Kumar, et al.
Pubblicazione: (2025)
Chain-of-Thought Reasoning in Streaming Full-Duplex End-to-End Spoken Dialogue Systems
di: Arora, Siddhant, et al.
Pubblicazione: (2025)
di: Arora, Siddhant, et al.
Pubblicazione: (2025)
Proactive for Uncertainty: Cause-Aware Error Diagnosis and Interactive Clarification for Spoken Dialogue Systems
di: Peng, Yizhou, et al.
Pubblicazione: (2026)
di: Peng, Yizhou, et al.
Pubblicazione: (2026)
Keyword Mamba: Spoken Keyword Spotting with State Space Models
di: Ding, Hanyu, et al.
Pubblicazione: (2025)
di: Ding, Hanyu, et al.
Pubblicazione: (2025)
Documenti analoghi
-
BEST-STD2.0: Balanced and Efficient Speech Tokenizer for Spoken Term Detection
di: Singh, Anup, et al.
Pubblicazione: (2025) -
Attention-Based Audio Embeddings for Query-by-Example
di: Singh, Anup, et al.
Pubblicazione: (2022) -
BEST-STD: Bidirectional Mamba-Enhanced Speech Tokenization for Spoken Term Detection
di: Singh, Anup, et al.
Pubblicazione: (2024) -
Cross-Lingual Query-by-Example Spoken Term Detection: A Transformer-Based Approach
di: Fatemeh, Allahdadi, et al.
Pubblicazione: (2024) -
Improving Active Learning for Melody Estimation by Disentangling Uncertainties
di: Jaiswal, Aayush, et al.
Pubblicazione: (2025)