H-QuEST: Accelerating Query-by-Example Spoken Term Detection with Hierarchical Indexing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Singh, Akanksha, Chen, Yi-Ping Phoebe, Arora, Vipul |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BEST-STD2.0: Balanced and Efficient Speech Tokenizer for Spoken Term Detection
von: Singh, Anup, et al.
Veröffentlicht: (2025)
von: Singh, Anup, et al.
Veröffentlicht: (2025)
Attention-Based Audio Embeddings for Query-by-Example
von: Singh, Anup, et al.
Veröffentlicht: (2022)
von: Singh, Anup, et al.
Veröffentlicht: (2022)
BEST-STD: Bidirectional Mamba-Enhanced Speech Tokenization for Spoken Term Detection
von: Singh, Anup, et al.
Veröffentlicht: (2024)
von: Singh, Anup, et al.
Veröffentlicht: (2024)
Cross-Lingual Query-by-Example Spoken Term Detection: A Transformer-Based Approach
von: Fatemeh, Allahdadi, et al.
Veröffentlicht: (2024)
von: Fatemeh, Allahdadi, et al.
Veröffentlicht: (2024)
Improving Active Learning for Melody Estimation by Disentangling Uncertainties
von: Jaiswal, Aayush, et al.
Veröffentlicht: (2025)
von: Jaiswal, Aayush, et al.
Veröffentlicht: (2025)
Explainable Deep Learning Analysis for Raga Identification in Indian Art Music
von: Singh, Parampreet, et al.
Veröffentlicht: (2024)
von: Singh, Parampreet, et al.
Veröffentlicht: (2024)
AudioNet: Supervised Deep Hashing for Retrieval of Similar Audio Events
von: Dutta, Sagar, et al.
Veröffentlicht: (2025)
von: Dutta, Sagar, et al.
Veröffentlicht: (2025)
Identification and Clustering of Unseen Ragas in Indian Art Music
von: Singh, Parampreet, et al.
Veröffentlicht: (2024)
von: Singh, Parampreet, et al.
Veröffentlicht: (2024)
Uncertainty Quantification in Melody Estimation using Histogram Representation
von: Saxena, Kavya Ranjan, et al.
Veröffentlicht: (2025)
von: Saxena, Kavya Ranjan, et al.
Veröffentlicht: (2025)
Meta-learning-based percussion transcription and $t\bar{a}la$ identification from low-resource audio
von: Kodag, Rahul Bapusaheb, et al.
Veröffentlicht: (2025)
von: Kodag, Rahul Bapusaheb, et al.
Veröffentlicht: (2025)
Weakly Supervised Tabla Stroke Transcription via TI-SDRM: A Rhythm-Aware Lattice Rescoring Framework
von: Kodag, Rahul Bapusaheb, et al.
Veröffentlicht: (2026)
von: Kodag, Rahul Bapusaheb, et al.
Veröffentlicht: (2026)
SyncNet: correlating objective for time delay estimation in audio signals
von: Raina, Akshay, et al.
Veröffentlicht: (2022)
von: Raina, Akshay, et al.
Veröffentlicht: (2022)
Automatic Detection and Analysis of Singing Mistakes for Music Pedagogy
von: Kumar, Sumit, et al.
Veröffentlicht: (2026)
von: Kumar, Sumit, et al.
Veröffentlicht: (2026)
Written Term Detection Improves Spoken Term Detection
von: Yusuf, Bolaji, et al.
Veröffentlicht: (2024)
von: Yusuf, Bolaji, et al.
Veröffentlicht: (2024)
Interactive singing melody extraction based on active adaptation
von: Saxena, Kavya Ranjan, et al.
Veröffentlicht: (2024)
von: Saxena, Kavya Ranjan, et al.
Veröffentlicht: (2024)
$T\bar{a}laGen:$ A System for Automatic $T\bar{a}la$ Identification and Generation
von: Kodag, Rahul Bapusaheb, et al.
Veröffentlicht: (2024)
von: Kodag, Rahul Bapusaheb, et al.
Veröffentlicht: (2024)
Spoken-Term Discovery using Discrete Speech Units
von: van Niekerk, Benjamin, et al.
Veröffentlicht: (2024)
von: van Niekerk, Benjamin, et al.
Veröffentlicht: (2024)
HiPPO: Exploring A Novel Hierarchical Pronunciation Assessment Approach for Spoken Languages
von: Yan, Bi-Cheng, et al.
Veröffentlicht: (2025)
von: Yan, Bi-Cheng, et al.
Veröffentlicht: (2025)
Learning from Limited Labels: Transductive Graph Label Propagation for Indian Music Analysis
von: Singh, Parampreet, et al.
Veröffentlicht: (2026)
von: Singh, Parampreet, et al.
Veröffentlicht: (2026)
Learning to Discover: A Generalized Framework for Raga Identification without Forgetting
von: Singh, Parampreet, et al.
Veröffentlicht: (2026)
von: Singh, Parampreet, et al.
Veröffentlicht: (2026)
Towards Hierarchical Spoken Language Dysfluency Modeling
von: Lian, Jiachen, et al.
Veröffentlicht: (2024)
von: Lian, Jiachen, et al.
Veröffentlicht: (2024)
Recognizing Ornaments in Vocal Indian Art Music with Active Annotation
von: Kumar, Sumit, et al.
Veröffentlicht: (2025)
von: Kumar, Sumit, et al.
Veröffentlicht: (2025)
TeLeS: Temporal Lexeme Similarity Score to Estimate Confidence in End-to-End ASR
von: Ravi, Nagarathna, et al.
Veröffentlicht: (2024)
von: Ravi, Nagarathna, et al.
Veröffentlicht: (2024)
Finding Task-specific Subnetworks in Multi-task Spoken Language Understanding Model
von: Futami, Hayato, et al.
Veröffentlicht: (2024)
von: Futami, Hayato, et al.
Veröffentlicht: (2024)
Acoustic and Semantic Modeling of Emotion in Spoken Language
von: Dutta, Soumya
Veröffentlicht: (2026)
von: Dutta, Soumya
Veröffentlicht: (2026)
Semantic-Aware Interruption Detection in Spoken Dialogue Systems: Benchmark, Metric, and Model
von: Xia, Kangxiang, et al.
Veröffentlicht: (2026)
von: Xia, Kangxiang, et al.
Veröffentlicht: (2026)
On The Landscape of Spoken Language Models: A Comprehensive Survey
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
Spoken Language Corpora Augmentation with Domain-Specific Voice-Cloned Speech
von: Czyżnikiewicz, Mateusz, et al.
Veröffentlicht: (2024)
von: Czyżnikiewicz, Mateusz, et al.
Veröffentlicht: (2024)
DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action
von: Zhang, Haoyang, et al.
Veröffentlicht: (2026)
von: Zhang, Haoyang, et al.
Veröffentlicht: (2026)
Prompting Whisper for QA-driven Zero-shot End-to-end Spoken Language Understanding
von: Li, Mohan, et al.
Veröffentlicht: (2024)
von: Li, Mohan, et al.
Veröffentlicht: (2024)
E-chat: Emotion-sensitive Spoken Dialogue System with Large Language Models
von: Xue, Hongfei, et al.
Veröffentlicht: (2023)
von: Xue, Hongfei, et al.
Veröffentlicht: (2023)
Spoken Stereoset: On Evaluating Social Bias Toward Speaker in Speech Large Language Models
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2024)
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2024)
Exploring Text-Queried Sound Event Detection with Audio Source Separation
von: Yin, Han, et al.
Veröffentlicht: (2024)
von: Yin, Han, et al.
Veröffentlicht: (2024)
On the Evaluation of Speech Foundation Models for Spoken Language Understanding
von: Arora, Siddhant, et al.
Veröffentlicht: (2024)
von: Arora, Siddhant, et al.
Veröffentlicht: (2024)
ESPnet-SDS: Unified Toolkit and Demo for Spoken Dialogue Systems
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
Evaluating Hallucinations in Audio-Visual Multimodal LLMs with Spoken Queries under Diverse Acoustic Conditions
von: Park, Hansol, et al.
Veröffentlicht: (2025)
von: Park, Hansol, et al.
Veröffentlicht: (2025)
On the use of Performer and Agent Attention for Spoken Language Identification
von: dhiman, Jitendra Kumar, et al.
Veröffentlicht: (2025)
von: dhiman, Jitendra Kumar, et al.
Veröffentlicht: (2025)
Chain-of-Thought Reasoning in Streaming Full-Duplex End-to-End Spoken Dialogue Systems
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
Proactive for Uncertainty: Cause-Aware Error Diagnosis and Interactive Clarification for Spoken Dialogue Systems
von: Peng, Yizhou, et al.
Veröffentlicht: (2026)
von: Peng, Yizhou, et al.
Veröffentlicht: (2026)
Keyword Mamba: Spoken Keyword Spotting with State Space Models
von: Ding, Hanyu, et al.
Veröffentlicht: (2025)
von: Ding, Hanyu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
BEST-STD2.0: Balanced and Efficient Speech Tokenizer for Spoken Term Detection
von: Singh, Anup, et al.
Veröffentlicht: (2025) -
Attention-Based Audio Embeddings for Query-by-Example
von: Singh, Anup, et al.
Veröffentlicht: (2022) -
BEST-STD: Bidirectional Mamba-Enhanced Speech Tokenization for Spoken Term Detection
von: Singh, Anup, et al.
Veröffentlicht: (2024) -
Cross-Lingual Query-by-Example Spoken Term Detection: A Transformer-Based Approach
von: Fatemeh, Allahdadi, et al.
Veröffentlicht: (2024) -
Improving Active Learning for Melody Estimation by Disentangling Uncertainties
von: Jaiswal, Aayush, et al.
Veröffentlicht: (2025)