BEST-STD: Bidirectional Mamba-Enhanced Speech Tokenization for Spoken Term Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Singh, Anup, Demuynck, Kris, Arora, Vipul |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BEST-STD2.0: Balanced and Efficient Speech Tokenizer for Spoken Term Detection
by: Singh, Anup, et al.
Published: (2025)
by: Singh, Anup, et al.
Published: (2025)
Attention-Based Audio Embeddings for Query-by-Example
by: Singh, Anup, et al.
Published: (2022)
by: Singh, Anup, et al.
Published: (2022)
SpeechDPR: End-to-End Spoken Passage Retrieval for Open-Domain Spoken Question Answering
by: Lin, Chyi-Jiunn, et al.
Published: (2024)
by: Lin, Chyi-Jiunn, et al.
Published: (2024)
Harmonic Summation-Based Robust Pitch Estimation in Noisy and Reverberant Environments
by: Singh, Anup, et al.
Published: (2025)
by: Singh, Anup, et al.
Published: (2025)
CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval
by: Abootorabi, Mohammad Mahdi, et al.
Published: (2024)
by: Abootorabi, Mohammad Mahdi, et al.
Published: (2024)
H-QuEST: Accelerating Query-by-Example Spoken Term Detection with Hierarchical Indexing
by: Singh, Akanksha, et al.
Published: (2025)
by: Singh, Akanksha, et al.
Published: (2025)
Written Term Detection Improves Spoken Term Detection
by: Yusuf, Bolaji, et al.
Published: (2024)
by: Yusuf, Bolaji, et al.
Published: (2024)
Scaling Spoken Language Models with Syllabic Speech Tokenization
by: Lee, Nicholas, et al.
Published: (2025)
by: Lee, Nicholas, et al.
Published: (2025)
Transforming LLMs into Cross-modal and Cross-lingual Retrieval Systems
by: Gomez, Frank Palma, et al.
Published: (2024)
by: Gomez, Frank Palma, et al.
Published: (2024)
Analyzing Byte-Pair Encoding on Monophonic and Polyphonic Symbolic Music: A Focus on Musical Phrase Segmentation
by: Le, Dinh-Viet-Toan, et al.
Published: (2024)
by: Le, Dinh-Viet-Toan, et al.
Published: (2024)
Multi-Modal Retrieval For Large Language Model Based Speech Recognition
by: Kolehmainen, Jari, et al.
Published: (2024)
by: Kolehmainen, Jari, et al.
Published: (2024)
VoxRAG: A Step Toward Transcription-Free RAG Systems in Spoken Question Answering
by: Rackauckas, Zackary, et al.
Published: (2025)
by: Rackauckas, Zackary, et al.
Published: (2025)
Multi-Axis Speech Similarity via Factor-Partitioned Embeddings
by: O'Regan, Jim, et al.
Published: (2026)
by: O'Regan, Jim, et al.
Published: (2026)
Navigating Speech Recording Collections with AI-Generated Illustrations
by: Håland, Sirina, et al.
Published: (2025)
by: Håland, Sirina, et al.
Published: (2025)
Speaker Embeddings With Weakly Supervised Voice Activity Detection For Efficient Speaker Diarization
by: Thienpondt, Jenthe, et al.
Published: (2024)
by: Thienpondt, Jenthe, et al.
Published: (2024)
ECAPA2: A Hybrid Neural Network Architecture and Training Strategy for Robust Speaker Embeddings
by: Thienpondt, Jenthe, et al.
Published: (2024)
by: Thienpondt, Jenthe, et al.
Published: (2024)
TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
by: Tseng, Liang-Hsuan, et al.
Published: (2025)
by: Tseng, Liang-Hsuan, et al.
Published: (2025)
Abstractive summarization from Audio Transcription
by: Derkach, Ilia
Published: (2024)
by: Derkach, Ilia
Published: (2024)
Evaluating Interval-based Tokenization for Pitch Representation in Symbolic Music Analysis
by: Le, Dinh-Viet-Toan, et al.
Published: (2025)
by: Le, Dinh-Viet-Toan, et al.
Published: (2025)
Weakly Supervised Phonological Features for Pathological Speech Analysis
by: Thienpondt, Jenthe, et al.
Published: (2025)
by: Thienpondt, Jenthe, et al.
Published: (2025)
On the Evaluation of Speech Foundation Models for Spoken Language Understanding
by: Arora, Siddhant, et al.
Published: (2024)
by: Arora, Siddhant, et al.
Published: (2024)
I can listen but cannot read: An evaluation of two-tower multimodal systems for instrument recognition
by: Vasilakis, Yannis, et al.
Published: (2024)
by: Vasilakis, Yannis, et al.
Published: (2024)
A GEN AI Framework for Medical Note Generation
by: Leong, Hui Yi, et al.
Published: (2024)
by: Leong, Hui Yi, et al.
Published: (2024)
More than words: Advancements and challenges in speech recognition for singing
by: Kruspe, Anna
Published: (2024)
by: Kruspe, Anna
Published: (2024)
WikiMuTe: A web-sourced dataset of semantic descriptions for music audio
by: Weck, Benno, et al.
Published: (2023)
by: Weck, Benno, et al.
Published: (2023)
Beyond Musical Descriptors: Extracting Preference-Bearing Intent in Music Queries
by: Baranes, Marion, et al.
Published: (2026)
by: Baranes, Marion, et al.
Published: (2026)
Technical Report on classification of literature related to children speech disorder
by: Wang, Ziang, et al.
Published: (2025)
by: Wang, Ziang, et al.
Published: (2025)
wav2graph: A Framework for Supervised Learning Knowledge Graph from Speech
by: Le-Duc, Khai, et al.
Published: (2024)
by: Le-Duc, Khai, et al.
Published: (2024)
Finding Task-specific Subnetworks in Multi-task Spoken Language Understanding Model
by: Futami, Hayato, et al.
Published: (2024)
by: Futami, Hayato, et al.
Published: (2024)
ZeroSyl: Simple Zero-Resource Syllable Tokenization for Spoken Language Modeling
by: Visser, Nicol, et al.
Published: (2026)
by: Visser, Nicol, et al.
Published: (2026)
PolySinger: Singing-Voice to Singing-Voice Translation from English to Japanese
by: Antonisen, Silas, et al.
Published: (2024)
by: Antonisen, Silas, et al.
Published: (2024)
Melody-Lyrics Matching with Contrastive Alignment Loss
by: Wang, Changhong, et al.
Published: (2025)
by: Wang, Changhong, et al.
Published: (2025)
Audio Prototypical Network For Controllable Music Recommendation
by: Öncel, Fırat, et al.
Published: (2025)
by: Öncel, Fırat, et al.
Published: (2025)
Evaluating Emotion Recognition in Spoken Language Models on Emotionally Incongruent Speech
by: Corrêa, Pedro, et al.
Published: (2025)
by: Corrêa, Pedro, et al.
Published: (2025)
Predictive Speech Recognition and End-of-Utterance Detection Towards Spoken Dialog Systems
by: Zink, Oswald, et al.
Published: (2024)
by: Zink, Oswald, et al.
Published: (2024)
A Large Dataset of Spontaneous Speech with the Accent Spoken in São Paulo for Automatic Speech Recognition Evaluation
by: Lima, Rodrigo, et al.
Published: (2024)
by: Lima, Rodrigo, et al.
Published: (2024)
MTLM: Incorporating Bidirectional Text Information to Enhance Language Model Training in Speech Recognition Systems
by: Meng, Qingliang, et al.
Published: (2025)
by: Meng, Qingliang, et al.
Published: (2025)
Hypothesis Clustering and Merging: Novel MultiTalker Speech Recognition with Speaker Tokens
by: Kashiwagi, Yosuke, et al.
Published: (2024)
by: Kashiwagi, Yosuke, et al.
Published: (2024)
Paralinguistics-Enhanced Large Language Modeling of Spoken Dialogue
by: Lin, Guan-Ting, et al.
Published: (2023)
by: Lin, Guan-Ting, et al.
Published: (2023)
Kanade: A Simple Disentangled Tokenizer for Spoken Language Modeling
by: Huang, Zhijie, et al.
Published: (2026)
by: Huang, Zhijie, et al.
Published: (2026)
Similar Items
-
BEST-STD2.0: Balanced and Efficient Speech Tokenizer for Spoken Term Detection
by: Singh, Anup, et al.
Published: (2025) -
Attention-Based Audio Embeddings for Query-by-Example
by: Singh, Anup, et al.
Published: (2022) -
SpeechDPR: End-to-End Spoken Passage Retrieval for Open-Domain Spoken Question Answering
by: Lin, Chyi-Jiunn, et al.
Published: (2024) -
Harmonic Summation-Based Robust Pitch Estimation in Noisy and Reverberant Environments
by: Singh, Anup, et al.
Published: (2025) -
CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval
by: Abootorabi, Mohammad Mahdi, et al.
Published: (2024)