Cross-Lingual Query-by-Example Spoken Term Detection: A Transformer-Based Approach
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fatemeh, Allahdadi, Rahil, Mahdian Toroghi, Hassan, Zareian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Enhancement of Dysarthric Speech Reconstruction by Contrastive Learning
von: Fatemeh, Keshvari, et al.
Veröffentlicht: (2024)
von: Fatemeh, Keshvari, et al.
Veröffentlicht: (2024)
A Layer-Anchoring Strategy for Enhancing Cross-Lingual Speech Emotion Recognition
von: Upadhyay, Shreya G., et al.
Veröffentlicht: (2024)
von: Upadhyay, Shreya G., et al.
Veröffentlicht: (2024)
Comparative Analysis of Mel-Frequency Cepstral Coefficients and Wavelet Based Audio Signal Processing for Emotion Detection and Mental Health Assessment in Spoken Speech
von: Agbo, Idoko, et al.
Veröffentlicht: (2024)
von: Agbo, Idoko, et al.
Veröffentlicht: (2024)
LMD: A Learnable Mask Network to Detect Adversarial Examples for Speaker Verification
von: Chen, Xing, et al.
Veröffentlicht: (2022)
von: Chen, Xing, et al.
Veröffentlicht: (2022)
COVID-19 Detection System: A Comparative Analysis of System Performance Based on Acoustic Features of Cough Audio Signals
von: Shati, Asmaa, et al.
Veröffentlicht: (2023)
von: Shati, Asmaa, et al.
Veröffentlicht: (2023)
Zero-Shot Multi-Lingual Speaker Verification in Clinical Trials
von: Akram, Ali, et al.
Veröffentlicht: (2024)
von: Akram, Ali, et al.
Veröffentlicht: (2024)
Attention-Based Audio Embeddings for Query-by-Example
von: Singh, Anup, et al.
Veröffentlicht: (2022)
von: Singh, Anup, et al.
Veröffentlicht: (2022)
Unsupervised ASR via Cross-Lingual Pseudo-Labeling
von: Likhomanenko, Tatiana, et al.
Veröffentlicht: (2023)
von: Likhomanenko, Tatiana, et al.
Veröffentlicht: (2023)
A Novel Stochastic Transformer-based Approach for Post-Traumatic Stress Disorder Detection using Audio Recording of Clinical Interviews
von: Dia, Mamadou, et al.
Veröffentlicht: (2024)
von: Dia, Mamadou, et al.
Veröffentlicht: (2024)
Identification of Cognitive Decline from Spoken Language through Feature Selection and the Bag of Acoustic Words Model
von: Niemelä, Marko, et al.
Veröffentlicht: (2024)
von: Niemelä, Marko, et al.
Veröffentlicht: (2024)
Description-based Controllable Text-to-Speech with Cross-Lingual Voice Control
von: Yamamoto, Ryuichi, et al.
Veröffentlicht: (2024)
von: Yamamoto, Ryuichi, et al.
Veröffentlicht: (2024)
Cross-Referencing Self-Training Network for Sound Event Detection in Audio Mixtures
von: Park, Sangwook, et al.
Veröffentlicht: (2021)
von: Park, Sangwook, et al.
Veröffentlicht: (2021)
YourMT3+: Multi-instrument Music Transcription with Enhanced Transformer Architectures and Cross-dataset Stem Augmentation
von: Chang, Sungkyun, et al.
Veröffentlicht: (2024)
von: Chang, Sungkyun, et al.
Veröffentlicht: (2024)
Exploring Transformer-Based Music Overpainting for Jazz Piano Variations
von: Row, Eleanor, et al.
Veröffentlicht: (2024)
von: Row, Eleanor, et al.
Veröffentlicht: (2024)
Blind Audio Bandwidth Extension: A Diffusion-Based Zero-Shot Approach
von: Moliner, Eloi, et al.
Veröffentlicht: (2023)
von: Moliner, Eloi, et al.
Veröffentlicht: (2023)
Towards Robust Overlapping Speech Detection: A Speaker-Aware Progressive Approach Using WavLM
von: Sun, Zhaokai, et al.
Veröffentlicht: (2025)
von: Sun, Zhaokai, et al.
Veröffentlicht: (2025)
MAIA: An Inpainting-Based Approach for Music Adversarial Attacks
von: Liu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Liu, Yuxuan, et al.
Veröffentlicht: (2025)
Mouth Articulation-Based Anchoring for Improved Cross-Corpus Speech Emotion Recognition
von: Upadhyay, Shreya G., et al.
Veröffentlicht: (2024)
von: Upadhyay, Shreya G., et al.
Veröffentlicht: (2024)
Classification of Short Segment Pediatric Heart Sounds Based on a Transformer-Based Convolutional Neural Network
von: Hassanuzzaman, Md, et al.
Veröffentlicht: (2024)
von: Hassanuzzaman, Md, et al.
Veröffentlicht: (2024)
SpokeN-100: A Cross-Lingual Benchmarking Dataset for The Classification of Spoken Numbers in Different Languages
von: Groh, René, et al.
Veröffentlicht: (2024)
von: Groh, René, et al.
Veröffentlicht: (2024)
Medical Spoken Named Entity Recognition
von: Le-Duc, Khai, et al.
Veröffentlicht: (2024)
von: Le-Duc, Khai, et al.
Veröffentlicht: (2024)
LMU-Based Sequential Learning and Posterior Ensemble Fusion for Cross-Domain Infant Cry Classification
von: Jazaeri, Niloofar, et al.
Veröffentlicht: (2026)
von: Jazaeri, Niloofar, et al.
Veröffentlicht: (2026)
Speech Emotion Detection Based on MFCC and CNN-LSTM Architecture
von: Ouyang, Qianhe
Veröffentlicht: (2025)
von: Ouyang, Qianhe
Veröffentlicht: (2025)
Temporal Convolution-based Hybrid Model Approach with Representation Learning for Real-Time Acoustic Anomaly Detection
von: Dissanayaka, Sahan, et al.
Veröffentlicht: (2024)
von: Dissanayaka, Sahan, et al.
Veröffentlicht: (2024)
End-to-End Spoken Grammatical Error Correction
von: Qian, Mengjie, et al.
Veröffentlicht: (2025)
von: Qian, Mengjie, et al.
Veröffentlicht: (2025)
An Attention Long Short-Term Memory based system for automatic classification of speech intelligibility
von: Fernández-Díaz, Miguel, et al.
Veröffentlicht: (2024)
von: Fernández-Díaz, Miguel, et al.
Veröffentlicht: (2024)
Aligning Spoken Dialogue Models from User Interactions
von: Wu, Anne, et al.
Veröffentlicht: (2025)
von: Wu, Anne, et al.
Veröffentlicht: (2025)
USM-SCD: Multilingual Speaker Change Detection Based on Large Pretrained Foundation Models
von: Zhao, Guanlong, et al.
Veröffentlicht: (2023)
von: Zhao, Guanlong, et al.
Veröffentlicht: (2023)
An Enhanced Audio Feature Tailored for Anomalous Sound Detection Based on Pre-trained Models
von: Zhong, Guirui, et al.
Veröffentlicht: (2025)
von: Zhong, Guirui, et al.
Veröffentlicht: (2025)
MiSTR: Multi-Modal iEEG-to-Speech Synthesis with Transformer-Based Prosody Prediction and Neural Phase Reconstruction
von: Al-Radhi, Mohammed Salah, et al.
Veröffentlicht: (2025)
von: Al-Radhi, Mohammed Salah, et al.
Veröffentlicht: (2025)
Anticipatory Music Transformer
von: Thickstun, John, et al.
Veröffentlicht: (2023)
von: Thickstun, John, et al.
Veröffentlicht: (2023)
Recurrence-Based Nonlinear Vocal Dynamics as Digital Biomarkers for Depression Detection from Conversational Speech
von: Samanta, Himadri S
Veröffentlicht: (2026)
von: Samanta, Himadri S
Veröffentlicht: (2026)
DC-Spin: A Speaker-invariant Speech Tokenizer for Spoken Language Models
von: Chang, Heng-Jui, et al.
Veröffentlicht: (2024)
von: Chang, Heng-Jui, et al.
Veröffentlicht: (2024)
Vocal Tract Length Warped Features for Spoken Keyword Spotting
von: Sarkar, Achintya kr., et al.
Veröffentlicht: (2025)
von: Sarkar, Achintya kr., et al.
Veröffentlicht: (2025)
Enhancing Neural Spoken Language Recognition: An Exploration with Multilingual Datasets
von: Anidjar, Or Haim, et al.
Veröffentlicht: (2025)
von: Anidjar, Or Haim, et al.
Veröffentlicht: (2025)
RingFormer: A Neural Vocoder with Ring Attention and Convolution-Augmented Transformer
von: Hong, Seongho, et al.
Veröffentlicht: (2025)
von: Hong, Seongho, et al.
Veröffentlicht: (2025)
A Recurrent Neural Network Approach to the Answering Machine Detection Problem
von: Altwlkany, Kemal, et al.
Veröffentlicht: (2024)
von: Altwlkany, Kemal, et al.
Veröffentlicht: (2024)
Zero-shot Voice Conversion with Diffusion Transformers
von: Liu, Songting
Veröffentlicht: (2024)
von: Liu, Songting
Veröffentlicht: (2024)
Zero-Shot End-To-End Spoken Question Answering In Medical Domain
von: Labrak, Yanis, et al.
Veröffentlicht: (2024)
von: Labrak, Yanis, et al.
Veröffentlicht: (2024)
Cross-domain Sound Recognition for Efficient Underwater Data Analysis
von: Park, Jeongsoo, et al.
Veröffentlicht: (2023)
von: Park, Jeongsoo, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Enhancement of Dysarthric Speech Reconstruction by Contrastive Learning
von: Fatemeh, Keshvari, et al.
Veröffentlicht: (2024) -
A Layer-Anchoring Strategy for Enhancing Cross-Lingual Speech Emotion Recognition
von: Upadhyay, Shreya G., et al.
Veröffentlicht: (2024) -
Comparative Analysis of Mel-Frequency Cepstral Coefficients and Wavelet Based Audio Signal Processing for Emotion Detection and Mental Health Assessment in Spoken Speech
von: Agbo, Idoko, et al.
Veröffentlicht: (2024) -
LMD: A Learnable Mask Network to Detect Adversarial Examples for Speaker Verification
von: Chen, Xing, et al.
Veröffentlicht: (2022) -
COVID-19 Detection System: A Comparative Analysis of System Performance Based on Acoustic Features of Cough Audio Signals
von: Shati, Asmaa, et al.
Veröffentlicht: (2023)