Gespeichert in:
| Hauptverfasser: | Baali, Massa, Aldoobi, Abdulhamid, Dhamyal, Hira, Singh, Rita, Raj, Bhiksha |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2409.05799 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
What and When to Learn: CURriculum Ranking Loss for Large-Scale Speaker Verification
von: Baali, Massa, et al.
Veröffentlicht: (2026)
von: Baali, Massa, et al.
Veröffentlicht: (2026)
DELULU: Discriminative Embedding Learning Using Latent Units for Speaker-Aware Self-Trained Speech Foundational Model
von: Baali, Massa, et al.
Veröffentlicht: (2025)
von: Baali, Massa, et al.
Veröffentlicht: (2025)
SVeritas: Benchmark for Robust Speaker Verification under Diverse Conditions
von: Baali, Massa, et al.
Veröffentlicht: (2025)
von: Baali, Massa, et al.
Veröffentlicht: (2025)
Improving Speaker Representations Using Contrastive Losses on Multi-scale Features
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
CoLMbo: Speaker Language Model for Descriptive Profiling
von: Baali, Massa, et al.
Veröffentlicht: (2025)
von: Baali, Massa, et al.
Veröffentlicht: (2025)
Speech vs. Transcript: Does It Matter for Human Annotators in Speech Summarization?
von: Sharma, Roshan, et al.
Veröffentlicht: (2024)
von: Sharma, Roshan, et al.
Veröffentlicht: (2024)
CAARMA: Class Augmentation with Adversarial Mixup Regularization
von: Baali, Massa, et al.
Veröffentlicht: (2025)
von: Baali, Massa, et al.
Veröffentlicht: (2025)
SELM: Enhancing Speech Emotion Recognition for Out-of-Domain Scenarios
von: Bukhari, Hazim, et al.
Veröffentlicht: (2024)
von: Bukhari, Hazim, et al.
Veröffentlicht: (2024)
Objective Measurements of Voice Quality
von: Dhamyal, Hira, et al.
Veröffentlicht: (2024)
von: Dhamyal, Hira, et al.
Veröffentlicht: (2024)
What Do Speech Foundation Models Not Learn About Speech?
von: Waheed, Abdul, et al.
Veröffentlicht: (2024)
von: Waheed, Abdul, et al.
Veröffentlicht: (2024)
Audio Language Model for Deepfake Detection Grounded in Acoustic Chain-of-Thought
von: Chen, Runkun, et al.
Veröffentlicht: (2026)
von: Chen, Runkun, et al.
Veröffentlicht: (2026)
Exploring Multilingual Unseen Speaker Emotion Recognition: Leveraging Co-Attention Cues in Multitask Learning
von: Goel, Arnav, et al.
Veröffentlicht: (2024)
von: Goel, Arnav, et al.
Veröffentlicht: (2024)
Revisiting Acoustic Features for Robust ASR
von: Shah, Muhammad A., et al.
Veröffentlicht: (2024)
von: Shah, Muhammad A., et al.
Veröffentlicht: (2024)
Human Voice is Unique
von: Singh, Rita, et al.
Veröffentlicht: (2025)
von: Singh, Rita, et al.
Veröffentlicht: (2025)
WhisperRT -- Turning Whisper into a Causal Streaming Model
von: Krichli, Tomer, et al.
Veröffentlicht: (2025)
von: Krichli, Tomer, et al.
Veröffentlicht: (2025)
Phonetic Segmentation of the UCLA Phonetics Lab Archive
von: Chodroff, Eleanor, et al.
Veröffentlicht: (2024)
von: Chodroff, Eleanor, et al.
Veröffentlicht: (2024)
uDistil-Whisper: Label-Free Data Filtering for Knowledge Distillation in Low-Data Regimes
von: Waheed, Abdul, et al.
Veröffentlicht: (2024)
von: Waheed, Abdul, et al.
Veröffentlicht: (2024)
Domain Adaptation for Contrastive Audio-Language Models
von: Deshmukh, Soham, et al.
Veröffentlicht: (2024)
von: Deshmukh, Soham, et al.
Veröffentlicht: (2024)
Phonetic Richness for Improved Automatic Speaker Verification
von: Klein, Nicholas, et al.
Veröffentlicht: (2024)
von: Klein, Nicholas, et al.
Veröffentlicht: (2024)
PhiNet: Speaker Verification with Phonetic Interpretability
von: Ma, Yi, et al.
Veröffentlicht: (2026)
von: Ma, Yi, et al.
Veröffentlicht: (2026)
VerLM: Explaining Face Verification Using Natural Language
von: Hannan, Syed Abdul, et al.
Veröffentlicht: (2026)
von: Hannan, Syed Abdul, et al.
Veröffentlicht: (2026)
On the Evaluation of Speech Foundation Models for Spoken Language Understanding
von: Arora, Siddhant, et al.
Veröffentlicht: (2024)
von: Arora, Siddhant, et al.
Veröffentlicht: (2024)
Evaluating and Improving Continual Learning in Spoken Language Understanding
von: Yang, Muqiao, et al.
Veröffentlicht: (2024)
von: Yang, Muqiao, et al.
Veröffentlicht: (2024)
Analysis of Speech Temporal Dynamics in the Context of Speaker Verification and Voice Anonymization
von: Tomashenko, Natalia, et al.
Veröffentlicht: (2024)
von: Tomashenko, Natalia, et al.
Veröffentlicht: (2024)
Psychoacoustic Challenges Of Speech Enhancement On VoIP Platforms
von: Konan, Joseph, et al.
Veröffentlicht: (2023)
von: Konan, Joseph, et al.
Veröffentlicht: (2023)
Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis
von: Zhou, Kun, et al.
Veröffentlicht: (2024)
von: Zhou, Kun, et al.
Veröffentlicht: (2024)
Speech Robust Bench: A Robustness Benchmark For Speech Recognition
von: Shah, Muhammad A., et al.
Veröffentlicht: (2024)
von: Shah, Muhammad A., et al.
Veröffentlicht: (2024)
A Technique for Isolating Lexically-Independent Phonetic Dependencies in Generative CNNs
von: Šegedin, Bruno Ferenc
Veröffentlicht: (2025)
von: Šegedin, Bruno Ferenc
Veröffentlicht: (2025)
PAST: Phonetic-Acoustic Speech Tokenizer
von: Har-Tuv, Nadav, et al.
Veröffentlicht: (2025)
von: Har-Tuv, Nadav, et al.
Veröffentlicht: (2025)
Multimodal Input Aids a Bayesian Model of Phonetic Learning
von: Zhi, Sophia, et al.
Veröffentlicht: (2024)
von: Zhi, Sophia, et al.
Veröffentlicht: (2024)
Disentangling Speaker Traits for Deepfake Source Verification via Chebyshev Polynomial and Riemannian Metric Learning
von: Xuan, Xi, et al.
Veröffentlicht: (2026)
von: Xuan, Xi, et al.
Veröffentlicht: (2026)
CryCeleb: A Speaker Verification Dataset Based on Infant Cry Sounds
von: Budaghyan, David, et al.
Veröffentlicht: (2023)
von: Budaghyan, David, et al.
Veröffentlicht: (2023)
LightCAM: A Fast and Light Implementation of Context-Aware Masking based D-TDNN for Speaker Verification
von: Cao, Di, et al.
Veröffentlicht: (2024)
von: Cao, Di, et al.
Veröffentlicht: (2024)
ADIFF: Explaining audio difference using natural language
von: Deshmukh, Soham, et al.
Veröffentlicht: (2025)
von: Deshmukh, Soham, et al.
Veröffentlicht: (2025)
Tessellated Linear Model for Age Prediction from Voice
von: Alharthi, Dareen, et al.
Veröffentlicht: (2025)
von: Alharthi, Dareen, et al.
Veröffentlicht: (2025)
Mellow: a small audio language model for reasoning
von: Deshmukh, Soham, et al.
Veröffentlicht: (2025)
von: Deshmukh, Soham, et al.
Veröffentlicht: (2025)
Multilingual Prosody Transfer: Comparing Supervised & Transfer Learning
von: Goel, Arnav, et al.
Veröffentlicht: (2024)
von: Goel, Arnav, et al.
Veröffentlicht: (2024)
CrossVoice: Crosslingual Prosody Preserving Cascade-S2ST using Transfer Learning
von: Hira, Medha, et al.
Veröffentlicht: (2024)
von: Hira, Medha, et al.
Veröffentlicht: (2024)
Deciphering GunType Hierarchy through Acoustic Analysis of Gunshot Recordings
von: Shah, Ankit, et al.
Veröffentlicht: (2025)
von: Shah, Ankit, et al.
Veröffentlicht: (2025)
ExPO: Explainable Phonetic Trait-Oriented Network for Speaker Verification
von: Ma, Yi, et al.
Veröffentlicht: (2025)
von: Ma, Yi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
What and When to Learn: CURriculum Ranking Loss for Large-Scale Speaker Verification
von: Baali, Massa, et al.
Veröffentlicht: (2026) -
DELULU: Discriminative Embedding Learning Using Latent Units for Speaker-Aware Self-Trained Speech Foundational Model
von: Baali, Massa, et al.
Veröffentlicht: (2025) -
SVeritas: Benchmark for Robust Speaker Verification under Diverse Conditions
von: Baali, Massa, et al.
Veröffentlicht: (2025) -
Improving Speaker Representations Using Contrastive Losses on Multi-scale Features
von: Dixit, Satvik, et al.
Veröffentlicht: (2024) -
CoLMbo: Speaker Language Model for Descriptive Profiling
von: Baali, Massa, et al.
Veröffentlicht: (2025)