Towards Speaker Identification with Minimal Dataset and Constrained Resources using 1D-Convolution Neural Network
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shahan, Irfan Nafiz, Auvi, Pulok Ahmed |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multi-Speaker Conversational Audio Deepfake: Taxonomy, Dataset and Pilot Study
von: Ahmed, Alabi, et al.
Veröffentlicht: (2026)
von: Ahmed, Alabi, et al.
Veröffentlicht: (2026)
Leveraging Speaker Embeddings in End-to-End Neural Diarization for Two-Speaker Scenarios
von: Alvarez-Trejos, Juan Ignacio, et al.
Veröffentlicht: (2024)
von: Alvarez-Trejos, Juan Ignacio, et al.
Veröffentlicht: (2024)
Deep Learning for Speaker Identification: Architectural Insights from AB-1 Corpus Analysis and Performance Evaluation
von: Bartolo, Matthias
Veröffentlicht: (2024)
von: Bartolo, Matthias
Veröffentlicht: (2024)
Acoustic Identification of Ae. aegypti Mosquitoes using Smartphone Apps and Residual Convolutional Neural Networks
von: Paim, Kayuã Oleques, et al.
Veröffentlicht: (2023)
von: Paim, Kayuã Oleques, et al.
Veröffentlicht: (2023)
Disentangling Age and Identity with a Mutual Information Minimization Approach for Cross-Age Speaker Verification
von: Zhang, Fengrun, et al.
Veröffentlicht: (2024)
von: Zhang, Fengrun, et al.
Veröffentlicht: (2024)
Raw Audio Classification with Cosine Convolutional Neural Network (CosCovNN)
von: Haque, Kazi Nazmul, et al.
Veröffentlicht: (2024)
von: Haque, Kazi Nazmul, et al.
Veröffentlicht: (2024)
Quranic Audio Dataset: Crowdsourced and Labeled Recitation from Non-Arabic Speakers
von: Salameh, Raghad, et al.
Veröffentlicht: (2024)
von: Salameh, Raghad, et al.
Veröffentlicht: (2024)
Planing It by Ear: Convolutional Neural Networks for Acoustic Anomaly Detection in Industrial Wood Planers
von: Deschênes, Anthony, et al.
Veröffentlicht: (2025)
von: Deschênes, Anthony, et al.
Veröffentlicht: (2025)
Developing an Effective Training Dataset to Enhance the Performance of AI-based Speaker Separation Systems
von: Melhem, Rawad, et al.
Veröffentlicht: (2024)
von: Melhem, Rawad, et al.
Veröffentlicht: (2024)
Real-Time Pitch/F0 Detection Using Spectrogram Images and Convolutional Neural Networks
von: Zhao, Xufang, et al.
Veröffentlicht: (2025)
von: Zhao, Xufang, et al.
Veröffentlicht: (2025)
Audio-to-Image Encoding for Improved Voice Characteristic Detection Using Deep Convolutional Neural Networks
von: Atif, Youness
Veröffentlicht: (2025)
von: Atif, Youness
Veröffentlicht: (2025)
Speaker Embeddings to Improve Tracking of Intermittent and Moving Speakers
von: Iatariene, Taous, et al.
Veröffentlicht: (2025)
von: Iatariene, Taous, et al.
Veröffentlicht: (2025)
Improving Neural Diarization through Speaker Attribute Attractors and Local Dependency Modeling
von: Palzer, David, et al.
Veröffentlicht: (2025)
von: Palzer, David, et al.
Veröffentlicht: (2025)
Towards Low-Latency Tracking of Multiple Speakers With Short-Context Speaker Embeddings
von: Iatariene, Taous, et al.
Veröffentlicht: (2025)
von: Iatariene, Taous, et al.
Veröffentlicht: (2025)
ExPO: Explainable Phonetic Trait-Oriented Network for Speaker Verification
von: Ma, Yi, et al.
Veröffentlicht: (2025)
von: Ma, Yi, et al.
Veröffentlicht: (2025)
Compressing Quaternion Convolutional Neural Networks for Audio Classification
von: Singh, Arshdeep, et al.
Veröffentlicht: (2025)
von: Singh, Arshdeep, et al.
Veröffentlicht: (2025)
Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTC
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
Memory-Efficient Training for Deep Speaker Embedding Learning in Speaker Verification
von: Liu, Bei, et al.
Veröffentlicht: (2024)
von: Liu, Bei, et al.
Veröffentlicht: (2024)
ML-SAN: Multi-Level Speaker-Adaptive Network for Emotion Recognition in Conversations
von: Wang, Kexue, et al.
Veröffentlicht: (2026)
von: Wang, Kexue, et al.
Veröffentlicht: (2026)
Pretraining Multi-Speaker Identification for Neural Speaker Diarization
von: Horiguchi, Shota, et al.
Veröffentlicht: (2025)
von: Horiguchi, Shota, et al.
Veröffentlicht: (2025)
A Multi-task Learning Balanced Attention Convolutional Neural Network Model for Few-shot Underwater Acoustic Target Recognition
von: Huang, Wei, et al.
Veröffentlicht: (2025)
von: Huang, Wei, et al.
Veröffentlicht: (2025)
Comprehensive Evaluation of CNN-Based Audio Tagging Models on Resource-Constrained Devices
von: Grau-Haro, Jordi, et al.
Veröffentlicht: (2025)
von: Grau-Haro, Jordi, et al.
Veröffentlicht: (2025)
Removing Speaker Information from Speech Representation using Variable-Length Soft Pooling
von: Hwang, Injune, et al.
Veröffentlicht: (2024)
von: Hwang, Injune, et al.
Veröffentlicht: (2024)
Speaker Diarization with Overlapping Community Detection Using Graph Attention Networks and Label Propagation Algorithm
von: Li, Zhaoyang, et al.
Veröffentlicht: (2025)
von: Li, Zhaoyang, et al.
Veröffentlicht: (2025)
Explainable Attribute-Based Speaker Verification
von: Wu, Xiaoliang, et al.
Veröffentlicht: (2024)
von: Wu, Xiaoliang, et al.
Veröffentlicht: (2024)
From Modular to End-to-End Speaker Diarization
von: Landini, Federico
Veröffentlicht: (2024)
von: Landini, Federico
Veröffentlicht: (2024)
Certification of Speaker Recognition Models to Additive Perturbations
von: Korzh, Dmitrii, et al.
Veröffentlicht: (2024)
von: Korzh, Dmitrii, et al.
Veröffentlicht: (2024)
Whisper Speaker Identification: Leveraging Pre-Trained Multilingual Transformers for Robust Speaker Embeddings
von: Emon, Jakaria Islam, et al.
Veröffentlicht: (2025)
von: Emon, Jakaria Islam, et al.
Veröffentlicht: (2025)
SDBench: A Comprehensive Benchmark Suite for Speaker Diarization
von: Pacheco, Eduardo, et al.
Veröffentlicht: (2025)
von: Pacheco, Eduardo, et al.
Veröffentlicht: (2025)
The VoxCeleb Speaker Recognition Challenge: A Retrospective
von: Huh, Jaesung, et al.
Veröffentlicht: (2024)
von: Huh, Jaesung, et al.
Veröffentlicht: (2024)
Bangla-WhisperDiar: Fine-Tuning Whisper and PyAnnote for Bangla Long-Form Speech Recognition and Speaker Diarization
von: Bhuiyan, Mohammed Aman, et al.
Veröffentlicht: (2026)
von: Bhuiyan, Mohammed Aman, et al.
Veröffentlicht: (2026)
End-to-End Supervised Hierarchical Graph Clustering for Speaker Diarization
von: Singh, Prachi, et al.
Veröffentlicht: (2024)
von: Singh, Prachi, et al.
Veröffentlicht: (2024)
LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec
von: Guo, Yiwei, et al.
Veröffentlicht: (2024)
von: Guo, Yiwei, et al.
Veröffentlicht: (2024)
Unispeaker: A Unified Approach for Multimodality-driven Speaker Generation
von: Sheng, Zhengyan, et al.
Veröffentlicht: (2025)
von: Sheng, Zhengyan, et al.
Veröffentlicht: (2025)
DART: Disentanglement of Accent and Speaker Representation in Multispeaker Text-to-Speech
von: Melechovsky, Jan, et al.
Veröffentlicht: (2024)
von: Melechovsky, Jan, et al.
Veröffentlicht: (2024)
Asynchronous Voice Anonymization Using Adversarial Perturbation On Speaker Embedding
von: Wang, Rui, et al.
Veröffentlicht: (2024)
von: Wang, Rui, et al.
Veröffentlicht: (2024)
Evaluating Speaker Identity Coding in Self-supervised Models and Humans
von: Elbanna, Gasser
Veröffentlicht: (2024)
von: Elbanna, Gasser
Veröffentlicht: (2024)
BanglaFake: Constructing and Evaluating a Specialized Bengali Deepfake Audio Dataset
von: Fahad, Istiaq Ahmed, et al.
Veröffentlicht: (2025)
von: Fahad, Istiaq Ahmed, et al.
Veröffentlicht: (2025)
Towards Improving Speaker Distance Estimation through Generative Impulse Response Augmentation
von: Ratnarajah, Anton, et al.
Veröffentlicht: (2026)
von: Ratnarajah, Anton, et al.
Veröffentlicht: (2026)
Speech-Forensics: Towards Comprehensive Synthetic Speech Dataset Establishment and Analysis
von: Ji, Zhoulin, et al.
Veröffentlicht: (2024)
von: Ji, Zhoulin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Multi-Speaker Conversational Audio Deepfake: Taxonomy, Dataset and Pilot Study
von: Ahmed, Alabi, et al.
Veröffentlicht: (2026) -
Leveraging Speaker Embeddings in End-to-End Neural Diarization for Two-Speaker Scenarios
von: Alvarez-Trejos, Juan Ignacio, et al.
Veröffentlicht: (2024) -
Deep Learning for Speaker Identification: Architectural Insights from AB-1 Corpus Analysis and Performance Evaluation
von: Bartolo, Matthias
Veröffentlicht: (2024) -
Acoustic Identification of Ae. aegypti Mosquitoes using Smartphone Apps and Residual Convolutional Neural Networks
von: Paim, Kayuã Oleques, et al.
Veröffentlicht: (2023) -
Disentangling Age and Identity with a Mutual Information Minimization Approach for Cross-Age Speaker Verification
von: Zhang, Fengrun, et al.
Veröffentlicht: (2024)