Gespeichert in:
| Hauptverfasser: | Rahman, Md. Abdur, Thuseethan, Selvarajah, Yeo, Kheng Cher, Mohamed, Reem E., Azam, Sami |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2510.00522 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
An Innovative Coverage Path Planning Approach for UAVs to Boost Precision Agriculture and Rescue Operations
von: Nur Mohammad Fahad, et al.
Veröffentlicht: (2025)
von: Nur Mohammad Fahad, et al.
Veröffentlicht: (2025)
BioAutoML-NAS: An End-to-End AutoML Framework for Multimodal Insect Classification via Neural Architecture Search on Large-Scale Biodiversity Data
von: Abian, Arefin Ittesafun, et al.
Veröffentlicht: (2025)
von: Abian, Arefin Ittesafun, et al.
Veröffentlicht: (2025)
DeepAgent: A Dual Stream Multi Agent Fusion for Robust Multimodal Deepfake Detection
von: Zaman, Sayeem Been, et al.
Veröffentlicht: (2025)
von: Zaman, Sayeem Been, et al.
Veröffentlicht: (2025)
Predicting Postresection Colorectal Liver Metastases Recurrence Using Advanced Graph Neural Networks with Explainability and Causal Inference
von: Jubair Ahmed, et al.
Veröffentlicht: (2025)
von: Jubair Ahmed, et al.
Veröffentlicht: (2025)
Learning to Weigh Waste: A Physics-Informed Multimodal Fusion Framework and Large-Scale Dataset for Commercial and Industrial Applications
von: Islam, Md. Adnanul, et al.
Veröffentlicht: (2026)
von: Islam, Md. Adnanul, et al.
Veröffentlicht: (2026)
Leveraging Self-supervised Audio Representations for Data-Efficient Acoustic Scene Classification
von: Cai, Yiqiang, et al.
Veröffentlicht: (2024)
von: Cai, Yiqiang, et al.
Veröffentlicht: (2024)
A Source-Free Approach for Domain Adaptation via Multiview Image Transformation and Latent Space Consistency
von: Sutradhar, Debopom, et al.
Veröffentlicht: (2026)
von: Sutradhar, Debopom, et al.
Veröffentlicht: (2026)
Masked Latent Prediction and Classification for Self-Supervised Audio Representation Learning
von: Quelennec, Aurian, et al.
Veröffentlicht: (2025)
von: Quelennec, Aurian, et al.
Veröffentlicht: (2025)
From Birdsong to Rumbles: Classifying Elephant Calls with Out-of-Species Embeddings
von: Geldenhuys, Christiaan M., et al.
Veröffentlicht: (2026)
von: Geldenhuys, Christiaan M., et al.
Veröffentlicht: (2026)
Deepfake Audio Detection Using Self-supervised Fusion Representations
von: Zaman, Khalid, et al.
Veröffentlicht: (2026)
von: Zaman, Khalid, et al.
Veröffentlicht: (2026)
WeCKD: Weakly-supervised Chained Distillation Network for Efficient Multimodal Medical Imaging
von: Rahman, Md. Abdur, et al.
Veröffentlicht: (2025)
von: Rahman, Md. Abdur, et al.
Veröffentlicht: (2025)
Self-supervised Learning for Acoustic Few-Shot Classification
von: Liang, Jingyong, et al.
Veröffentlicht: (2024)
von: Liang, Jingyong, et al.
Veröffentlicht: (2024)
Self-supervised Reflective Learning through Self-distillation and Online Clustering for Speaker Representation Learning
von: Cai, Danwei, et al.
Veröffentlicht: (2024)
von: Cai, Danwei, et al.
Veröffentlicht: (2024)
Self-supervised Multimodal Speech Representations for the Assessment of Schizophrenia Symptoms
von: Premananth, Gowtham, et al.
Veröffentlicht: (2024)
von: Premananth, Gowtham, et al.
Veröffentlicht: (2024)
Implicit Self-supervised Language Representation for Spoken Language Diarization
von: Mishra, Jagabandhu, et al.
Veröffentlicht: (2023)
von: Mishra, Jagabandhu, et al.
Veröffentlicht: (2023)
[b]=[d]-[t]+[p]: Self-supervised Speech Models Discover Phonological Vector Arithmetic
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026)
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026)
Voice Biomarker Analysis and Automated Severity Classification of Dysarthric Speech in a Multilingual Context
von: Yeo, Eunjung
Veröffentlicht: (2024)
von: Yeo, Eunjung
Veröffentlicht: (2024)
SCORE: Self-supervised Correspondence Fine-tuning for Improved Content Representations
von: Meghanani, Amit, et al.
Veröffentlicht: (2024)
von: Meghanani, Amit, et al.
Veröffentlicht: (2024)
Prosodic ABX: A Language-Agnostic Method for Measuring Prosodic Contrast in Speech Representations
von: Sun, Haitong, et al.
Veröffentlicht: (2026)
von: Sun, Haitong, et al.
Veröffentlicht: (2026)
Optimized Self-supervised Training with BEST-RQ for Speech Recognition
von: Baumann, Ilja, et al.
Veröffentlicht: (2025)
von: Baumann, Ilja, et al.
Veröffentlicht: (2025)
Self-supervised Speech Representations Still Struggle with African American Vernacular English
von: Chang, Kalvin, et al.
Veröffentlicht: (2024)
von: Chang, Kalvin, et al.
Veröffentlicht: (2024)
SSHR: Leveraging Self-supervised Hierarchical Representations for Multilingual Automatic Speech Recognition
von: Xue, Hongfei, et al.
Veröffentlicht: (2023)
von: Xue, Hongfei, et al.
Veröffentlicht: (2023)
SS-DPPN: A self-supervised dual-path foundation model for the generalizable cardiac audio representation
von: Muna, Ummy Maria, et al.
Veröffentlicht: (2025)
von: Muna, Ummy Maria, et al.
Veröffentlicht: (2025)
Causal Speech Enhancement with Predicting Semantics based on Quantized Self-supervised Learning Features
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2024)
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2024)
Learning Domain-Robust Bioacoustic Representations for Mosquito Species Classification with Contrastive Learning and Distribution Alignment
von: Hou, Yuanbo, et al.
Veröffentlicht: (2025)
von: Hou, Yuanbo, et al.
Veröffentlicht: (2025)
Causal Self-supervised Pretrained Frontend with Predictive Code for Speech Separation
von: Wang, Wupeng, et al.
Veröffentlicht: (2025)
von: Wang, Wupeng, et al.
Veröffentlicht: (2025)
SoundPlot: An Open-Source Framework for Birdsong Acoustic Analysis and Neural Synthesis with Interactive 3D Visualization
von: Mehdi, Naqcho Ali, et al.
Veröffentlicht: (2026)
von: Mehdi, Naqcho Ali, et al.
Veröffentlicht: (2026)
STONE: Self-supervised Tonality Estimator
von: Kong, Yuexuan, et al.
Veröffentlicht: (2024)
von: Kong, Yuexuan, et al.
Veröffentlicht: (2024)
LASER: Learning by Aligning Self-supervised Representations of Speech for Improving Content-related Tasks
von: Meghanani, Amit, et al.
Veröffentlicht: (2024)
von: Meghanani, Amit, et al.
Veröffentlicht: (2024)
Revisiting Self-supervised Learning of Speech Representation from a Mutual Information Perspective
von: Liu, Alexander H., et al.
Veröffentlicht: (2024)
von: Liu, Alexander H., et al.
Veröffentlicht: (2024)
Position-invariant Fine-tuning of Speech Enhancement Models with Self-supervised Speech Representations
von: Meghanani, Amit, et al.
Veröffentlicht: (2026)
von: Meghanani, Amit, et al.
Veröffentlicht: (2026)
Improving Acoustic Word Embeddings through Correspondence Training of Self-supervised Speech Representations
von: Meghanani, Amit, et al.
Veröffentlicht: (2024)
von: Meghanani, Amit, et al.
Veröffentlicht: (2024)
Less Forgetting for Better Generalization: Exploring Continual-learning Fine-tuning Methods for Speech Self-supervised Representations
von: Zaiem, Salah, et al.
Veröffentlicht: (2024)
von: Zaiem, Salah, et al.
Veröffentlicht: (2024)
ZMM-TTS: Zero-shot Multilingual and Multispeaker Speech Synthesis Conditioned on Self-supervised Discrete Speech Representations
von: Gong, Cheng, et al.
Veröffentlicht: (2023)
von: Gong, Cheng, et al.
Veröffentlicht: (2023)
Selection of Layers from Self-supervised Learning Models for Predicting Mean-Opinion-Score of Speech
von: Liang, Xinyu, et al.
Veröffentlicht: (2025)
von: Liang, Xinyu, et al.
Veröffentlicht: (2025)
The Effect of Batch Size on Contrastive Self-Supervised Speech Representation Learning
von: Vaessen, Nik, et al.
Veröffentlicht: (2024)
von: Vaessen, Nik, et al.
Veröffentlicht: (2024)
Multi-Class-Token Transformer for Multitask Self-supervised Music Information Retrieval
von: Kong, Yuexuan, et al.
Veröffentlicht: (2025)
von: Kong, Yuexuan, et al.
Veröffentlicht: (2025)
MMM: Multi-Layer Multi-Residual Multi-Stream Discrete Speech Representation from Self-supervised Learning Model
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
Boosting Multi-Speaker Expressive Speech Synthesis with Semi-supervised Contrastive Learning
von: Zhu, Xinfa, et al.
Veröffentlicht: (2023)
von: Zhu, Xinfa, et al.
Veröffentlicht: (2023)
oboVox Far Field Speaker Recognition: A Novel Data Augmentation Approach with Pretrained Models
von: Dip, Muhammad Sudipto Siam, et al.
Veröffentlicht: (2024)
von: Dip, Muhammad Sudipto Siam, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
An Innovative Coverage Path Planning Approach for UAVs to Boost Precision Agriculture and Rescue Operations
von: Nur Mohammad Fahad, et al.
Veröffentlicht: (2025) -
BioAutoML-NAS: An End-to-End AutoML Framework for Multimodal Insect Classification via Neural Architecture Search on Large-Scale Biodiversity Data
von: Abian, Arefin Ittesafun, et al.
Veröffentlicht: (2025) -
DeepAgent: A Dual Stream Multi Agent Fusion for Robust Multimodal Deepfake Detection
von: Zaman, Sayeem Been, et al.
Veröffentlicht: (2025) -
Predicting Postresection Colorectal Liver Metastases Recurrence Using Advanced Graph Neural Networks with Explainability and Causal Inference
von: Jubair Ahmed, et al.
Veröffentlicht: (2025) -
Learning to Weigh Waste: A Physics-Informed Multimodal Fusion Framework and Large-Scale Dataset for Commercial and Industrial Applications
von: Islam, Md. Adnanul, et al.
Veröffentlicht: (2026)