Guardado en:
| Autores principales: | Rahman, Md. Abdur, Thuseethan, Selvarajah, Yeo, Kheng Cher, Mohamed, Reem E., Azam, Sami |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2510.00522 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
An Innovative Coverage Path Planning Approach for UAVs to Boost Precision Agriculture and Rescue Operations
por: Nur Mohammad Fahad, et al.
Publicado: (2025)
por: Nur Mohammad Fahad, et al.
Publicado: (2025)
BioAutoML-NAS: An End-to-End AutoML Framework for Multimodal Insect Classification via Neural Architecture Search on Large-Scale Biodiversity Data
por: Abian, Arefin Ittesafun, et al.
Publicado: (2025)
por: Abian, Arefin Ittesafun, et al.
Publicado: (2025)
DeepAgent: A Dual Stream Multi Agent Fusion for Robust Multimodal Deepfake Detection
por: Zaman, Sayeem Been, et al.
Publicado: (2025)
por: Zaman, Sayeem Been, et al.
Publicado: (2025)
Predicting Postresection Colorectal Liver Metastases Recurrence Using Advanced Graph Neural Networks with Explainability and Causal Inference
por: Jubair Ahmed, et al.
Publicado: (2025)
por: Jubair Ahmed, et al.
Publicado: (2025)
Learning to Weigh Waste: A Physics-Informed Multimodal Fusion Framework and Large-Scale Dataset for Commercial and Industrial Applications
por: Islam, Md. Adnanul, et al.
Publicado: (2026)
por: Islam, Md. Adnanul, et al.
Publicado: (2026)
Leveraging Self-supervised Audio Representations for Data-Efficient Acoustic Scene Classification
por: Cai, Yiqiang, et al.
Publicado: (2024)
por: Cai, Yiqiang, et al.
Publicado: (2024)
A Source-Free Approach for Domain Adaptation via Multiview Image Transformation and Latent Space Consistency
por: Sutradhar, Debopom, et al.
Publicado: (2026)
por: Sutradhar, Debopom, et al.
Publicado: (2026)
Masked Latent Prediction and Classification for Self-Supervised Audio Representation Learning
por: Quelennec, Aurian, et al.
Publicado: (2025)
por: Quelennec, Aurian, et al.
Publicado: (2025)
From Birdsong to Rumbles: Classifying Elephant Calls with Out-of-Species Embeddings
por: Geldenhuys, Christiaan M., et al.
Publicado: (2026)
por: Geldenhuys, Christiaan M., et al.
Publicado: (2026)
Deepfake Audio Detection Using Self-supervised Fusion Representations
por: Zaman, Khalid, et al.
Publicado: (2026)
por: Zaman, Khalid, et al.
Publicado: (2026)
WeCKD: Weakly-supervised Chained Distillation Network for Efficient Multimodal Medical Imaging
por: Rahman, Md. Abdur, et al.
Publicado: (2025)
por: Rahman, Md. Abdur, et al.
Publicado: (2025)
Self-supervised Learning for Acoustic Few-Shot Classification
por: Liang, Jingyong, et al.
Publicado: (2024)
por: Liang, Jingyong, et al.
Publicado: (2024)
Self-supervised Reflective Learning through Self-distillation and Online Clustering for Speaker Representation Learning
por: Cai, Danwei, et al.
Publicado: (2024)
por: Cai, Danwei, et al.
Publicado: (2024)
Self-supervised Multimodal Speech Representations for the Assessment of Schizophrenia Symptoms
por: Premananth, Gowtham, et al.
Publicado: (2024)
por: Premananth, Gowtham, et al.
Publicado: (2024)
Implicit Self-supervised Language Representation for Spoken Language Diarization
por: Mishra, Jagabandhu, et al.
Publicado: (2023)
por: Mishra, Jagabandhu, et al.
Publicado: (2023)
[b]=[d]-[t]+[p]: Self-supervised Speech Models Discover Phonological Vector Arithmetic
por: Choi, Kwanghee, et al.
Publicado: (2026)
por: Choi, Kwanghee, et al.
Publicado: (2026)
Voice Biomarker Analysis and Automated Severity Classification of Dysarthric Speech in a Multilingual Context
por: Yeo, Eunjung
Publicado: (2024)
por: Yeo, Eunjung
Publicado: (2024)
SCORE: Self-supervised Correspondence Fine-tuning for Improved Content Representations
por: Meghanani, Amit, et al.
Publicado: (2024)
por: Meghanani, Amit, et al.
Publicado: (2024)
Prosodic ABX: A Language-Agnostic Method for Measuring Prosodic Contrast in Speech Representations
por: Sun, Haitong, et al.
Publicado: (2026)
por: Sun, Haitong, et al.
Publicado: (2026)
Optimized Self-supervised Training with BEST-RQ for Speech Recognition
por: Baumann, Ilja, et al.
Publicado: (2025)
por: Baumann, Ilja, et al.
Publicado: (2025)
Self-supervised Speech Representations Still Struggle with African American Vernacular English
por: Chang, Kalvin, et al.
Publicado: (2024)
por: Chang, Kalvin, et al.
Publicado: (2024)
SSHR: Leveraging Self-supervised Hierarchical Representations for Multilingual Automatic Speech Recognition
por: Xue, Hongfei, et al.
Publicado: (2023)
por: Xue, Hongfei, et al.
Publicado: (2023)
SS-DPPN: A self-supervised dual-path foundation model for the generalizable cardiac audio representation
por: Muna, Ummy Maria, et al.
Publicado: (2025)
por: Muna, Ummy Maria, et al.
Publicado: (2025)
Causal Speech Enhancement with Predicting Semantics based on Quantized Self-supervised Learning Features
por: Tsunoo, Emiru, et al.
Publicado: (2024)
por: Tsunoo, Emiru, et al.
Publicado: (2024)
Learning Domain-Robust Bioacoustic Representations for Mosquito Species Classification with Contrastive Learning and Distribution Alignment
por: Hou, Yuanbo, et al.
Publicado: (2025)
por: Hou, Yuanbo, et al.
Publicado: (2025)
Causal Self-supervised Pretrained Frontend with Predictive Code for Speech Separation
por: Wang, Wupeng, et al.
Publicado: (2025)
por: Wang, Wupeng, et al.
Publicado: (2025)
SoundPlot: An Open-Source Framework for Birdsong Acoustic Analysis and Neural Synthesis with Interactive 3D Visualization
por: Mehdi, Naqcho Ali, et al.
Publicado: (2026)
por: Mehdi, Naqcho Ali, et al.
Publicado: (2026)
STONE: Self-supervised Tonality Estimator
por: Kong, Yuexuan, et al.
Publicado: (2024)
por: Kong, Yuexuan, et al.
Publicado: (2024)
LASER: Learning by Aligning Self-supervised Representations of Speech for Improving Content-related Tasks
por: Meghanani, Amit, et al.
Publicado: (2024)
por: Meghanani, Amit, et al.
Publicado: (2024)
Revisiting Self-supervised Learning of Speech Representation from a Mutual Information Perspective
por: Liu, Alexander H., et al.
Publicado: (2024)
por: Liu, Alexander H., et al.
Publicado: (2024)
Position-invariant Fine-tuning of Speech Enhancement Models with Self-supervised Speech Representations
por: Meghanani, Amit, et al.
Publicado: (2026)
por: Meghanani, Amit, et al.
Publicado: (2026)
Improving Acoustic Word Embeddings through Correspondence Training of Self-supervised Speech Representations
por: Meghanani, Amit, et al.
Publicado: (2024)
por: Meghanani, Amit, et al.
Publicado: (2024)
Less Forgetting for Better Generalization: Exploring Continual-learning Fine-tuning Methods for Speech Self-supervised Representations
por: Zaiem, Salah, et al.
Publicado: (2024)
por: Zaiem, Salah, et al.
Publicado: (2024)
ZMM-TTS: Zero-shot Multilingual and Multispeaker Speech Synthesis Conditioned on Self-supervised Discrete Speech Representations
por: Gong, Cheng, et al.
Publicado: (2023)
por: Gong, Cheng, et al.
Publicado: (2023)
Selection of Layers from Self-supervised Learning Models for Predicting Mean-Opinion-Score of Speech
por: Liang, Xinyu, et al.
Publicado: (2025)
por: Liang, Xinyu, et al.
Publicado: (2025)
The Effect of Batch Size on Contrastive Self-Supervised Speech Representation Learning
por: Vaessen, Nik, et al.
Publicado: (2024)
por: Vaessen, Nik, et al.
Publicado: (2024)
Multi-Class-Token Transformer for Multitask Self-supervised Music Information Retrieval
por: Kong, Yuexuan, et al.
Publicado: (2025)
por: Kong, Yuexuan, et al.
Publicado: (2025)
MMM: Multi-Layer Multi-Residual Multi-Stream Discrete Speech Representation from Self-supervised Learning Model
por: Shi, Jiatong, et al.
Publicado: (2024)
por: Shi, Jiatong, et al.
Publicado: (2024)
Boosting Multi-Speaker Expressive Speech Synthesis with Semi-supervised Contrastive Learning
por: Zhu, Xinfa, et al.
Publicado: (2023)
por: Zhu, Xinfa, et al.
Publicado: (2023)
oboVox Far Field Speaker Recognition: A Novel Data Augmentation Approach with Pretrained Models
por: Dip, Muhammad Sudipto Siam, et al.
Publicado: (2024)
por: Dip, Muhammad Sudipto Siam, et al.
Publicado: (2024)
Ejemplares similares
-
An Innovative Coverage Path Planning Approach for UAVs to Boost Precision Agriculture and Rescue Operations
por: Nur Mohammad Fahad, et al.
Publicado: (2025) -
BioAutoML-NAS: An End-to-End AutoML Framework for Multimodal Insect Classification via Neural Architecture Search on Large-Scale Biodiversity Data
por: Abian, Arefin Ittesafun, et al.
Publicado: (2025) -
DeepAgent: A Dual Stream Multi Agent Fusion for Robust Multimodal Deepfake Detection
por: Zaman, Sayeem Been, et al.
Publicado: (2025) -
Predicting Postresection Colorectal Liver Metastases Recurrence Using Advanced Graph Neural Networks with Explainability and Causal Inference
por: Jubair Ahmed, et al.
Publicado: (2025) -
Learning to Weigh Waste: A Physics-Informed Multimodal Fusion Framework and Large-Scale Dataset for Commercial and Industrial Applications
por: Islam, Md. Adnanul, et al.
Publicado: (2026)