Deep Hierarchical Knowledge Loss for Fault Intensity Diagnosis
Fuente:
arXiv
Guardado en:
| Autores principales: | Sha, Yu, Gou, Shuiping, Liu, Bo, Lu, Haofan, Liu, Ningtao, Fu, Jiahui, Stoecker, Horst, Vnucec, Domagoj, Wetzstein, Nadine, Widl, Andreas, Zhou, Kai |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A Review of Machine Learning for Cavitation Intensity Recognition in Complex Industrial Systems
por: Sha, Yu, et al.
Publicado: (2025)
por: Sha, Yu, et al.
Publicado: (2025)
FC-TTS: Style and Timbre Control in Zero-Shot Text-to-Speech with Disentangled Speech Representations
por: Lee, Yoonhyung, et al.
Publicado: (2026)
por: Lee, Yoonhyung, et al.
Publicado: (2026)
Enhancing Audio Generation Diversity with Visual Information
por: Xie, Zeyu, et al.
Publicado: (2024)
por: Xie, Zeyu, et al.
Publicado: (2024)
Deep Dubbing: End-to-End Auto-Audiobook System with Text-to-Timbre and Context-Aware Instruct-TTS
por: Dai, Ziqi, et al.
Publicado: (2025)
por: Dai, Ziqi, et al.
Publicado: (2025)
An open-source voice type classifier for child-centered daylong recordings
por: Lavechin, Marvin, et al.
Publicado: (2020)
por: Lavechin, Marvin, et al.
Publicado: (2020)
Beyond Deep Learning: Speech Segmentation and Phone Classification with Neural Assemblies
por: Adelson, Trevor, et al.
Publicado: (2026)
por: Adelson, Trevor, et al.
Publicado: (2026)
Exploring rhythm formant analysis for Indic language classification
por: Gogoi, Parismita, et al.
Publicado: (2024)
por: Gogoi, Parismita, et al.
Publicado: (2024)
SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech
por: Cheng, Zhuangfei, et al.
Publicado: (2025)
por: Cheng, Zhuangfei, et al.
Publicado: (2025)
Joint Feature and Output Distillation for Low-complexity Acoustic Scene Classification
por: Li, Haowen, et al.
Publicado: (2025)
por: Li, Haowen, et al.
Publicado: (2025)
Quantifying the effect of speech pathology on automatic and human speaker verification
por: Halpern, Bence Mark, et al.
Publicado: (2024)
por: Halpern, Bence Mark, et al.
Publicado: (2024)
Boundary Regression for Leitmotif Detection in Music Audio
por: Lee, Sihun, et al.
Publicado: (2025)
por: Lee, Sihun, et al.
Publicado: (2025)
A Voice-based Triage for Type 2 Diabetes using a Conversational Virtual Assistant in the Home Environment
por: Summoogum, Kelvin, et al.
Publicado: (2024)
por: Summoogum, Kelvin, et al.
Publicado: (2024)
Dereverberation Using Binary Residual Masking with Time-Domain Consistency
por: Williams, Daniel G.
Publicado: (2025)
por: Williams, Daniel G.
Publicado: (2025)
BAST: Binaural Audio Spectrogram Transformer for Binaural Sound Localization
por: Kuang, Sheng, et al.
Publicado: (2022)
por: Kuang, Sheng, et al.
Publicado: (2022)
Automatic Speech Recognition (ASR) for the Diagnosis of pronunciation of Speech Sound Disorders in Korean children
por: Ahn, Taekyung, et al.
Publicado: (2024)
por: Ahn, Taekyung, et al.
Publicado: (2024)
Representation Loss Minimization with Randomized Selection Strategy for Efficient Environmental Fake Audio Detection
por: Phukan, Orchid Chetia, et al.
Publicado: (2024)
por: Phukan, Orchid Chetia, et al.
Publicado: (2024)
A Robust Classification Method using Hybrid Word Embedding for Early Diagnosis of Alzheimer's Disease
por: Li, Yangyang
Publicado: (2025)
por: Li, Yangyang
Publicado: (2025)
STAR: Speech-to-Audio Generation via Representation Learning
por: Xie, Zeyu, et al.
Publicado: (2025)
por: Xie, Zeyu, et al.
Publicado: (2025)
M2D-CLAP: Exploring General-purpose Audio-Language Representations Beyond CLAP
por: Niizumi, Daisuke, et al.
Publicado: (2025)
por: Niizumi, Daisuke, et al.
Publicado: (2025)
FakeSound2: A Benchmark for Explainable and Generalizable Deepfake Sound Detection
por: Xie, Zeyu, et al.
Publicado: (2025)
por: Xie, Zeyu, et al.
Publicado: (2025)
SynParaSpeech: Automated Synthesis of Paralinguistic Datasets for Speech Generation and Understanding
por: Bai, Bingsong, et al.
Publicado: (2025)
por: Bai, Bingsong, et al.
Publicado: (2025)
PicoAudio: Enabling Precise Timestamp and Frequency Controllability of Audio Events in Text-to-audio Generation
por: Xie, Zeyu, et al.
Publicado: (2024)
por: Xie, Zeyu, et al.
Publicado: (2024)
Non-Invasive Suicide Risk Prediction Through Speech Analysis
por: Amiriparian, Shahin, et al.
Publicado: (2024)
por: Amiriparian, Shahin, et al.
Publicado: (2024)
An Empirical Analysis of Speech Self-Supervised Learning at Multiple Resolutions
por: Clark, Theo, et al.
Publicado: (2024)
por: Clark, Theo, et al.
Publicado: (2024)
AudioTime: A Temporally-aligned Audio-text Benchmark Dataset
por: Xie, Zeyu, et al.
Publicado: (2024)
por: Xie, Zeyu, et al.
Publicado: (2024)
Towards Multi-Level Transcript Segmentation: LoRA Fine-Tuning for Table-of-Contents Generation
por: Freisinger, Steffen, et al.
Publicado: (2026)
por: Freisinger, Steffen, et al.
Publicado: (2026)
Investigation into respiratory sound classification for an imbalanced data set using hybrid LSTM-KAN architectures
por: K. V, Nithinkumar, et al.
Publicado: (2026)
por: K. V, Nithinkumar, et al.
Publicado: (2026)
Whisper Speaker Identification: Leveraging Pre-Trained Multilingual Transformers for Robust Speaker Embeddings
por: Emon, Jakaria Islam, et al.
Publicado: (2025)
por: Emon, Jakaria Islam, et al.
Publicado: (2025)
CAST-TTS: A Simple Cross-Attention Framework for Unified Timbre Control in TTS
por: Zheng, Zihao, et al.
Publicado: (2026)
por: Zheng, Zihao, et al.
Publicado: (2026)
PicoAudio2: Temporal Controllable Text-to-Audio Generation with Natural Language Description
por: Zheng, Zihao, et al.
Publicado: (2025)
por: Zheng, Zihao, et al.
Publicado: (2025)
FakeSound: Deepfake General Audio Detection
por: Xie, Zeyu, et al.
Publicado: (2024)
por: Xie, Zeyu, et al.
Publicado: (2024)
U2++ MoE: Scaling 4.7x parameters with minimal impact on RTF
por: Song, Xingchen, et al.
Publicado: (2024)
por: Song, Xingchen, et al.
Publicado: (2024)
Beyond Speech and More: Investigating the Emergent Ability of Speech Foundation Models for Classifying Physiological Time-Series Signals
por: Phukan, Orchid Chetia, et al.
Publicado: (2024)
por: Phukan, Orchid Chetia, et al.
Publicado: (2024)
SigWavNet: Learning Multiresolution Signal Wavelet Network for Speech Emotion Recognition
por: Nfissi, Alaa, et al.
Publicado: (2025)
por: Nfissi, Alaa, et al.
Publicado: (2025)
Knowledge Distillation for Real-Time Classification of Early Media in Voice Communications
por: Altwlkany, Kemal, et al.
Publicado: (2024)
por: Altwlkany, Kemal, et al.
Publicado: (2024)
An Audio-centric Multi-task Learning Framework for Streaming Ads Targeting on Spotify
por: Verma, Shivam, et al.
Publicado: (2025)
por: Verma, Shivam, et al.
Publicado: (2025)
Deepfake audio as a data augmentation technique for training automatic speech to text transcription models
por: Ferreira, Alexandre R., et al.
Publicado: (2023)
por: Ferreira, Alexandre R., et al.
Publicado: (2023)
TuneGenie: Reasoning-based LLM agents for preferential music generation
por: Pandey, Amitesh, et al.
Publicado: (2025)
por: Pandey, Amitesh, et al.
Publicado: (2025)
SW-ASR: A Context-Aware Hybrid ASR Pipeline for Robust Single Word Speech Recognition
por: Sharma, Manali, et al.
Publicado: (2026)
por: Sharma, Manali, et al.
Publicado: (2026)
Investigating Prosodic Signatures via Speech Pre-Trained Models for Audio Deepfake Source Attribution
por: Phukan, Orchid Chetia, et al.
Publicado: (2024)
por: Phukan, Orchid Chetia, et al.
Publicado: (2024)
Ejemplares similares
-
A Review of Machine Learning for Cavitation Intensity Recognition in Complex Industrial Systems
por: Sha, Yu, et al.
Publicado: (2025) -
FC-TTS: Style and Timbre Control in Zero-Shot Text-to-Speech with Disentangled Speech Representations
por: Lee, Yoonhyung, et al.
Publicado: (2026) -
Enhancing Audio Generation Diversity with Visual Information
por: Xie, Zeyu, et al.
Publicado: (2024) -
Deep Dubbing: End-to-End Auto-Audiobook System with Text-to-Timbre and Context-Aware Instruct-TTS
por: Dai, Ziqi, et al.
Publicado: (2025) -
An open-source voice type classifier for child-centered daylong recordings
por: Lavechin, Marvin, et al.
Publicado: (2020)