A Novel Score-CAM based Denoiser for Spectrographic Signature Extraction without Ground Truth
Fuente:
arXiv
Guardado en:
| Autor principal: | Elias, Noel |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Audio Classification of Low Feature Spectrograms Utilizing Convolutional Neural Networks
por: Elias, Noel
Publicado: (2024)
por: Elias, Noel
Publicado: (2024)
HRTF Estimation using a Score-based Prior
por: Thuillier, Etienne, et al.
Publicado: (2024)
por: Thuillier, Etienne, et al.
Publicado: (2024)
Automatic Contextual Audio Denoising
por: Luong, Diep, et al.
Publicado: (2026)
por: Luong, Diep, et al.
Publicado: (2026)
Denoising by neural network for muzzle blast detection
por: Pujol, Hadrien, et al.
Publicado: (2025)
por: Pujol, Hadrien, et al.
Publicado: (2025)
Are Deep Speech Denoising Models Robust to Adversarial Noise?
por: Schwarzer, Will, et al.
Publicado: (2025)
por: Schwarzer, Will, et al.
Publicado: (2025)
SoundCTM: Unifying Score-based and Consistency Models for Full-band Text-to-Sound Generation
por: Saito, Koichi, et al.
Publicado: (2024)
por: Saito, Koichi, et al.
Publicado: (2024)
Knowledge Distillation for Speech Denoising by Latent Representation Alignment with Cosine Distance
por: Luong, Diep, et al.
Publicado: (2025)
por: Luong, Diep, et al.
Publicado: (2025)
Contextual Speech Extraction: Leveraging Textual History as an Implicit Cue for Target Speech Extraction
por: Kim, Minsu, et al.
Publicado: (2025)
por: Kim, Minsu, et al.
Publicado: (2025)
Neural Speech Extraction with Human Feedback
por: Itani, Malek, et al.
Publicado: (2025)
por: Itani, Malek, et al.
Publicado: (2025)
Uncertainty-Aware Mean Opinion Score Prediction
por: Wang, Hui, et al.
Publicado: (2024)
por: Wang, Hui, et al.
Publicado: (2024)
Cosine Scoring with Uncertainty for Neural Speaker Embedding
por: Wang, Qiongqiong, et al.
Publicado: (2024)
por: Wang, Qiongqiong, et al.
Publicado: (2024)
PBSCR: The Piano Bootleg Score Composer Recognition Dataset
por: Jain, Arhan, et al.
Publicado: (2024)
por: Jain, Arhan, et al.
Publicado: (2024)
Score-Based Training for Energy-Based TTS Models
por: Sun, Wanli, et al.
Publicado: (2025)
por: Sun, Wanli, et al.
Publicado: (2025)
Adversarial Training of Denoising Diffusion Model Using Dual Discriminators for High-Fidelity Multi-Speaker TTS
por: Ko, Myeongjin, et al.
Publicado: (2023)
por: Ko, Myeongjin, et al.
Publicado: (2023)
Identifying birdsong syllables without labelled data
por: Teng, Mélisande, et al.
Publicado: (2025)
por: Teng, Mélisande, et al.
Publicado: (2025)
End-to-end Piano Performance-MIDI to Score Conversion with Transformers
por: Beyer, Tim, et al.
Publicado: (2024)
por: Beyer, Tim, et al.
Publicado: (2024)
FlowTSE: Target Speaker Extraction with Flow Matching
por: Navon, Aviv, et al.
Publicado: (2025)
por: Navon, Aviv, et al.
Publicado: (2025)
DDTSE: Discriminative Diffusion Model for Target Speech Extraction
por: Zhang, Leying, et al.
Publicado: (2023)
por: Zhang, Leying, et al.
Publicado: (2023)
Beat this! Accurate beat tracking without DBN postprocessing
por: Foscarin, Francesco, et al.
Publicado: (2024)
por: Foscarin, Francesco, et al.
Publicado: (2024)
A Novel Stochastic Transformer-based Approach for Post-Traumatic Stress Disorder Detection using Audio Recording of Clinical Interviews
por: Dia, Mamadou, et al.
Publicado: (2024)
por: Dia, Mamadou, et al.
Publicado: (2024)
ASTRA: Aligning Speech and Text Representations for Asr without Sampling
por: Gaur, Neeraj, et al.
Publicado: (2024)
por: Gaur, Neeraj, et al.
Publicado: (2024)
Scoring Time Intervals using Non-Hierarchical Transformer For Automatic Piano Transcription
por: Yan, Yujia, et al.
Publicado: (2024)
por: Yan, Yujia, et al.
Publicado: (2024)
TSELM: Target Speaker Extraction using Discrete Tokens and Language Models
por: Tang, Beilong, et al.
Publicado: (2024)
por: Tang, Beilong, et al.
Publicado: (2024)
SoundSculpt: Direction and Semantics Driven Ambisonic Target Sound Extraction
por: Chen, Tuochao, et al.
Publicado: (2025)
por: Chen, Tuochao, et al.
Publicado: (2025)
TeLeS: Temporal Lexeme Similarity Score to Estimate Confidence in End-to-End ASR
por: Ravi, Nagarathna, et al.
Publicado: (2024)
por: Ravi, Nagarathna, et al.
Publicado: (2024)
SeMaScore : a new evaluation metric for automatic speech recognition tasks
por: Sasindran, Zitha, et al.
Publicado: (2024)
por: Sasindran, Zitha, et al.
Publicado: (2024)
An Effective Strategy for Modeling Score Ordinality and Non-uniform Intervals in Automated Speaking Assessment
por: Lo, Tien-Hong, et al.
Publicado: (2025)
por: Lo, Tien-Hong, et al.
Publicado: (2025)
Score-informed Music Source Separation: Improving Synthetic-to-real Generalization in Classical Music
por: Tunturi, Eetu, et al.
Publicado: (2025)
por: Tunturi, Eetu, et al.
Publicado: (2025)
MGE-LDM: Joint Latent Diffusion for Simultaneous Music Generation and Source Extraction
por: Chae, Yunkee, et al.
Publicado: (2025)
por: Chae, Yunkee, et al.
Publicado: (2025)
Frame-Level Internal Tool Use for Temporal Grounding in Audio LMs
por: An, Joesph, et al.
Publicado: (2026)
por: An, Joesph, et al.
Publicado: (2026)
Adaptive Rotary Steering with Joint Autoregression for Robust Extraction of Closely Moving Speakers in Dynamic Scenarios
por: Kienegger, Jakob, et al.
Publicado: (2026)
por: Kienegger, Jakob, et al.
Publicado: (2026)
Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
por: Kienegger, Jakob, et al.
Publicado: (2026)
por: Kienegger, Jakob, et al.
Publicado: (2026)
Steering Deep Non-Linear Spatially Selective Filters for Weakly Guided Extraction of Moving Speakers in Dynamic Scenarios
por: Kienegger, Jakob, et al.
Publicado: (2025)
por: Kienegger, Jakob, et al.
Publicado: (2025)
Multi-Representation Attention Framework for Underwater Bioacoustic Denoising and Recognition
por: Razig, Amine, et al.
Publicado: (2025)
por: Razig, Amine, et al.
Publicado: (2025)
Self-Steering Deep Non-Linear Spatially Selective Filters for Efficient Extraction of Moving Speakers under Weak Guidance
por: Kienegger, Jakob, et al.
Publicado: (2025)
por: Kienegger, Jakob, et al.
Publicado: (2025)
Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities
por: Kong, Zhifeng, et al.
Publicado: (2024)
por: Kong, Zhifeng, et al.
Publicado: (2024)
Unrolled Creative Adversarial Network For Generating Novel Musical Pieces
por: Nag, Pratik
Publicado: (2024)
por: Nag, Pratik
Publicado: (2024)
Bringing the Discussion of Minima Sharpness to the Audio Domain: a Filter-Normalised Evaluation for Acoustic Scene Classification
por: Milling, Manuel, et al.
Publicado: (2023)
por: Milling, Manuel, et al.
Publicado: (2023)
Improving BERT for Symbolic Music Understanding Using Token Denoising and Pianoroll Prediction
por: Wang, Jun-You, et al.
Publicado: (2025)
por: Wang, Jun-You, et al.
Publicado: (2025)
Biodenoising: Animal Vocalization Denoising without Access to Clean Data
por: Miron, Marius, et al.
Publicado: (2024)
por: Miron, Marius, et al.
Publicado: (2024)
Ejemplares similares
-
Audio Classification of Low Feature Spectrograms Utilizing Convolutional Neural Networks
por: Elias, Noel
Publicado: (2024) -
HRTF Estimation using a Score-based Prior
por: Thuillier, Etienne, et al.
Publicado: (2024) -
Automatic Contextual Audio Denoising
por: Luong, Diep, et al.
Publicado: (2026) -
Denoising by neural network for muzzle blast detection
por: Pujol, Hadrien, et al.
Publicado: (2025) -
Are Deep Speech Denoising Models Robust to Adversarial Noise?
por: Schwarzer, Will, et al.
Publicado: (2025)