Model as Loss: A Self-Consistent Training Paradigm
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Phaye, Saisamarth Rajesh, Cernak, Milos, Harper, Andrew |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DeepFilterGAN: A Full-band Real-time Speech Enhancement System with GAN-based Stochastic Regeneration
von: Serbest, Sanberk, et al.
Veröffentlicht: (2025)
von: Serbest, Sanberk, et al.
Veröffentlicht: (2025)
Differentiable Time-Varying IIR Filtering for Real-Time Speech Denoising
von: Rota, Riccardo, et al.
Veröffentlicht: (2026)
von: Rota, Riccardo, et al.
Veröffentlicht: (2026)
Real-time Timbre Remapping with Differentiable DSP
von: Shier, Jordie, et al.
Veröffentlicht: (2024)
von: Shier, Jordie, et al.
Veröffentlicht: (2024)
Lightweight Self-Supervised Detection of Fundamental Frequency and Accurate Probability of Voicing in Monophonic Music
von: Bitra, Venkat Suprabath, et al.
Veröffentlicht: (2026)
von: Bitra, Venkat Suprabath, et al.
Veröffentlicht: (2026)
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training
von: Wang, Yuancheng, et al.
Veröffentlicht: (2025)
von: Wang, Yuancheng, et al.
Veröffentlicht: (2025)
T-FOLEY: A Controllable Waveform-Domain Diffusion Model for Temporal-Event-Guided Foley Sound Synthesis
von: Chung, Yoonjin, et al.
Veröffentlicht: (2024)
von: Chung, Yoonjin, et al.
Veröffentlicht: (2024)
An Explainable Proxy Model for Multiabel Audio Segmentation
von: Mariotte, Théo, et al.
Veröffentlicht: (2024)
von: Mariotte, Théo, et al.
Veröffentlicht: (2024)
MaskSR: Masked Language Model for Full-band Speech Restoration
von: Li, Xu, et al.
Veröffentlicht: (2024)
von: Li, Xu, et al.
Veröffentlicht: (2024)
TinyChirp: Bird Song Recognition Using TinyML Models on Low-power Wireless Acoustic Sensors
von: Huang, Zhaolan, et al.
Veröffentlicht: (2024)
von: Huang, Zhaolan, et al.
Veröffentlicht: (2024)
Gull: A Generative Multifunctional Audio Codec
von: Luo, Yi, et al.
Veröffentlicht: (2024)
von: Luo, Yi, et al.
Veröffentlicht: (2024)
BUET Multi-disease Heart Sound Dataset: A Comprehensive Auscultation Dataset for Developing Computer-Aided Diagnostic Systems
von: Ali, Shams Nafisa, et al.
Veröffentlicht: (2024)
von: Ali, Shams Nafisa, et al.
Veröffentlicht: (2024)
Multi-Channel MOSRA: Mean Opinion Score and Room Acoustics Estimation Using Simulated Data and a Teacher Model
von: Coldenhoff, Jozef, et al.
Veröffentlicht: (2023)
von: Coldenhoff, Jozef, et al.
Veröffentlicht: (2023)
Deep Active Speech Cancellation with Mamba-Masking Network
von: Mishaly, Yehuda, et al.
Veröffentlicht: (2025)
von: Mishaly, Yehuda, et al.
Veröffentlicht: (2025)
ANIRA: An Architecture for Neural Network Inference in Real-Time Audio Applications
von: Ackva, Valentin, et al.
Veröffentlicht: (2025)
von: Ackva, Valentin, et al.
Veröffentlicht: (2025)
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning
von: Tuncay, Ludovic, et al.
Veröffentlicht: (2025)
von: Tuncay, Ludovic, et al.
Veröffentlicht: (2025)
Gaussian Process Regression of Steering Vectors With Physics-Aware Deep Composite Kernels for Augmented Listening
von: Di Carlo, Diego, et al.
Veröffentlicht: (2025)
von: Di Carlo, Diego, et al.
Veröffentlicht: (2025)
Pitch-Conditioned Instrument Sound Synthesis From an Interactive Timbre Latent Space
von: Limberg, Christian, et al.
Veröffentlicht: (2025)
von: Limberg, Christian, et al.
Veröffentlicht: (2025)
Can Layer-wise SSL Features Improve Zero-Shot ASR Performance for Children's Speech?
von: Sinha, Abhijit, et al.
Veröffentlicht: (2025)
von: Sinha, Abhijit, et al.
Veröffentlicht: (2025)
AudioFuse: Unified Spectral-Temporal Learning via a Hybrid ViT-1D CNN Architecture for Robust Phonocardiogram Classification
von: Siddiqui, Md. Saiful Bari, et al.
Veröffentlicht: (2025)
von: Siddiqui, Md. Saiful Bari, et al.
Veröffentlicht: (2025)
When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds
von: Kang, Minsu, et al.
Veröffentlicht: (2025)
von: Kang, Minsu, et al.
Veröffentlicht: (2025)
PeriodWave: Multi-Period Flow Matching for High-Fidelity Waveform Generation
von: Lee, Sang-Hoon, et al.
Veröffentlicht: (2024)
von: Lee, Sang-Hoon, et al.
Veröffentlicht: (2024)
Scaling Transformers for Low-Bitrate High-Quality Speech Coding
von: Parker, Julian D, et al.
Veröffentlicht: (2024)
von: Parker, Julian D, et al.
Veröffentlicht: (2024)
LMAC-TD: Producing Time Domain Explanations for Audio Classifiers
von: Mancini, Eleonora, et al.
Veröffentlicht: (2024)
von: Mancini, Eleonora, et al.
Veröffentlicht: (2024)
Design Of Rubble Analyzer Probe Using ML For Earthquake
von: Sebastian, Abhishek, et al.
Veröffentlicht: (2023)
von: Sebastian, Abhishek, et al.
Veröffentlicht: (2023)
Expressive Acoustic Guitar Sound Synthesis with an Instrument-Specific Input Representation and Diffusion Outpainting
von: Kim, Hounsu, et al.
Veröffentlicht: (2024)
von: Kim, Hounsu, et al.
Veröffentlicht: (2024)
ViolinDiff: Enhancing Expressive Violin Synthesis with Pitch Bend Conditioning
von: Kim, Daewoong, et al.
Veröffentlicht: (2024)
von: Kim, Daewoong, et al.
Veröffentlicht: (2024)
Accelerating High-Fidelity Waveform Generation via Adversarial Flow Matching Optimization
von: Lee, Sang-Hoon, et al.
Veröffentlicht: (2024)
von: Lee, Sang-Hoon, et al.
Veröffentlicht: (2024)
Mismatch-Robust Underwater Acoustic Localization Using A Differentiable Modular Forward Model
von: Kari, Dariush, et al.
Veröffentlicht: (2025)
von: Kari, Dariush, et al.
Veröffentlicht: (2025)
Semi-intrusive audio evaluation: Casting non-intrusive assessment as a multi-modal text prediction task
von: Coldenhoff, Jozef, et al.
Veröffentlicht: (2024)
von: Coldenhoff, Jozef, et al.
Veröffentlicht: (2024)
Joint Source-Environment Adaptation of Data-Driven Underwater Acoustic Source Ranging Based on Model Uncertainty
von: Kari, Dariush, et al.
Veröffentlicht: (2025)
von: Kari, Dariush, et al.
Veröffentlicht: (2025)
Self-Tuning Spectral Clustering for Speaker Diarization
von: Raghav, Nikhil, et al.
Veröffentlicht: (2024)
von: Raghav, Nikhil, et al.
Veröffentlicht: (2024)
Diffusion-based Speech Enhancement with Schrödinger Bridge and Symmetric Noise Schedule
von: Wang, Siyi, et al.
Veröffentlicht: (2024)
von: Wang, Siyi, et al.
Veröffentlicht: (2024)
A Hybrid Model for Weakly-Supervised Speech Dereverberation
von: Bahrman, Louis, et al.
Veröffentlicht: (2025)
von: Bahrman, Louis, et al.
Veröffentlicht: (2025)
SWIM: Short-Window CNN Integrated with Mamba for EEG-Based Auditory Spatial Attention Decoding
von: Zhang, Ziyang, et al.
Veröffentlicht: (2024)
von: Zhang, Ziyang, et al.
Veröffentlicht: (2024)
Joint Source-Environment Adaptation for Deep Learning-Based Underwater Acoustic Source Ranging
von: Kari, Dariush, et al.
Veröffentlicht: (2025)
von: Kari, Dariush, et al.
Veröffentlicht: (2025)
Speech Self-Supervised Representations Benchmarking: a Case for Larger Probing Heads
von: Zaiem, Salah, et al.
Veröffentlicht: (2023)
von: Zaiem, Salah, et al.
Veröffentlicht: (2023)
CochCeps-Augment: A Novel Self-Supervised Contrastive Learning Using Cochlear Cepstrum-based Masking for Speech Emotion Recognition
von: Ziogas, Ioannis, et al.
Veröffentlicht: (2024)
von: Ziogas, Ioannis, et al.
Veröffentlicht: (2024)
EEG-Based Speech Decoding: A Novel Approach Using Multi-Kernel Ensemble Diffusion Models
von: Kim, Soowon, et al.
Veröffentlicht: (2024)
von: Kim, Soowon, et al.
Veröffentlicht: (2024)
A Domain-Knowledge-Inspired Music Embedding Space and a Novel Attention Mechanism for Symbolic Music Modeling
von: Guo, Z., et al.
Veröffentlicht: (2022)
von: Guo, Z., et al.
Veröffentlicht: (2022)
Wavelet GPT: Wavelet Inspired Large Language Models
von: Verma, Prateek
Veröffentlicht: (2024)
von: Verma, Prateek
Veröffentlicht: (2024)
Ähnliche Einträge
-
DeepFilterGAN: A Full-band Real-time Speech Enhancement System with GAN-based Stochastic Regeneration
von: Serbest, Sanberk, et al.
Veröffentlicht: (2025) -
Differentiable Time-Varying IIR Filtering for Real-Time Speech Denoising
von: Rota, Riccardo, et al.
Veröffentlicht: (2026) -
Real-time Timbre Remapping with Differentiable DSP
von: Shier, Jordie, et al.
Veröffentlicht: (2024) -
Lightweight Self-Supervised Detection of Fundamental Frequency and Accurate Probability of Voicing in Monophonic Music
von: Bitra, Venkat Suprabath, et al.
Veröffentlicht: (2026) -
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training
von: Wang, Yuancheng, et al.
Veröffentlicht: (2025)