Preserving Russek's "Summermood" Using Reality Check and a DeltaLab DL-4 Approximation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hyrkas, Jeremy, Carrillo, Pablo Dodero, Sánchez, Teresa Díaz de Cossio |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Network Modulation Synthesis: New Algorithms for Generating Musical Audio Using Autoencoder Networks
von: Hyrkas, Jeremy
Veröffentlicht: (2021)
von: Hyrkas, Jeremy
Veröffentlicht: (2021)
Real-time implementation of vibrato transfer as an audio effect
von: Hyrkas, Jeremy
Veröffentlicht: (2025)
von: Hyrkas, Jeremy
Veröffentlicht: (2025)
823-OLT @ BUET DL Sprint 4.0: Context-Aware Windowing for ASR and Fine-Tuned Speaker Diarization in Bengali Long Form Audio
von: Dhar, Ratnajit, et al.
Veröffentlicht: (2026)
von: Dhar, Ratnajit, et al.
Veröffentlicht: (2026)
Towards Privacy-Preserving Audio Classification Systems
von: Chhaglani, Bhawana, et al.
Veröffentlicht: (2024)
von: Chhaglani, Bhawana, et al.
Veröffentlicht: (2024)
UTDUSS: UTokyo-SaruLab System for Interspeech2024 Speech Processing Using Discrete Speech Unit Challenge
von: Nakata, Wataru, et al.
Veröffentlicht: (2024)
von: Nakata, Wataru, et al.
Veröffentlicht: (2024)
Balancing Information Preservation and Disentanglement in Self-Supervised Music Representation Learning
von: Wilkins, Julia, et al.
Veröffentlicht: (2025)
von: Wilkins, Julia, et al.
Veröffentlicht: (2025)
Mind the Shift: Using Delta SSL Embeddings to Enhance Child ASR
von: Wang, Zilai, et al.
Veröffentlicht: (2026)
von: Wang, Zilai, et al.
Veröffentlicht: (2026)
The SJTU X-LANCE Lab System for MSR Challenge 2025
von: Zhu, Jinxuan, et al.
Veröffentlicht: (2026)
von: Zhu, Jinxuan, et al.
Veröffentlicht: (2026)
Emotion-Aware Quantization for Discrete Speech Representations: An Analysis of Emotion Preservation
von: Zhou, Haoguang, et al.
Veröffentlicht: (2026)
von: Zhou, Haoguang, et al.
Veröffentlicht: (2026)
AVENet: Disentangling Features by Approximating Average Features for Voice Conversion
von: Wang, Wenyu, et al.
Veröffentlicht: (2025)
von: Wang, Wenyu, et al.
Veröffentlicht: (2025)
AffectCodec: Emotion-Preserving Neural Speech Codec for Expressive Speech Modeling
von: Shi, Jiacheng, et al.
Veröffentlicht: (2026)
von: Shi, Jiacheng, et al.
Veröffentlicht: (2026)
AffectCodec: Emotion-Preserving Neural Speech Codec with Block-Diagonal Residual FSQ
von: Meng, Zhaoyang, et al.
Veröffentlicht: (2026)
von: Meng, Zhaoyang, et al.
Veröffentlicht: (2026)
Phonetic Segmentation of the UCLA Phonetics Lab Archive
von: Chodroff, Eleanor, et al.
Veröffentlicht: (2024)
von: Chodroff, Eleanor, et al.
Veröffentlicht: (2024)
The Affective Bridge: Preserving Speech Representations while Enhancing Deepfake Detection vian emotional Constraints
von: Li, Yupei, et al.
Veröffentlicht: (2025)
von: Li, Yupei, et al.
Veröffentlicht: (2025)
SonoTraceLab -- A Raytracing-Based Acoustic Modelling System for Simulating Echolocation Behavior of Bats
von: Jansen, Wouter, et al.
Veröffentlicht: (2024)
von: Jansen, Wouter, et al.
Veröffentlicht: (2024)
Do We Need EMA for Diffusion-Based Speech Enhancement? Toward a Magnitude-Preserving Network Architecture
von: Richter, Julius, et al.
Veröffentlicht: (2025)
von: Richter, Julius, et al.
Veröffentlicht: (2025)
Sound Check: Auditing Audio Datasets
von: Agnew, William, et al.
Veröffentlicht: (2024)
von: Agnew, William, et al.
Veröffentlicht: (2024)
Quality of Automatic Speech Recognition -- Polish Language case study -- from Wav2Vec to Scribe ElevenLabs
von: Pietroń, Marcin, et al.
Veröffentlicht: (2026)
von: Pietroń, Marcin, et al.
Veröffentlicht: (2026)
AR&D: A Framework for Retrieving and Describing Concepts for Interpreting AudioLLMs
von: Chowdhury, Townim Faisal, et al.
Veröffentlicht: (2026)
von: Chowdhury, Townim Faisal, et al.
Veröffentlicht: (2026)
Padé Approximant Neural Networks for Enhanced Electric Motor Fault Diagnosis Using Vibration and Acoustic Data
von: Kilickaya, Sertac, et al.
Veröffentlicht: (2025)
von: Kilickaya, Sertac, et al.
Veröffentlicht: (2025)
Exploring the Feasibility of LLMs for Automated Music Emotion Annotation
von: Yang, Meng, et al.
Veröffentlicht: (2025)
von: Yang, Meng, et al.
Veröffentlicht: (2025)
AnchorSteer: Self-Discovered Concept Injection for Structure-Preserving Music Editing
von: Chang, Chih-Heng, et al.
Veröffentlicht: (2026)
von: Chang, Chih-Heng, et al.
Veröffentlicht: (2026)
Acoustic evaluation of a neural network dedicated to the detection of animal vocalisations
von: Rouch, Jérémy, et al.
Veröffentlicht: (2025)
von: Rouch, Jérémy, et al.
Veröffentlicht: (2025)
ClearMask: Noise-Free and Naturalness-Preserving Protection Against Voice Deepfake Attacks
von: Wang, Yuanda, et al.
Veröffentlicht: (2025)
von: Wang, Yuanda, et al.
Veröffentlicht: (2025)
MuseCPBench: an Empirical Study of Music Editing Methods through Music Context Preservation
von: Vishe, Yash, et al.
Veröffentlicht: (2025)
von: Vishe, Yash, et al.
Veröffentlicht: (2025)
EchoVoices: Preserving Generational Voices and Memories for Seniors and Children
von: Xu, Haiying, et al.
Veröffentlicht: (2025)
von: Xu, Haiying, et al.
Veröffentlicht: (2025)
Universal Score-based Speech Enhancement with High Content Preservation
von: Scheibler, Robin, et al.
Veröffentlicht: (2024)
von: Scheibler, Robin, et al.
Veröffentlicht: (2024)
Latent Multi-view Learning for Robust Environmental Sound Representations
von: Ding, Sivan, et al.
Veröffentlicht: (2025)
von: Ding, Sivan, et al.
Veröffentlicht: (2025)
Robust Training of Singing Voice Synthesis Using Prior and Posterior Uncertainty
von: Zhao, Yiwen, et al.
Veröffentlicht: (2025)
von: Zhao, Yiwen, et al.
Veröffentlicht: (2025)
Adapting General Disentanglement-Based Speaker Anonymization for Enhanced Emotion Preservation
von: Miao, Xiaoxiao, et al.
Veröffentlicht: (2024)
von: Miao, Xiaoxiao, et al.
Veröffentlicht: (2024)
An Explicit Consistency-Preserving Loss Function for Phase Reconstruction and Speech Enhancement
von: Ku, Pin-Jui, et al.
Veröffentlicht: (2024)
von: Ku, Pin-Jui, et al.
Veröffentlicht: (2024)
Facilitating Personalized TTS for Dysarthric Speakers Using Knowledge Anchoring and Curriculum Learning
von: Jeon, Yejin, et al.
Veröffentlicht: (2025)
von: Jeon, Yejin, et al.
Veröffentlicht: (2025)
Enhancing Self-Supervised Speaker Verification Using Similarity-Connected Graphs and GCN
von: Sun, Zhaorui, et al.
Veröffentlicht: (2025)
von: Sun, Zhaorui, et al.
Veröffentlicht: (2025)
Adversarial Domain Adaptation for Metal Cutting Sound Detection: Leveraging Abundant Lab Data for Scarce Industry Data
von: Mostafiz, Mir Imtiaz, et al.
Veröffentlicht: (2024)
von: Mostafiz, Mir Imtiaz, et al.
Veröffentlicht: (2024)
Lend me an Ear: Speech Enhancement Using a Robotic Arm with a Microphone Array
von: Turcotte, Zachary, et al.
Veröffentlicht: (2026)
von: Turcotte, Zachary, et al.
Veröffentlicht: (2026)
End-to-End User-Defined Keyword Spotting using Shifted Delta Coefficients
von: V, Kesavaraj, et al.
Veröffentlicht: (2024)
von: V, Kesavaraj, et al.
Veröffentlicht: (2024)
Vector Signal Reconstruction Sparse and Parametric Approach of direction of arrival Using Single Vector Hydrophone
von: Guo, Jiabin
Veröffentlicht: (2024)
von: Guo, Jiabin
Veröffentlicht: (2024)
Oral Tradition-Encoded NanyinHGNN: Integrating Nanyin Music Preservation and Generation through a Pipa-Centric Dataset
von: Xiahou, Jianbing, et al.
Veröffentlicht: (2025)
von: Xiahou, Jianbing, et al.
Veröffentlicht: (2025)
Improving DF-Conformer Using Hydra For High-Fidelity Generative Speech Enhancement on Discrete Codec Token
von: Seki, Shogo, et al.
Veröffentlicht: (2025)
von: Seki, Shogo, et al.
Veröffentlicht: (2025)
A Lightweight Fourier-based Network for Binaural Speech Enhancement with Spatial Cue Preservation
von: Lu, Xikun, et al.
Veröffentlicht: (2025)
von: Lu, Xikun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Network Modulation Synthesis: New Algorithms for Generating Musical Audio Using Autoencoder Networks
von: Hyrkas, Jeremy
Veröffentlicht: (2021) -
Real-time implementation of vibrato transfer as an audio effect
von: Hyrkas, Jeremy
Veröffentlicht: (2025) -
823-OLT @ BUET DL Sprint 4.0: Context-Aware Windowing for ASR and Fine-Tuned Speaker Diarization in Bengali Long Form Audio
von: Dhar, Ratnajit, et al.
Veröffentlicht: (2026) -
Towards Privacy-Preserving Audio Classification Systems
von: Chhaglani, Bhawana, et al.
Veröffentlicht: (2024) -
UTDUSS: UTokyo-SaruLab System for Interspeech2024 Speech Processing Using Discrete Speech Unit Challenge
von: Nakata, Wataru, et al.
Veröffentlicht: (2024)