Quality Audio Prototyping: a prototype system for unified sound retrieval and procedural generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Garcia, Nelly, Bhattacharjee, Aditya, Mason-Williams, Gabryel, Mason-Williams, Israel, Benetos, Emmanouil, Reiss, Joshua |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ST-ITO: Controlling Audio Effects for Style Transfer with Inference-Time Optimization
von: Steinmetz, Christian J., et al.
Veröffentlicht: (2024)
von: Steinmetz, Christian J., et al.
Veröffentlicht: (2024)
GraFPrint: A GNN-Based Approach for Audio Identification
von: Bhattacharjee, Aditya, et al.
Veröffentlicht: (2024)
von: Bhattacharjee, Aditya, et al.
Veröffentlicht: (2024)
Scalable Evaluation for Audio Identification via Synthetic Latent Fingerprint Generation
von: Bhattacharjee, Aditya, et al.
Veröffentlicht: (2025)
von: Bhattacharjee, Aditya, et al.
Veröffentlicht: (2025)
Learning Music Audio Representations With Limited Data
von: Plachouras, Christos, et al.
Veröffentlicht: (2025)
von: Plachouras, Christos, et al.
Veröffentlicht: (2025)
LC-Protonets: Multi-Label Few-Shot Learning for World Music Audio Tagging
von: Papaioannou, Charilaos, et al.
Veröffentlicht: (2024)
von: Papaioannou, Charilaos, et al.
Veröffentlicht: (2024)
LHGNN: Local-Higher Order Graph Neural Networks For Audio Classification and Tagging
von: Singh, Shubhr, et al.
Veröffentlicht: (2025)
von: Singh, Shubhr, et al.
Veröffentlicht: (2025)
Towards Building an End-to-End Multilingual Automatic Lyrics Transcription Model
von: Huang, Jiawen, et al.
Veröffentlicht: (2024)
von: Huang, Jiawen, et al.
Veröffentlicht: (2024)
Human Perception of Audio Deepfakes
von: Müller, Nicolas M., et al.
Veröffentlicht: (2021)
von: Müller, Nicolas M., et al.
Veröffentlicht: (2021)
Classification of Spontaneous and Scripted Speech for Multilingual Audio
von: Elisha, Shahar, et al.
Veröffentlicht: (2024)
von: Elisha, Shahar, et al.
Veröffentlicht: (2024)
Domain-Invariant Representation Learning of Bird Sounds
von: Moummad, Ilyass, et al.
Veröffentlicht: (2024)
von: Moummad, Ilyass, et al.
Veröffentlicht: (2024)
RUMAA: Repeat-Aware Unified Music Audio Analysis for Score-Performance Alignment, Transcription, and Mistake Detection
von: Chang, Sungkyun, et al.
Veröffentlicht: (2025)
von: Chang, Sungkyun, et al.
Veröffentlicht: (2025)
Exploring Perceptual Audio Quality Measurement on Stereo Processing Using the Open Dataset of Audio Quality
von: Delgado, Pablo M., et al.
Veröffentlicht: (2025)
von: Delgado, Pablo M., et al.
Veröffentlicht: (2025)
ExSampling: a system for the real-time ensemble performance of field-recorded environmental sounds
von: Kobayashi, Atsuya, et al.
Veröffentlicht: (2020)
von: Kobayashi, Atsuya, et al.
Veröffentlicht: (2020)
Differentiable Black-box and Gray-box Modeling of Nonlinear Audio Effects
von: Comunità, Marco, et al.
Veröffentlicht: (2025)
von: Comunità, Marco, et al.
Veröffentlicht: (2025)
Enhancing Lyrics Transcription on Music Mixtures with Consistency Loss
von: Huang, Jiawen, et al.
Veröffentlicht: (2025)
von: Huang, Jiawen, et al.
Veröffentlicht: (2025)
Visual-based spatial audio generation system for multi-speaker environments
von: Liu, Xiaojing, et al.
Veröffentlicht: (2025)
von: Liu, Xiaojing, et al.
Veröffentlicht: (2025)
Prototype: A Keyword Spotting-Based Intelligent Audio SoC for IoT
von: Liang, Huihong, et al.
Veröffentlicht: (2025)
von: Liang, Huihong, et al.
Veröffentlicht: (2025)
6KSFx Synth Dataset
von: Garcia, Nelly, et al.
Veröffentlicht: (2025)
von: Garcia, Nelly, et al.
Veröffentlicht: (2025)
The effect of self-motion and room familiarity on sound source localization in virtual environments
von: Isserstedt, Niklas, et al.
Veröffentlicht: (2024)
von: Isserstedt, Niklas, et al.
Veröffentlicht: (2024)
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning
von: Tuncay, Ludovic, et al.
Veröffentlicht: (2025)
von: Tuncay, Ludovic, et al.
Veröffentlicht: (2025)
Learning Control of Neural Sound Effects Synthesis from Physically Inspired Models
von: Zong, Yisu, et al.
Veröffentlicht: (2025)
von: Zong, Yisu, et al.
Veröffentlicht: (2025)
Acoustic identification of individual animals with hierarchical contrastive learning
von: Nolasco, Ines, et al.
Veröffentlicht: (2024)
von: Nolasco, Ines, et al.
Veröffentlicht: (2024)
YourMT3+: Multi-instrument Music Transcription with Enhanced Transformer Architectures and Cross-dataset Stem Augmentation
von: Chang, Sungkyun, et al.
Veröffentlicht: (2024)
von: Chang, Sungkyun, et al.
Veröffentlicht: (2024)
SCRAPL: Scattering Transform with Random Paths for Machine Learning
von: Mitcheltree, Christopher, et al.
Veröffentlicht: (2026)
von: Mitcheltree, Christopher, et al.
Veröffentlicht: (2026)
Hidden bawls, whispers, and yelps: can text be made to sound more than just its words?
von: Pataca, Caluã de Lacerda, et al.
Veröffentlicht: (2022)
von: Pataca, Caluã de Lacerda, et al.
Veröffentlicht: (2022)
Seeing Beyond Sound: Visualization and Abstraction in Audio Data Representation
von: Blum'e, Ashlae
Veröffentlicht: (2025)
von: Blum'e, Ashlae
Veröffentlicht: (2025)
SACM: SEEG-Audio Contrastive Matching for Chinese Speech Decoding
von: Wang, Hongbin, et al.
Veröffentlicht: (2025)
von: Wang, Hongbin, et al.
Veröffentlicht: (2025)
FeatureSense: Protecting Speaker Attributes in Always-On Audio Sensing System
von: Chhaglani, Bhawana, et al.
Veröffentlicht: (2025)
von: Chhaglani, Bhawana, et al.
Veröffentlicht: (2025)
Learning Vocal-Tract Area and Radiation with a Physics-Informed Webster Model
von: Lu, Minhui, et al.
Veröffentlicht: (2026)
von: Lu, Minhui, et al.
Veröffentlicht: (2026)
Improving Neural Pitch Estimation with SWIPE Kernels
von: Marttila, David, et al.
Veröffentlicht: (2025)
von: Marttila, David, et al.
Veröffentlicht: (2025)
A Mapping Strategy for Interacting with Latent Audio Synthesis Using Artistic Materials
von: Zheng, Shuoyang, et al.
Veröffentlicht: (2024)
von: Zheng, Shuoyang, et al.
Veröffentlicht: (2024)
Sound2Hap: Learning Audio-to-Vibrotactile Haptic Generation from Human Ratings
von: Li, Yinan, et al.
Veröffentlicht: (2026)
von: Li, Yinan, et al.
Veröffentlicht: (2026)
How Private is Low-Frequency Speech Audio in the Wild? An Analysis of Verbal Intelligibility by Humans and Machines
von: Liu, Ailin, et al.
Veröffentlicht: (2024)
von: Liu, Ailin, et al.
Veröffentlicht: (2024)
DESAMO: A Device for Elder-Friendly Smart Homes Powered by Embedded LLM with Audio Modality
von: Choi, Youngwon, et al.
Veröffentlicht: (2025)
von: Choi, Youngwon, et al.
Veröffentlicht: (2025)
chatter: a Python library for applying information theory and AI/ML models to animal communication
von: Youngblood, Mason
Veröffentlicht: (2025)
von: Youngblood, Mason
Veröffentlicht: (2025)
Universal Music Representations? Evaluating Foundation Models on World Music Corpora
von: Papaioannou, Charilaos, et al.
Veröffentlicht: (2025)
von: Papaioannou, Charilaos, et al.
Veröffentlicht: (2025)
An automatic mixing speech enhancement system for multi-track audio
von: Liu, Xiaojing, et al.
Veröffentlicht: (2024)
von: Liu, Xiaojing, et al.
Veröffentlicht: (2024)
Generalized Multi-Source Inference for Text Conditioned Music Diffusion Models
von: Postolache, Emilian, et al.
Veröffentlicht: (2024)
von: Postolache, Emilian, et al.
Veröffentlicht: (2024)
A Data-Driven Analysis of Robust Automatic Piano Transcription
von: Edwards, Drew, et al.
Veröffentlicht: (2024)
von: Edwards, Drew, et al.
Veröffentlicht: (2024)
Accelerating Audio Research with Robotic Dummy Heads
von: Lu, Austin, et al.
Veröffentlicht: (2025)
von: Lu, Austin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ST-ITO: Controlling Audio Effects for Style Transfer with Inference-Time Optimization
von: Steinmetz, Christian J., et al.
Veröffentlicht: (2024) -
GraFPrint: A GNN-Based Approach for Audio Identification
von: Bhattacharjee, Aditya, et al.
Veröffentlicht: (2024) -
Scalable Evaluation for Audio Identification via Synthetic Latent Fingerprint Generation
von: Bhattacharjee, Aditya, et al.
Veröffentlicht: (2025) -
Learning Music Audio Representations With Limited Data
von: Plachouras, Christos, et al.
Veröffentlicht: (2025) -
LC-Protonets: Multi-Label Few-Shot Learning for World Music Audio Tagging
von: Papaioannou, Charilaos, et al.
Veröffentlicht: (2024)