Investigating Design Choices in Joint-Embedding Predictive Architectures for General Audio Representation Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Riou, Alain, Lattner, Stefan, Hadjeres, Gaëtan, Peeters, Geoffroy |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Zero-shot Musical Stem Retrieval with Joint-Embedding Predictive Architectures
von: Riou, Alain, et al.
Veröffentlicht: (2024)
von: Riou, Alain, et al.
Veröffentlicht: (2024)
Stem-JEPA: A Joint-Embedding Predictive Architecture for Musical Stem Compatibility Estimation
von: Riou, Alain, et al.
Veröffentlicht: (2024)
von: Riou, Alain, et al.
Veröffentlicht: (2024)
PESTO: Pitch Estimation with Self-supervised Transposition-equivariant Objective
von: Riou, Alain, et al.
Veröffentlicht: (2023)
von: Riou, Alain, et al.
Veröffentlicht: (2023)
PESTO: Real-Time Pitch Estimation with Self-supervised Transposition-equivariant Objective
von: Riou, Alain, et al.
Veröffentlicht: (2025)
von: Riou, Alain, et al.
Veröffentlicht: (2025)
Translation-Equivariant Self-Supervised Learning for Pitch Estimation with Optimal Transport
von: Torres, Bernardo, et al.
Veröffentlicht: (2025)
von: Torres, Bernardo, et al.
Veröffentlicht: (2025)
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning
von: Tuncay, Ludovic, et al.
Veröffentlicht: (2025)
von: Tuncay, Ludovic, et al.
Veröffentlicht: (2025)
Music2Latent2: Audio Compression with Summary Embeddings and Autoregressive Decoding
von: Pasini, Marco, et al.
Veröffentlicht: (2025)
von: Pasini, Marco, et al.
Veröffentlicht: (2025)
CoDiCodec: Unifying Continuous and Discrete Compressed Representations of Audio
von: Pasini, Marco, et al.
Veröffentlicht: (2025)
von: Pasini, Marco, et al.
Veröffentlicht: (2025)
Estimating Musical Surprisal from Audio in Autoregressive Diffusion Model Noise Spaces
von: Bjare, Mathias Rose, et al.
Veröffentlicht: (2025)
von: Bjare, Mathias Rose, et al.
Veröffentlicht: (2025)
Estimating Musical Surprisal in Audio
von: Bjare, Mathias Rose, et al.
Veröffentlicht: (2025)
von: Bjare, Mathias Rose, et al.
Veröffentlicht: (2025)
Assessing the Alignment of Audio Representations with Timbre Similarity Ratings
von: Tian, Haokun, et al.
Veröffentlicht: (2025)
von: Tian, Haokun, et al.
Veröffentlicht: (2025)
Automatic Music Sample Identification with Multi-Track Contrastive Learning
von: Riou, Alain, et al.
Veröffentlicht: (2025)
von: Riou, Alain, et al.
Veröffentlicht: (2025)
EnCLAP: Combining Neural Audio Codec and Audio-Text Joint Embedding for Automated Audio Captioning
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
Soft Clustering Anchors for Self-Supervised Speech Representation Learning in Joint Embedding Prediction Architectures
von: Ioannides, Georgios, et al.
Veröffentlicht: (2026)
von: Ioannides, Georgios, et al.
Veröffentlicht: (2026)
Do Joint Language-Audio Embeddings Encode Perceptual Timbre Semantics?
von: Deng, Qixin, et al.
Veröffentlicht: (2025)
von: Deng, Qixin, et al.
Veröffentlicht: (2025)
Unify Variables in Neural Scaling Laws for General Audio Representations via Embedding Effective Rank
von: Deng, Xuyao, et al.
Veröffentlicht: (2025)
von: Deng, Xuyao, et al.
Veröffentlicht: (2025)
Diff-A-Riff: Musical Accompaniment Co-creation via Latent Diffusion Models
von: Nistal, Javier, et al.
Veröffentlicht: (2024)
von: Nistal, Javier, et al.
Veröffentlicht: (2024)
Controlling Surprisal in Music Generation via Information Content Curve Matching
von: Bjare, Mathias Rose, et al.
Veröffentlicht: (2024)
von: Bjare, Mathias Rose, et al.
Veröffentlicht: (2024)
Diffusion Timbre Transfer Via Mutual Information Guided Inpainting
von: Lee, Ching Ho, et al.
Veröffentlicht: (2026)
von: Lee, Ching Ho, et al.
Veröffentlicht: (2026)
Audio Mamba: Bidirectional State Space Model for Audio Representation Learning
von: Erol, Mehmet Hamza, et al.
Veröffentlicht: (2024)
von: Erol, Mehmet Hamza, et al.
Veröffentlicht: (2024)
SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding
von: Guinot, Julien, et al.
Veröffentlicht: (2025)
von: Guinot, Julien, et al.
Veröffentlicht: (2025)
Embedding Alignment in Code Generation for Audio
von: Kouteili, Sam, et al.
Veröffentlicht: (2025)
von: Kouteili, Sam, et al.
Veröffentlicht: (2025)
Representation-Regularized Convolutional Audio Transformer for Audio Understanding
von: Han, Bing, et al.
Veröffentlicht: (2026)
von: Han, Bing, et al.
Veröffentlicht: (2026)
Audio Mamba: Selective State Spaces for Self-Supervised Audio Representations
von: Yadav, Sarthak, et al.
Veröffentlicht: (2024)
von: Yadav, Sarthak, et al.
Veröffentlicht: (2024)
Blind estimation of audio effects using an auto-encoder approach and differentiable digital signal processing
von: Peladeau, Côme, et al.
Veröffentlicht: (2023)
von: Peladeau, Côme, et al.
Veröffentlicht: (2023)
Continuous Autoregressive Models with Noise Augmentation Avoid Error Accumulation
von: Pasini, Marco, et al.
Veröffentlicht: (2024)
von: Pasini, Marco, et al.
Veröffentlicht: (2024)
Audio Deepfake Attribution: An Initial Dataset and Investigation
von: Yan, Xinrui, et al.
Veröffentlicht: (2022)
von: Yan, Xinrui, et al.
Veröffentlicht: (2022)
BrewCLIP: A Bifurcated Representation Learning Framework for Audio-Visual Retrieval
von: Lu, Zhenyu, et al.
Veröffentlicht: (2024)
von: Lu, Zhenyu, et al.
Veröffentlicht: (2024)
Exploring Musical Roots: Applying Audio Embeddings to Empower Influence Attribution for a Generative Music Model
von: Barnett, Julia, et al.
Veröffentlicht: (2024)
von: Barnett, Julia, et al.
Veröffentlicht: (2024)
CoughViT: A Self-Supervised Vision Transformer for Cough Audio Representation Learning
von: Luong, Justin, et al.
Veröffentlicht: (2025)
von: Luong, Justin, et al.
Veröffentlicht: (2025)
OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model
von: Li, Maomao, et al.
Veröffentlicht: (2026)
von: Li, Maomao, et al.
Veröffentlicht: (2026)
A Contrastive Self-Supervised Learning scheme for beat tracking amenable to few-shot learning
von: Gagnere, Antonin, et al.
Veröffentlicht: (2024)
von: Gagnere, Antonin, et al.
Veröffentlicht: (2024)
Controlling Contrastive Self-Supervised Learning with Knowledge-Driven Multiple Hypothesis: Application to Beat Tracking
von: Gagnere, Antonin, et al.
Veröffentlicht: (2025)
von: Gagnere, Antonin, et al.
Veröffentlicht: (2025)
Joint Learning of Emotions in Music and Generalized Sounds
von: Simonetta, Federico, et al.
Veröffentlicht: (2024)
von: Simonetta, Federico, et al.
Veröffentlicht: (2024)
Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation
von: Yan, Canxiang, et al.
Veröffentlicht: (2025)
von: Yan, Canxiang, et al.
Veröffentlicht: (2025)
Learning Semantic Information from Raw Audio Signal Using Both Contextual and Phonetic Representations
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
WavShape: Information-Theoretic Speech Representation Learning for Fair and Privacy-Aware Audio Processing
von: Baser, Oguzhan, et al.
Veröffentlicht: (2025)
von: Baser, Oguzhan, et al.
Veröffentlicht: (2025)
Towards Leveraging Contrastively Pretrained Neural Audio Embeddings for Recommender Tasks
von: Grötschla, Florian, et al.
Veröffentlicht: (2024)
von: Grötschla, Florian, et al.
Veröffentlicht: (2024)
Investigation of Whisper ASR Hallucinations Induced by Non-Speech Audio
von: Barański, Mateusz, et al.
Veröffentlicht: (2025)
von: Barański, Mateusz, et al.
Veröffentlicht: (2025)
AudioTurbo: Fast Text-to-Audio Generation with Rectified Diffusion
von: Zhao, Junqi, et al.
Veröffentlicht: (2025)
von: Zhao, Junqi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Zero-shot Musical Stem Retrieval with Joint-Embedding Predictive Architectures
von: Riou, Alain, et al.
Veröffentlicht: (2024) -
Stem-JEPA: A Joint-Embedding Predictive Architecture for Musical Stem Compatibility Estimation
von: Riou, Alain, et al.
Veröffentlicht: (2024) -
PESTO: Pitch Estimation with Self-supervised Transposition-equivariant Objective
von: Riou, Alain, et al.
Veröffentlicht: (2023) -
PESTO: Real-Time Pitch Estimation with Self-supervised Transposition-equivariant Objective
von: Riou, Alain, et al.
Veröffentlicht: (2025) -
Translation-Equivariant Self-Supervised Learning for Pitch Estimation with Optimal Transport
von: Torres, Bernardo, et al.
Veröffentlicht: (2025)