Do we need more complex representations for structure? A comparison of note duration representation for Music Transformers
Fuente:
arXiv
Guardado en:
| Autores principales: | Souza, Gabriel, Figueiredo, Flavio, Machado, Alexei, Guimarães, Deborah |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Transformation of audio embeddings into interpretable, concept-based representations
por: Zhang, Alice, et al.
Publicado: (2025)
por: Zhang, Alice, et al.
Publicado: (2025)
LISTEN: Lightweight Industrial Sound-representable Transformer for Edge Notification
por: Han, Changheon, et al.
Publicado: (2025)
por: Han, Changheon, et al.
Publicado: (2025)
Houston we have a Divergence: A Subgroup Performance Analysis of ASR Models
por: Koudounas, Alkis, et al.
Publicado: (2024)
por: Koudounas, Alkis, et al.
Publicado: (2024)
Anticipatory Music Transformer
por: Thickstun, John, et al.
Publicado: (2023)
por: Thickstun, John, et al.
Publicado: (2023)
Rank-based loss for learning hierarchical representations
por: Nolasco, Ines, et al.
Publicado: (2021)
por: Nolasco, Ines, et al.
Publicado: (2021)
Do Foundational Audio Encoders Understand Music Structure?
por: Toyama, Keisuke, et al.
Publicado: (2025)
por: Toyama, Keisuke, et al.
Publicado: (2025)
Late fusion ensembles for speech recognition on diverse input audio representations
por: Jezidžić, Marin, et al.
Publicado: (2024)
por: Jezidžić, Marin, et al.
Publicado: (2024)
Exploring synthetic data for cross-speaker style transfer in style representation based TTS
por: Ueda, Lucas H., et al.
Publicado: (2024)
por: Ueda, Lucas H., et al.
Publicado: (2024)
EnCodecMAE: Leveraging neural codecs for universal audio representation learning
por: Pepino, Leonardo, et al.
Publicado: (2023)
por: Pepino, Leonardo, et al.
Publicado: (2023)
Hierarchical speaker representation for target speaker extraction
por: He, Shulin, et al.
Publicado: (2022)
por: He, Shulin, et al.
Publicado: (2022)
Acoustic-to-articulatory inversion for dysarthric speech: Are pre-trained self-supervised representations favorable?
por: Maharana, Sarthak Kumar, et al.
Publicado: (2023)
por: Maharana, Sarthak Kumar, et al.
Publicado: (2023)
Do we really need Self-Attention for Streaming Automatic Speech Recognition?
por: Dkhissi, Youness, et al.
Publicado: (2026)
por: Dkhissi, Youness, et al.
Publicado: (2026)
Exploring Transformer-Based Music Overpainting for Jazz Piano Variations
por: Row, Eleanor, et al.
Publicado: (2024)
por: Row, Eleanor, et al.
Publicado: (2024)
An Experimental Comparison Of Multi-view Self-supervised Methods For Music Tagging
por: Meseguer-Brocal, Gabriel, et al.
Publicado: (2024)
por: Meseguer-Brocal, Gabriel, et al.
Publicado: (2024)
AxLSTMs: learning self-supervised audio representations with xLSTMs
por: Yadav, Sarthak, et al.
Publicado: (2024)
por: Yadav, Sarthak, et al.
Publicado: (2024)
Towards generalisable and calibrated synthetic speech detection with self-supervised representations
por: Pascu, Octavian, et al.
Publicado: (2023)
por: Pascu, Octavian, et al.
Publicado: (2023)
The role of direct sound spherical harmonics representation in externalization using binaural reproduction
por: Miller, Eran, et al.
Publicado: (2024)
por: Miller, Eran, et al.
Publicado: (2024)
Building speech corpus with diverse voice characteristics for its prompt-based representation
por: Watanabe, Aya, et al.
Publicado: (2024)
por: Watanabe, Aya, et al.
Publicado: (2024)
A state-space representation of the boundary integral equation for room acoustic modelling
por: Ali, Randall, et al.
Publicado: (2026)
por: Ali, Randall, et al.
Publicado: (2026)
MusicRL: Aligning Music Generation to Human Preferences
por: Cideron, Geoffrey, et al.
Publicado: (2024)
por: Cideron, Geoffrey, et al.
Publicado: (2024)
A multimodal dynamical variational autoencoder for audiovisual speech representation learning
por: Sadok, Samir, et al.
Publicado: (2023)
por: Sadok, Samir, et al.
Publicado: (2023)
AI-Generated Music Detection and its Challenges
por: Afchar, Darius, et al.
Publicado: (2025)
por: Afchar, Darius, et al.
Publicado: (2025)
Zero-Shot Sing Voice Conversion: built upon clustering-based phoneme representations
por: Zhou, Wangjin, et al.
Publicado: (2024)
por: Zhou, Wangjin, et al.
Publicado: (2024)
Do neonates hear what we measure? Assessing neonatal ward soundscapes at the neonates ears
por: Lam, Bhan, et al.
Publicado: (2025)
por: Lam, Bhan, et al.
Publicado: (2025)
YourMT3+: Multi-instrument Music Transcription with Enhanced Transformer Architectures and Cross-dataset Stem Augmentation
por: Chang, Sungkyun, et al.
Publicado: (2024)
por: Chang, Sungkyun, et al.
Publicado: (2024)
Improving Musical Accompaniment Co-creation via Diffusion Transformers
por: Nistal, Javier, et al.
Publicado: (2024)
por: Nistal, Javier, et al.
Publicado: (2024)
Leave-One-EquiVariant: Alleviating invariance-related information loss in contrastive music representations
por: Guinot, Julien, et al.
Publicado: (2024)
por: Guinot, Julien, et al.
Publicado: (2024)
Self-supervised learning method using multiple sampling strategies for general-purpose audio representation
por: Kuroyanagi, Ibuki, et al.
Publicado: (2025)
por: Kuroyanagi, Ibuki, et al.
Publicado: (2025)
ProGress: Structured Music Generation via Graph Diffusion and Hierarchical Music Analysis
por: Ni-Hahn, Stephen, et al.
Publicado: (2025)
por: Ni-Hahn, Stephen, et al.
Publicado: (2025)
Score-informed Music Source Separation: Improving Synthetic-to-real Generalization in Classical Music
por: Tunturi, Eetu, et al.
Publicado: (2025)
por: Tunturi, Eetu, et al.
Publicado: (2025)
Integrating Text-to-Music Models with Language Models: Composing Long Structured Music Pieces
por: Atassi, Lilac
Publicado: (2024)
por: Atassi, Lilac
Publicado: (2024)
SMITIN: Self-Monitored Inference-Time INtervention for Generative Music Transformers
por: Koo, Junghyun, et al.
Publicado: (2024)
por: Koo, Junghyun, et al.
Publicado: (2024)
Score-Informed Transformer for Refining MIDI Velocity in Automatic Music Transcription
por: He, Zhanhong, et al.
Publicado: (2025)
por: He, Zhanhong, et al.
Publicado: (2025)
Editing Music with Melody and Text: Using ControlNet for Diffusion Transformer
por: Hou, Siyuan, et al.
Publicado: (2024)
por: Hou, Siyuan, et al.
Publicado: (2024)
Dirichlet process mixture model based on topologically augmented signal representation for clustering infant vocalizations
por: Bonafos, Guillem, et al.
Publicado: (2024)
por: Bonafos, Guillem, et al.
Publicado: (2024)
Generating Piano Music with Transformers: A Comparative Study of Scale, Data, and Metrics
por: Lehmkuhl, Jonathan, et al.
Publicado: (2025)
por: Lehmkuhl, Jonathan, et al.
Publicado: (2025)
MelodyT5: A Unified Score-to-Score Transformer for Symbolic Music Processing
por: Wu, Shangda, et al.
Publicado: (2024)
por: Wu, Shangda, et al.
Publicado: (2024)
Implicit neural representation with physics-informed neural networks for the reconstruction of the early part of room impulse responses
por: Pezzoli, Mirco, et al.
Publicado: (2023)
por: Pezzoli, Mirco, et al.
Publicado: (2023)
Evaluating Disentangled Representations for Controllable Music Generation
por: Ibáñez-Martínez, Laura, et al.
Publicado: (2026)
por: Ibáñez-Martínez, Laura, et al.
Publicado: (2026)
Tune It Up: Music Genre Transfer and Prediction
por: Samet, Fidan, et al.
Publicado: (2025)
por: Samet, Fidan, et al.
Publicado: (2025)
Ejemplares similares
-
Transformation of audio embeddings into interpretable, concept-based representations
por: Zhang, Alice, et al.
Publicado: (2025) -
LISTEN: Lightweight Industrial Sound-representable Transformer for Edge Notification
por: Han, Changheon, et al.
Publicado: (2025) -
Houston we have a Divergence: A Subgroup Performance Analysis of ASR Models
por: Koudounas, Alkis, et al.
Publicado: (2024) -
Anticipatory Music Transformer
por: Thickstun, John, et al.
Publicado: (2023) -
Rank-based loss for learning hierarchical representations
por: Nolasco, Ines, et al.
Publicado: (2021)