Multimodal Lyrics-Rhythm Matching
Fuente:
arXiv
Saved in:
| Main Authors: | Liao, Callie C., Liao, Duoduo, Guessford, Jesse |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Relationships between Keywords and Strong Beats in Lyrical Music
by: Liao, Callie C., et al.
Published: (2024)
by: Liao, Callie C., et al.
Published: (2024)
Automatic Time Signature Determination for New Scores Using Lyrics for Latent Rhythmic Structure
by: Liao, Callie C., et al.
Published: (2023)
by: Liao, Callie C., et al.
Published: (2023)
Dynamic Multi-Species Bird Soundscape Generation with Acoustic Patterning and 3D Spatialization
by: Zhang, Ellie L., et al.
Published: (2025)
by: Zhang, Ellie L., et al.
Published: (2025)
Leveraging Whisper Embeddings for Audio-based Lyrics Matching
by: Mancini, Eleonora, et al.
Published: (2025)
by: Mancini, Eleonora, et al.
Published: (2025)
Vocal Melody Construction for Persian Lyrics Using LSTM Recurrent Neural Networks
by: Jafari, Farshad, et al.
Published: (2024)
by: Jafari, Farshad, et al.
Published: (2024)
A Real-Time Lyrics Alignment System Using Chroma And Phonetic Features For Classical Vocal Performance
by: Park, Jiyun, et al.
Published: (2024)
by: Park, Jiyun, et al.
Published: (2024)
Lyrics Transcription for Humans: A Readability-Aware Benchmark
by: Cífka, Ondřej, et al.
Published: (2024)
by: Cífka, Ondřej, et al.
Published: (2024)
An Investigation of Test-time Adaptation for Audio Classification under Background Noise
by: Shao, Weichuang, et al.
Published: (2025)
by: Shao, Weichuang, et al.
Published: (2025)
Audio Match Cutting: Finding and Creating Matching Audio Transitions in Movies and Videos
by: Fedorishin, Dennis, et al.
Published: (2024)
by: Fedorishin, Dennis, et al.
Published: (2024)
DiarizationLM: Speaker Diarization Post-Processing with Large Language Models
by: Wang, Quan, et al.
Published: (2024)
by: Wang, Quan, et al.
Published: (2024)
Adversarial Speaker Distillation for Countermeasure Model on Automatic Speaker Verification
by: Liao, Yen-Lun, et al.
Published: (2022)
by: Liao, Yen-Lun, et al.
Published: (2022)
Drax: Speech Recognition with Discrete Flow Matching
by: Navon, Aviv, et al.
Published: (2025)
by: Navon, Aviv, et al.
Published: (2025)
Synthesizer Sound Matching Using Audio Spectrogram Transformers
by: Bruford, Fred, et al.
Published: (2024)
by: Bruford, Fred, et al.
Published: (2024)
FlowTSE: Target Speaker Extraction with Flow Matching
by: Navon, Aviv, et al.
Published: (2025)
by: Navon, Aviv, et al.
Published: (2025)
USM-SCD: Multilingual Speaker Change Detection Based on Large Pretrained Foundation Models
by: Zhao, Guanlong, et al.
Published: (2023)
by: Zhao, Guanlong, et al.
Published: (2023)
Improving Inference-Time Optimisation for Vocal Effects Style Transfer with a Gaussian Prior
by: Yu, Chin-Yun, et al.
Published: (2025)
by: Yu, Chin-Yun, et al.
Published: (2025)
Keyword Spotting with Hyper-Matched Filters for Small Footprint Devices
by: Segal-Feldman, Yael, et al.
Published: (2025)
by: Segal-Feldman, Yael, et al.
Published: (2025)
UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching
by: Glazer, Neta, et al.
Published: (2025)
by: Glazer, Neta, et al.
Published: (2025)
Timbre-Trap: A Low-Resource Framework for Instrument-Agnostic Music Transcription
by: Cwitkowitz, Frank, et al.
Published: (2023)
by: Cwitkowitz, Frank, et al.
Published: (2023)
Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching
by: Zuo, Jialong, et al.
Published: (2025)
by: Zuo, Jialong, et al.
Published: (2025)
FlowMAC: Conditional Flow Matching for Audio Coding at Low Bit Rates
by: Pia, Nicola, et al.
Published: (2024)
by: Pia, Nicola, et al.
Published: (2024)
Unsupervised Rhythm and Voice Conversion to Improve ASR on Dysarthric Speech
by: Hajal, Karl El, et al.
Published: (2025)
by: Hajal, Karl El, et al.
Published: (2025)
Unsupervised Rhythm and Voice Conversion of Dysarthric to Healthy Speech for ASR
by: Hajal, Karl El, et al.
Published: (2025)
by: Hajal, Karl El, et al.
Published: (2025)
LM2D: Lyrics- and Music-Driven Dance Synthesis
by: Yin, Wenjie, et al.
Published: (2024)
by: Yin, Wenjie, et al.
Published: (2024)
LIWhiz: A Non-Intrusive Lyric Intelligibility Prediction System for the Cadenza Challenge
by: Shekar, Ram C. M. C., et al.
Published: (2025)
by: Shekar, Ram C. M. C., et al.
Published: (2025)
Diffusion-Based Speech Enhancement in Matched and Mismatched Conditions Using a Heun-Based Sampler
by: Gonzalez, Philippe, et al.
Published: (2023)
by: Gonzalez, Philippe, et al.
Published: (2023)
Semantic-Aware Interpretable Multimodal Music Auto-Tagging
by: Patakis, Andreas, et al.
Published: (2025)
by: Patakis, Andreas, et al.
Published: (2025)
RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching
by: Yang, Jinhyeok, et al.
Published: (2026)
by: Yang, Jinhyeok, et al.
Published: (2026)
LLark: A Multimodal Instruction-Following Language Model for Music
by: Gardner, Josh, et al.
Published: (2023)
by: Gardner, Josh, et al.
Published: (2023)
Enhancing Lyrics Transcription on Music Mixtures with Consistency Loss
by: Huang, Jiawen, et al.
Published: (2025)
by: Huang, Jiawen, et al.
Published: (2025)
Multimodal Attention Merging for Improved Speech Recognition and Audio Event Classification
by: Sundar, Anirudh S., et al.
Published: (2023)
by: Sundar, Anirudh S., et al.
Published: (2023)
Bone-conduction Guided Multimodal Speech Enhancement with Conditional Diffusion Models
by: Khanagha, Sina, et al.
Published: (2026)
by: Khanagha, Sina, et al.
Published: (2026)
Variable Bitrate Residual Vector Quantization for Audio Coding
by: Chae, Yunkee, et al.
Published: (2024)
by: Chae, Yunkee, et al.
Published: (2024)
MERGE -- A Bimodal Audio-Lyrics Dataset for Static Music Emotion Recognition
by: Louro, Pedro Lima, et al.
Published: (2024)
by: Louro, Pedro Lima, et al.
Published: (2024)
Exploiting Music Source Separation for Automatic Lyrics Transcription with Whisper
by: Syed, Jaza, et al.
Published: (2025)
by: Syed, Jaza, et al.
Published: (2025)
CogniVoice: Multimodal and Multilingual Fusion Networks for Mild Cognitive Impairment Assessment from Spontaneous Speech
by: Cheng, Jiali, et al.
Published: (2024)
by: Cheng, Jiali, et al.
Published: (2024)
EVA-GAN: Enhanced Various Audio Generation via Scalable Generative Adversarial Networks
by: Liao, Shijia, et al.
Published: (2024)
by: Liao, Shijia, et al.
Published: (2024)
Speech Rhythm-Based Speaker Embeddings Extraction from Phonemes and Phoneme Duration for Multi-Speaker Speech Synthesis
by: Fujita, Kenichi, et al.
Published: (2024)
by: Fujita, Kenichi, et al.
Published: (2024)
Generating Rhythm Game Music with Jukebox
by: Yan, Nicholas
Published: (2023)
by: Yan, Nicholas
Published: (2023)
Real-Time Streaming Mel Vocoding with Generative Flow Matching
by: Welker, Simon, et al.
Published: (2025)
by: Welker, Simon, et al.
Published: (2025)
Similar Items
-
Relationships between Keywords and Strong Beats in Lyrical Music
by: Liao, Callie C., et al.
Published: (2024) -
Automatic Time Signature Determination for New Scores Using Lyrics for Latent Rhythmic Structure
by: Liao, Callie C., et al.
Published: (2023) -
Dynamic Multi-Species Bird Soundscape Generation with Acoustic Patterning and 3D Spatialization
by: Zhang, Ellie L., et al.
Published: (2025) -
Leveraging Whisper Embeddings for Audio-based Lyrics Matching
by: Mancini, Eleonora, et al.
Published: (2025) -
Vocal Melody Construction for Persian Lyrics Using LSTM Recurrent Neural Networks
by: Jafari, Farshad, et al.
Published: (2024)