Audio-to-Score Conversion Model Based on Whisper methodology
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Hongyao, Sun, Bohang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
WhisperRT -- Turning Whisper into a Causal Streaming Model
por: Krichli, Tomer, et al.
Publicado: (2025)
por: Krichli, Tomer, et al.
Publicado: (2025)
Whisper-GPT: A Hybrid Representation Audio Large Language Model
por: Verma, Prateek
Publicado: (2024)
por: Verma, Prateek
Publicado: (2024)
RUMAA: Repeat-Aware Unified Music Audio Analysis for Score-Performance Alignment, Transcription, and Mistake Detection
por: Chang, Sungkyun, et al.
Publicado: (2025)
por: Chang, Sungkyun, et al.
Publicado: (2025)
Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities
por: Ghosh, Sreyan, et al.
Publicado: (2025)
por: Ghosh, Sreyan, et al.
Publicado: (2025)
Improving Text-To-Audio Models with Synthetic Captions
por: Kong, Zhifeng, et al.
Publicado: (2024)
por: Kong, Zhifeng, et al.
Publicado: (2024)
ETTA: Elucidating the Design Space of Text-to-Audio Models
por: Lee, Sang-gil, et al.
Publicado: (2024)
por: Lee, Sang-gil, et al.
Publicado: (2024)
tinyCLAP: Distilling Constrastive Language-Audio Pretrained Models
por: Paissan, Francesco, et al.
Publicado: (2023)
por: Paissan, Francesco, et al.
Publicado: (2023)
AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension
por: Yang, Qian, et al.
Publicado: (2024)
por: Yang, Qian, et al.
Publicado: (2024)
Simul-Whisper: Attention-Guided Streaming Whisper with Truncation Detection
por: Wang, Haoyu, et al.
Publicado: (2024)
por: Wang, Haoyu, et al.
Publicado: (2024)
Understanding Sounds, Missing the Questions: The Challenge of Object Hallucination in Large Audio-Language Models
por: Kuan, Chun-Yi, et al.
Publicado: (2024)
por: Kuan, Chun-Yi, et al.
Publicado: (2024)
TTSDS -- Text-to-Speech Distribution Score
por: Minixhofer, Christoph, et al.
Publicado: (2024)
por: Minixhofer, Christoph, et al.
Publicado: (2024)
ChipChat: Low-Latency Cascaded Conversational Agent in MLX
por: Likhomanenko, Tatiana, et al.
Publicado: (2025)
por: Likhomanenko, Tatiana, et al.
Publicado: (2025)
T-CLAP: Temporal-Enhanced Contrastive Language-Audio Pretraining
por: Yuan, Yi, et al.
Publicado: (2024)
por: Yuan, Yi, et al.
Publicado: (2024)
Cascaded Cross-Modal Transformer for Audio-Textual Classification
por: Ristea, Nicolae-Catalin, et al.
Publicado: (2024)
por: Ristea, Nicolae-Catalin, et al.
Publicado: (2024)
Exploring Fine-Tuning of Large Audio Language Models for Spoken Language Understanding under Limited Speech Data
por: Choi, Youngwon, et al.
Publicado: (2025)
por: Choi, Youngwon, et al.
Publicado: (2025)
Spontaneous Speech-Based Suicide Risk Detection Using Whisper and Large Language Models
por: Cui, Ziyun, et al.
Publicado: (2024)
por: Cui, Ziyun, et al.
Publicado: (2024)
AVCap: Leveraging Audio-Visual Features as Text Tokens for Captioning
por: Kim, Jongsuk, et al.
Publicado: (2024)
por: Kim, Jongsuk, et al.
Publicado: (2024)
Score-Based Training for Energy-Based TTS Models
por: Sun, Wanli, et al.
Publicado: (2025)
por: Sun, Wanli, et al.
Publicado: (2025)
Multiple-Instance, Cascaded Classification for Keyword Spotting in Narrow-Band Audio
por: AbdulKader, Ahmad, et al.
Publicado: (2017)
por: AbdulKader, Ahmad, et al.
Publicado: (2017)
Remastering Divide and Remaster: A Cinematic Audio Source Separation Dataset with Multilingual Support
por: Watcharasupat, Karn N., et al.
Publicado: (2024)
por: Watcharasupat, Karn N., et al.
Publicado: (2024)
Conversational Rubert for Detecting Competitive Interruptions in ASR-Transcribed Dialogues
por: Galimzianov, Dmitrii, et al.
Publicado: (2024)
por: Galimzianov, Dmitrii, et al.
Publicado: (2024)
Towards General-Purpose Text-Instruction-Guided Voice Conversion
por: Kuan, Chun-Yi, et al.
Publicado: (2023)
por: Kuan, Chun-Yi, et al.
Publicado: (2023)
Predicting User Intents and Musical Attributes from Music Discovery Conversations
por: Kwon, Daeyong, et al.
Publicado: (2024)
por: Kwon, Daeyong, et al.
Publicado: (2024)
AlignAtt: Using Attention-based Audio-Translation Alignments as a Guide for Simultaneous Speech Translation
por: Papi, Sara, et al.
Publicado: (2023)
por: Papi, Sara, et al.
Publicado: (2023)
TelME: Teacher-leading Multimodal Fusion Network for Emotion Recognition in Conversation
por: Yun, Taeyang, et al.
Publicado: (2024)
por: Yun, Taeyang, et al.
Publicado: (2024)
Cross-lingual Text-To-Speech with Flow-based Voice Conversion for Improved Pronunciation
por: Ellinas, Nikolaos, et al.
Publicado: (2022)
por: Ellinas, Nikolaos, et al.
Publicado: (2022)
TVTSyn: Content-Synchronous Time-Varying Timbre for Streaming Voice Conversion and Anonymization
por: Quamer, Waris, et al.
Publicado: (2026)
por: Quamer, Waris, et al.
Publicado: (2026)
Improving Whisper's Recognition Performance for Under-Represented Language Kazakh Leveraging Unpaired Speech and Text
por: Li, Jinpeng, et al.
Publicado: (2024)
por: Li, Jinpeng, et al.
Publicado: (2024)
The ART of Conversation: Measuring Phonetic Convergence and Deliberate Imitation in L2-Speech with a Siamese RNN
por: Yuan, Zheng, et al.
Publicado: (2023)
por: Yuan, Zheng, et al.
Publicado: (2023)
MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models
por: Weck, Benno, et al.
Publicado: (2024)
por: Weck, Benno, et al.
Publicado: (2024)
Energy-Based Models with Applications to Speech and Language Processing
por: Ou, Zhijian
Publicado: (2024)
por: Ou, Zhijian
Publicado: (2024)
Can Whisper perform speech-based in-context learning?
por: Wang, Siyin, et al.
Publicado: (2023)
por: Wang, Siyin, et al.
Publicado: (2023)
On the Semantic Latent Space of Diffusion-Based Text-to-Speech Models
por: Varshavsky-Hassid, Miri, et al.
Publicado: (2024)
por: Varshavsky-Hassid, Miri, et al.
Publicado: (2024)
Whispy: Adapting STT Whisper Models to Real-Time Environments
por: Bevilacqua, Antonio, et al.
Publicado: (2024)
por: Bevilacqua, Antonio, et al.
Publicado: (2024)
A Joint Spectro-Temporal Relational Thinking Based Acoustic Modeling Framework
por: Nan, Zheng, et al.
Publicado: (2024)
por: Nan, Zheng, et al.
Publicado: (2024)
AudioBench: A Universal Benchmark for Audio Large Language Models
por: Wang, Bin, et al.
Publicado: (2024)
por: Wang, Bin, et al.
Publicado: (2024)
Spontaneous Style Text-to-Speech Synthesis with Controllable Spontaneous Behaviors Based on Language Models
por: Li, Weiqin, et al.
Publicado: (2024)
por: Li, Weiqin, et al.
Publicado: (2024)
Muting Whisper: A Universal Acoustic Adversarial Attack on Speech Foundation Models
por: Raina, Vyas, et al.
Publicado: (2024)
por: Raina, Vyas, et al.
Publicado: (2024)
Controlling Whisper: Universal Acoustic Adversarial Attacks to Control Speech Foundation Models
por: Raina, Vyas, et al.
Publicado: (2024)
por: Raina, Vyas, et al.
Publicado: (2024)
Leveraging Whisper Embeddings for Audio-based Lyrics Matching
por: Mancini, Eleonora, et al.
Publicado: (2025)
por: Mancini, Eleonora, et al.
Publicado: (2025)
Ejemplares similares
-
WhisperRT -- Turning Whisper into a Causal Streaming Model
por: Krichli, Tomer, et al.
Publicado: (2025) -
Whisper-GPT: A Hybrid Representation Audio Large Language Model
por: Verma, Prateek
Publicado: (2024) -
RUMAA: Repeat-Aware Unified Music Audio Analysis for Score-Performance Alignment, Transcription, and Mistake Detection
por: Chang, Sungkyun, et al.
Publicado: (2025) -
Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities
por: Ghosh, Sreyan, et al.
Publicado: (2025) -
Improving Text-To-Audio Models with Synthetic Captions
por: Kong, Zhifeng, et al.
Publicado: (2024)