A Differentiable Alignment Framework for Sequence-to-Sequence Modeling via Optimal Transport
Fuente:
arXiv
Guardado en:
| Autores principales: | Kaloga, Yacouba, Kumar, Shashi, Motlicek, Petr, Kodrasi, Ina |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Towards interpretable emotion recognition: Identifying key features with machine learning
por: Kaloga, Yacouba, et al.
Publicado: (2025)
por: Kaloga, Yacouba, et al.
Publicado: (2025)
Multiview Canonical Correlation Analysis for Automatic Pathological Speech Detection
por: Kaloga, Yacouba, et al.
Publicado: (2024)
por: Kaloga, Yacouba, et al.
Publicado: (2024)
CLAP-Based Automatic Word Naming Recognition in Post-Stroke Aphasia
por: Kaloga, Yacouba, et al.
Publicado: (2026)
por: Kaloga, Yacouba, et al.
Publicado: (2026)
Graph Neural Networks for Parkinsons Disease Detection
por: Sheikh, Shakeel A., et al.
Publicado: (2024)
por: Sheikh, Shakeel A., et al.
Publicado: (2024)
Impact of Speech Mode in Automatic Pathological Speech Detection
por: Sheikh, Shakeel A., et al.
Publicado: (2024)
por: Sheikh, Shakeel A., et al.
Publicado: (2024)
Variational Autoencoder for Personalized Pathological Speech Enhancement
por: Hou, Mingchi, et al.
Publicado: (2025)
por: Hou, Mingchi, et al.
Publicado: (2025)
Suppressing Noise Disparity in Training Data for Automatic Pathological Speech Detection
por: Amiri, Mahdi, et al.
Publicado: (2024)
por: Amiri, Mahdi, et al.
Publicado: (2024)
Exploring In-Context Learning Capabilities of ChatGPT for Pathological Speech Detection
por: Amiri, Mahdi, et al.
Publicado: (2025)
por: Amiri, Mahdi, et al.
Publicado: (2025)
Sequence-Level Unsupervised Training in Speech Recognition: A Theoretical Study
por: Yang, Zijian, et al.
Publicado: (2026)
por: Yang, Zijian, et al.
Publicado: (2026)
Text-only adaptation in LLM-based ASR through text denoising
por: Carofilis, Andrés, et al.
Publicado: (2026)
por: Carofilis, Andrés, et al.
Publicado: (2026)
Discrete Optimal Transport and Voice Conversion
por: Selitskiy, Anton, et al.
Publicado: (2025)
por: Selitskiy, Anton, et al.
Publicado: (2025)
Unsupervised Harmonic Parameter Estimation Using Differentiable DSP and Spectral Optimal Transport
por: Torres, Bernardo, et al.
Publicado: (2023)
por: Torres, Bernardo, et al.
Publicado: (2023)
Optimal Transport Maps are Good Voice Converters
por: Asadulaev, Arip, et al.
Publicado: (2024)
por: Asadulaev, Arip, et al.
Publicado: (2024)
SCORE-SET: A dataset of GuitarPro files for Music Phrase Generation and Sequence Learning
por: Begari, Vishakh
Publicado: (2025)
por: Begari, Vishakh
Publicado: (2025)
What Does an Audio Deepfake Detector Focus on? A Study in the Time Domain
por: Grinberg, Petr, et al.
Publicado: (2025)
por: Grinberg, Petr, et al.
Publicado: (2025)
Sequence-to-Sequence Multi-Modal Speech In-Painting
por: Elyaderani, Mahsa Kadkhodaei, et al.
Publicado: (2024)
por: Elyaderani, Mahsa Kadkhodaei, et al.
Publicado: (2024)
Robust Multi-Modal Speech In-Painting: A Sequence-to-Sequence Approach
por: Elyaderani, Mahsa Kadkhodaei, et al.
Publicado: (2024)
por: Elyaderani, Mahsa Kadkhodaei, et al.
Publicado: (2024)
Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis
por: Lin, Weiwei, et al.
Publicado: (2025)
por: Lin, Weiwei, et al.
Publicado: (2025)
Fast Streaming Transducer ASR Prototyping via Knowledge Distillation with Whisper
por: Thorbecke, Iuliia, et al.
Publicado: (2024)
por: Thorbecke, Iuliia, et al.
Publicado: (2024)
On the Relation between Internal Language Model and Sequence Discriminative Training for Neural Transducers
por: Yang, Zijian, et al.
Publicado: (2023)
por: Yang, Zijian, et al.
Publicado: (2023)
TokenVerse: Towards Unifying Speech and NLP Tasks via Transducer-based ASR
por: Kumar, Shashi, et al.
Publicado: (2024)
por: Kumar, Shashi, et al.
Publicado: (2024)
Post-Training Embedding Alignment for Decoupling Enrollment and Runtime Speaker Recognition Models
por: Gao, Chenyang, et al.
Publicado: (2024)
por: Gao, Chenyang, et al.
Publicado: (2024)
Patient-Aware Feature Alignment for Robust Lung Sound Classification:Cohesion-Separation and Global Alignment Losses
por: Jeong, Seung Gyu, et al.
Publicado: (2025)
por: Jeong, Seung Gyu, et al.
Publicado: (2025)
Data-Driven Room Acoustic Modeling Via Differentiable Feedback Delay Networks With Learnable Delay Lines
por: Mezza, Alessandro Ilic, et al.
Publicado: (2024)
por: Mezza, Alessandro Ilic, et al.
Publicado: (2024)
Online Symbolic Music Alignment with Offline Reinforcement Learning
por: Peter, Silvan David
Publicado: (2023)
por: Peter, Silvan David
Publicado: (2023)
Latent Space Factorization in LoRA
por: Kumar, Shashi, et al.
Publicado: (2025)
por: Kumar, Shashi, et al.
Publicado: (2025)
Self-Supervised Frameworks for Speaker Verification via Bootstrapped Positive Sampling
por: Lepage, Theo, et al.
Publicado: (2025)
por: Lepage, Theo, et al.
Publicado: (2025)
Modulation Discovery with Differentiable Digital Signal Processing
por: Mitcheltree, Christopher, et al.
Publicado: (2025)
por: Mitcheltree, Christopher, et al.
Publicado: (2025)
BenSParX: A Robust Explainable Machine Learning Framework for Parkinson's Disease Detection from Bengali Conversational Speech
por: Hossain, Riad, et al.
Publicado: (2025)
por: Hossain, Riad, et al.
Publicado: (2025)
Knowledge Distillation for Speech Denoising by Latent Representation Alignment with Cosine Distance
por: Luong, Diep, et al.
Publicado: (2025)
por: Luong, Diep, et al.
Publicado: (2025)
A Real-Time Lyrics Alignment System Using Chroma And Phonetic Features For Classical Vocal Performance
por: Park, Jiyun, et al.
Publicado: (2024)
por: Park, Jiyun, et al.
Publicado: (2024)
TheGlueNote: Learned Representations for Robust and Flexible Note Alignment
por: Peter, Silvan David, et al.
Publicado: (2024)
por: Peter, Silvan David, et al.
Publicado: (2024)
Do Audio-Language Models Understand Linguistic Variations?
por: Selvakumar, Ramaneswaran, et al.
Publicado: (2024)
por: Selvakumar, Ramaneswaran, et al.
Publicado: (2024)
Differentiable All-pole Filters for Time-varying Audio Systems
por: Yu, Chin-Yun, et al.
Publicado: (2024)
por: Yu, Chin-Yun, et al.
Publicado: (2024)
Enhancing Speech Emotion Recognition Through Differentiable Architecture Search
por: Rajapakshe, Thejan, et al.
Publicado: (2023)
por: Rajapakshe, Thejan, et al.
Publicado: (2023)
Symbolic Music Generation with Non-Differentiable Rule Guided Diffusion
por: Huang, Yujia, et al.
Publicado: (2024)
por: Huang, Yujia, et al.
Publicado: (2024)
Development of Large Annotated Music Datasets using HMM-based Forced Viterbi Alignment
por: Joysingh, S. Johanan, et al.
Publicado: (2024)
por: Joysingh, S. Johanan, et al.
Publicado: (2024)
Ultra-lightweight Neural Differential DSP Vocoder For High Quality Speech Synthesis
por: Agrawal, Prabhav, et al.
Publicado: (2024)
por: Agrawal, Prabhav, et al.
Publicado: (2024)
UniPET-SPK: A Unified Framework for Parameter-Efficient Tuning of Pre-trained Speech Models for Robust Speaker Verification
por: Sang, Mufan, et al.
Publicado: (2025)
por: Sang, Mufan, et al.
Publicado: (2025)
MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis
por: Jiang, Ziyue, et al.
Publicado: (2025)
por: Jiang, Ziyue, et al.
Publicado: (2025)
Ejemplares similares
-
Towards interpretable emotion recognition: Identifying key features with machine learning
por: Kaloga, Yacouba, et al.
Publicado: (2025) -
Multiview Canonical Correlation Analysis for Automatic Pathological Speech Detection
por: Kaloga, Yacouba, et al.
Publicado: (2024) -
CLAP-Based Automatic Word Naming Recognition in Post-Stroke Aphasia
por: Kaloga, Yacouba, et al.
Publicado: (2026) -
Graph Neural Networks for Parkinsons Disease Detection
por: Sheikh, Shakeel A., et al.
Publicado: (2024) -
Impact of Speech Mode in Automatic Pathological Speech Detection
por: Sheikh, Shakeel A., et al.
Publicado: (2024)