Mixer is more than just a model
Fuente:
arXiv
Salvato in:
| Autori principali: | Ji, Qingfeng, Wang, Yuxin, Sun, Letong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ASM: Audio Spectrogram Mixer
di: Ji, Qingfeng, et al.
Pubblicazione: (2024)
di: Ji, Qingfeng, et al.
Pubblicazione: (2024)
Sparse Autoencoders Make Audio Foundation Models more Explainable
di: Mariotte, Théo, et al.
Pubblicazione: (2025)
di: Mariotte, Théo, et al.
Pubblicazione: (2025)
Mitigating Unauthorized Speech Synthesis for Voice Protection
di: Zhang, Zhisheng, et al.
Pubblicazione: (2024)
di: Zhang, Zhisheng, et al.
Pubblicazione: (2024)
Let There Be Sound: Reconstructing High Quality Speech from Silent Videos
di: Kim, Ji-Hoon, et al.
Pubblicazione: (2023)
di: Kim, Ji-Hoon, et al.
Pubblicazione: (2023)
Symbotunes: unified hub for symbolic music generative models
di: Skierś, Paweł, et al.
Pubblicazione: (2024)
di: Skierś, Paweł, et al.
Pubblicazione: (2024)
Self-Supervised Disentangled Representation Learning for Robust Target Speech Extraction
di: Mu, Zhaoxi, et al.
Pubblicazione: (2023)
di: Mu, Zhaoxi, et al.
Pubblicazione: (2023)
Collaborative Watermarking for Adversarial Speech Synthesis
di: Juvela, Lauri, et al.
Pubblicazione: (2023)
di: Juvela, Lauri, et al.
Pubblicazione: (2023)
F-StrIPE: Fast Structure-Informed Positional Encoding for Symbolic Music Generation
di: Agarwal, Manvi, et al.
Pubblicazione: (2025)
di: Agarwal, Manvi, et al.
Pubblicazione: (2025)
AnyEnhance: A Unified Generative Model with Prompt-Guidance and Self-Critic for Voice Enhancement
di: Zhang, Junan, et al.
Pubblicazione: (2025)
di: Zhang, Junan, et al.
Pubblicazione: (2025)
A Conditioned UNet for Music Source Separation
di: O'Hanlon, Ken, et al.
Pubblicazione: (2025)
di: O'Hanlon, Ken, et al.
Pubblicazione: (2025)
Efficient Multi-Model Fusion with Adversarial Complementary Representation Learning
di: Kang, Zuheng, et al.
Pubblicazione: (2024)
di: Kang, Zuheng, et al.
Pubblicazione: (2024)
Schrödinger Bridge Mamba for One-Step Speech Enhancement
di: Yang, Jing, et al.
Pubblicazione: (2025)
di: Yang, Jing, et al.
Pubblicazione: (2025)
Scaling Ambiguity: Augmenting Human Annotation in Speech Emotion Recognition with Audio-Language Models
di: Zhang, Wenda, et al.
Pubblicazione: (2026)
di: Zhang, Wenda, et al.
Pubblicazione: (2026)
MagiCodec: Simple Masked Gaussian-Injected Codec for High-Fidelity Reconstruction and Generation
di: Song, Yakun, et al.
Pubblicazione: (2025)
di: Song, Yakun, et al.
Pubblicazione: (2025)
Scaling Speech Tokenizers with Diffusion Autoencoders
di: Wang, Yuancheng, et al.
Pubblicazione: (2026)
di: Wang, Yuancheng, et al.
Pubblicazione: (2026)
CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
di: Du, Zhihao, et al.
Pubblicazione: (2024)
di: Du, Zhihao, et al.
Pubblicazione: (2024)
Self-Supervised Speech Quality Estimation and Enhancement Using Only Clean Speech
di: Fu, Szu-Wei, et al.
Pubblicazione: (2024)
di: Fu, Szu-Wei, et al.
Pubblicazione: (2024)
Multi-Metric Preference Alignment for Generative Speech Restoration
di: Zhang, Junan, et al.
Pubblicazione: (2025)
di: Zhang, Junan, et al.
Pubblicazione: (2025)
Repurposing Image Diffusion Models for Training-Free Music Style Transfer on Mel-spectrograms
di: Wang, Heehwan, et al.
Pubblicazione: (2024)
di: Wang, Heehwan, et al.
Pubblicazione: (2024)
Machine listening in a neonatal intensive care unit
di: Tailleur, Modan, et al.
Pubblicazione: (2024)
di: Tailleur, Modan, et al.
Pubblicazione: (2024)
Whisfusion: Parallel ASR Decoding via a Diffusion Transformer
di: Kwon, Taeyoun, et al.
Pubblicazione: (2025)
di: Kwon, Taeyoun, et al.
Pubblicazione: (2025)
Robust Cross-Etiology and Speaker-Independent Dysarthric Speech Recognition
di: Singh, Satwinder, et al.
Pubblicazione: (2025)
di: Singh, Satwinder, et al.
Pubblicazione: (2025)
Personalized Speech Enhancement Without a Separate Speaker Embedding Model
di: Pärnamaa, Tanel, et al.
Pubblicazione: (2024)
di: Pärnamaa, Tanel, et al.
Pubblicazione: (2024)
Masked Audio Generation using a Single Non-Autoregressive Transformer
di: Ziv, Alon, et al.
Pubblicazione: (2024)
di: Ziv, Alon, et al.
Pubblicazione: (2024)
Music Plagiarism Detection: Problem Formulation and a Segment-based Solution
di: Go, Seonghyeon, et al.
Pubblicazione: (2026)
di: Go, Seonghyeon, et al.
Pubblicazione: (2026)
Prompt-Singer: Controllable Singing-Voice-Synthesis with Natural Language Prompt
di: Wang, Yongqi, et al.
Pubblicazione: (2024)
di: Wang, Yongqi, et al.
Pubblicazione: (2024)
aTENNuate: Optimized Real-time Speech Enhancement with Deep SSMs on Raw Audio
di: Pei, Yan Ru, et al.
Pubblicazione: (2024)
di: Pei, Yan Ru, et al.
Pubblicazione: (2024)
JEPA as a Neural Tokenizer: Learning Robust Speech Representations with Density Adaptive Attention
di: Ioannides, Georgios, et al.
Pubblicazione: (2025)
di: Ioannides, Georgios, et al.
Pubblicazione: (2025)
DiceHuBERT: Distilling HuBERT with a Self-Supervised Learning Objective
di: Chi, Hyung Gun, et al.
Pubblicazione: (2025)
di: Chi, Hyung Gun, et al.
Pubblicazione: (2025)
MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
di: Wang, Yuancheng, et al.
Pubblicazione: (2024)
di: Wang, Yuancheng, et al.
Pubblicazione: (2024)
Enhancement of a Text-Independent Speaker Verification System by using Feature Combination and Parallel-Structure Classifiers
di: Abdalmalak, Kerlos Atia, et al.
Pubblicazione: (2024)
di: Abdalmalak, Kerlos Atia, et al.
Pubblicazione: (2024)
Auto-Regressive vs Flow-Matching: a Comparative Study of Modeling Paradigms for Text-to-Music Generation
di: Tal, Or, et al.
Pubblicazione: (2025)
di: Tal, Or, et al.
Pubblicazione: (2025)
Multimodal Audio-based Disease Prediction with Transformer-based Hierarchical Fusion Network
di: Cai, Jinjin, et al.
Pubblicazione: (2024)
di: Cai, Jinjin, et al.
Pubblicazione: (2024)
SAGE-Music: Low-Latency Symbolic Music Generation via Attribute-Specialized Key-Value Head Sharing
di: Tan, Jiaye, et al.
Pubblicazione: (2025)
di: Tan, Jiaye, et al.
Pubblicazione: (2025)
Speech Emotion Recognition Using CNN and Its Use Case in Digital Healthcare
di: Nigar, Nishargo
Pubblicazione: (2024)
di: Nigar, Nishargo
Pubblicazione: (2024)
DisMix: Disentangling Mixtures of Musical Instruments for Source-level Pitch and Timbre Manipulation
di: Luo, Yin-Jyun, et al.
Pubblicazione: (2024)
di: Luo, Yin-Jyun, et al.
Pubblicazione: (2024)
Run-Time Adaptation of Neural Beamforming for Robust Speech Dereverberation and Denoising
di: Fujita, Yoto, et al.
Pubblicazione: (2024)
di: Fujita, Yoto, et al.
Pubblicazione: (2024)
Towards Open Respiratory Acoustic Foundation Models: Pretraining and Benchmarking
di: Zhang, Yuwei, et al.
Pubblicazione: (2024)
di: Zhang, Yuwei, et al.
Pubblicazione: (2024)
Exploring and Applying Audio-Based Sentiment Analysis in Music
di: Jhanji, Etash
Pubblicazione: (2024)
di: Jhanji, Etash
Pubblicazione: (2024)
Bridging Modalities: Knowledge Distillation and Masked Training for Translating Multi-Modal Emotion Recognition to Uni-Modal, Speech-Only Emotion Recognition
di: Muaz, Muhammad, et al.
Pubblicazione: (2024)
di: Muaz, Muhammad, et al.
Pubblicazione: (2024)
Documenti analoghi
-
ASM: Audio Spectrogram Mixer
di: Ji, Qingfeng, et al.
Pubblicazione: (2024) -
Sparse Autoencoders Make Audio Foundation Models more Explainable
di: Mariotte, Théo, et al.
Pubblicazione: (2025) -
Mitigating Unauthorized Speech Synthesis for Voice Protection
di: Zhang, Zhisheng, et al.
Pubblicazione: (2024) -
Let There Be Sound: Reconstructing High Quality Speech from Silent Videos
di: Kim, Ji-Hoon, et al.
Pubblicazione: (2023) -
Symbotunes: unified hub for symbolic music generative models
di: Skierś, Paweł, et al.
Pubblicazione: (2024)