Episodic fine-tuning prototypical networks for optimization-based few-shot learning: Application to audio classification
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhuang, Xuanyu, Peeters, Geoffroy, Richard, Gaël |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The Inverse Drum Machine: Source Separation Through Joint Transcription and Analysis-by-Synthesis
di: Torres, Bernardo, et al.
Pubblicazione: (2025)
di: Torres, Bernardo, et al.
Pubblicazione: (2025)
Unsupervised Harmonic Parameter Estimation Using Differentiable DSP and Spectral Optimal Transport
di: Torres, Bernardo, et al.
Pubblicazione: (2023)
di: Torres, Bernardo, et al.
Pubblicazione: (2023)
FoVNet: Configurable Field-of-View Speech Enhancement with Low Computation and Distortion for Smart Glasses
di: Xu, Zhongweiyang, et al.
Pubblicazione: (2024)
di: Xu, Zhongweiyang, et al.
Pubblicazione: (2024)
Stimulus Modality Matters: Impact of Perceptual Evaluations from Different Modalities on Speech Emotion Recognition System Performance
di: Chou, Huang-Cheng, et al.
Pubblicazione: (2024)
di: Chou, Huang-Cheng, et al.
Pubblicazione: (2024)
A Contrastive Self-Supervised Learning scheme for beat tracking amenable to few-shot learning
di: Gagnere, Antonin, et al.
Pubblicazione: (2024)
di: Gagnere, Antonin, et al.
Pubblicazione: (2024)
Visual-based spatial audio generation system for multi-speaker environments
di: Liu, Xiaojing, et al.
Pubblicazione: (2025)
di: Liu, Xiaojing, et al.
Pubblicazione: (2025)
An automatic mixing speech enhancement system for multi-track audio
di: Liu, Xiaojing, et al.
Pubblicazione: (2024)
di: Liu, Xiaojing, et al.
Pubblicazione: (2024)
LRS-VoxMM: A benchmark for in-the-wild audio-visual speech recognition
di: Kwak, Doyeop, et al.
Pubblicazione: (2026)
di: Kwak, Doyeop, et al.
Pubblicazione: (2026)
Versatile audio-visual learning for emotion recognition
di: Goncalves, Lucas, et al.
Pubblicazione: (2023)
di: Goncalves, Lucas, et al.
Pubblicazione: (2023)
Improved symbolic drum style classification with grammar-based hierarchical representations
di: Géré, Léo, et al.
Pubblicazione: (2024)
di: Géré, Léo, et al.
Pubblicazione: (2024)
Episode-specific Fine-tuning for Metric-based Few-shot Learners with Optimization-based Training
di: Zhuang, Xuanyu, et al.
Pubblicazione: (2025)
di: Zhuang, Xuanyu, et al.
Pubblicazione: (2025)
Blind estimation of audio effects using an auto-encoder approach and differentiable digital signal processing
di: Peladeau, Côme, et al.
Pubblicazione: (2023)
di: Peladeau, Côme, et al.
Pubblicazione: (2023)
Synthetic training set generation using text-to-audio models for environmental sound classification
di: Ronchini, Francesca, et al.
Pubblicazione: (2024)
di: Ronchini, Francesca, et al.
Pubblicazione: (2024)
Learning Temporal Resolution in Spectrogram for Audio Classification
di: Liu, Haohe, et al.
Pubblicazione: (2022)
di: Liu, Haohe, et al.
Pubblicazione: (2022)
AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining
di: Liu, Haohe, et al.
Pubblicazione: (2023)
di: Liu, Haohe, et al.
Pubblicazione: (2023)
SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound
di: Liu, Haohe, et al.
Pubblicazione: (2024)
di: Liu, Haohe, et al.
Pubblicazione: (2024)
SynthTab: Leveraging Synthesized Data for Guitar Tablature Transcription
di: Zang, Yongyi, et al.
Pubblicazione: (2023)
di: Zang, Yongyi, et al.
Pubblicazione: (2023)
ArrayDPS: Unsupervised Blind Speech Separation with a Diffusion Prior
di: Xu, Zhongweiyang, et al.
Pubblicazione: (2025)
di: Xu, Zhongweiyang, et al.
Pubblicazione: (2025)
Recent Advances in Discrete Speech Tokens: A Review
di: Guo, Yiwei, et al.
Pubblicazione: (2025)
di: Guo, Yiwei, et al.
Pubblicazione: (2025)
Analyzable Chain-of-Musical-Thought Prompting for High-Fidelity Music Generation
di: Lam, Max W. Y., et al.
Pubblicazione: (2025)
di: Lam, Max W. Y., et al.
Pubblicazione: (2025)
WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling
di: Ji, Shengpeng, et al.
Pubblicazione: (2024)
di: Ji, Shengpeng, et al.
Pubblicazione: (2024)
A multi-modal approach for identifying schizophrenia using cross-modal attention
di: Premananth, Gowtham, et al.
Pubblicazione: (2023)
di: Premananth, Gowtham, et al.
Pubblicazione: (2023)
SteerMusic: Enhanced Musical Consistency for Zero-shot Text-guided and Personalized Music Editing
di: Niu, Xinlei, et al.
Pubblicazione: (2025)
di: Niu, Xinlei, et al.
Pubblicazione: (2025)
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text
di: Liu, Haohe, et al.
Pubblicazione: (2024)
di: Liu, Haohe, et al.
Pubblicazione: (2024)
Deep learning classification system for coconut maturity levels based on acoustic signals
di: Caladcad, June Anne, et al.
Pubblicazione: (2024)
di: Caladcad, June Anne, et al.
Pubblicazione: (2024)
Enhancing Expressiveness in Dance Generation via Integrating Frequency and Music Style Information
di: Huang, Qiaochu, et al.
Pubblicazione: (2024)
di: Huang, Qiaochu, et al.
Pubblicazione: (2024)
Conformer-based Ultrasound-to-Speech Conversion
di: Ibrahimov, Ibrahim, et al.
Pubblicazione: (2025)
di: Ibrahimov, Ibrahim, et al.
Pubblicazione: (2025)
Speech dereverberation constrained on room impulse response characteristics
di: Bahrman, Louis, et al.
Pubblicazione: (2024)
di: Bahrman, Louis, et al.
Pubblicazione: (2024)
Dance-to-Music Generation with Encoder-based Textual Inversion
di: Li, Sifei, et al.
Pubblicazione: (2024)
di: Li, Sifei, et al.
Pubblicazione: (2024)
Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions
di: Yuan, Yi, et al.
Pubblicazione: (2024)
di: Yuan, Yi, et al.
Pubblicazione: (2024)
Exploring compressibility of transformer based text-to-music (TTM) models
di: Moschopoulos, Vasileios, et al.
Pubblicazione: (2024)
di: Moschopoulos, Vasileios, et al.
Pubblicazione: (2024)
Attentive-based Multi-level Feature Fusion for Voice Disorder Diagnosis
di: Shen, Lipeng, et al.
Pubblicazione: (2024)
di: Shen, Lipeng, et al.
Pubblicazione: (2024)
REWIND: Speech Time Reversal for Enhancing Speaker Representations in Diffusion-based Voice Conversion
di: Biyani, Ishan D., et al.
Pubblicazione: (2025)
di: Biyani, Ishan D., et al.
Pubblicazione: (2025)
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding
di: Wang, Ziqian, et al.
Pubblicazione: (2025)
di: Wang, Ziqian, et al.
Pubblicazione: (2025)
Deep learning-based filtering of cross-spectral matrices using generative adversarial networks
di: Puhle, Christof
Pubblicazione: (2025)
di: Puhle, Christof
Pubblicazione: (2025)
Robust LLM-based Audio-Visual Speech Recognition with Sparse Modality Alignment and Visual Unit-Guided Refinement
di: Su, Fei, et al.
Pubblicazione: (2026)
di: Su, Fei, et al.
Pubblicazione: (2026)
Reverse the auditory processing pathway: Coarse-to-fine audio reconstruction from fMRI
di: Liu, Che, et al.
Pubblicazione: (2024)
di: Liu, Che, et al.
Pubblicazione: (2024)
Why some audio signal short-time Fourier transform coefficients have nonuniform phase distributions
di: Voran, Stephen D.
Pubblicazione: (2024)
di: Voran, Stephen D.
Pubblicazione: (2024)
Visual and audio scene classification for detecting discrepancies in video: a baseline method and experimental protocol
di: Apostolidis, Konstantinos, et al.
Pubblicazione: (2024)
di: Apostolidis, Konstantinos, et al.
Pubblicazione: (2024)
Audio-Visual Speaker Tracking: Progress, Challenges, and Future Directions
di: Zhao, Jinzheng, et al.
Pubblicazione: (2023)
di: Zhao, Jinzheng, et al.
Pubblicazione: (2023)
Documenti analoghi
-
The Inverse Drum Machine: Source Separation Through Joint Transcription and Analysis-by-Synthesis
di: Torres, Bernardo, et al.
Pubblicazione: (2025) -
Unsupervised Harmonic Parameter Estimation Using Differentiable DSP and Spectral Optimal Transport
di: Torres, Bernardo, et al.
Pubblicazione: (2023) -
FoVNet: Configurable Field-of-View Speech Enhancement with Low Computation and Distortion for Smart Glasses
di: Xu, Zhongweiyang, et al.
Pubblicazione: (2024) -
Stimulus Modality Matters: Impact of Perceptual Evaluations from Different Modalities on Speech Emotion Recognition System Performance
di: Chou, Huang-Cheng, et al.
Pubblicazione: (2024) -
A Contrastive Self-Supervised Learning scheme for beat tracking amenable to few-shot learning
di: Gagnere, Antonin, et al.
Pubblicazione: (2024)