EH-MAM: Easy-to-Hard Masked Acoustic Modeling for Self-Supervised Speech Representation Learning
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Seth, Ashish, Selvakumar, Ramaneswaran, Sakshi, S, Kumar, Sonal, Ghosh, Sreyan, Manocha, Dinesh |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
PAT: Parameter-Free Audio-Text Aligner to Boost Zero-Shot Audio Classification
par: Seth, Ashish, et autres
Publié: (2024)
par: Seth, Ashish, et autres
Publié: (2024)
Do Audio-Language Models Understand Linguistic Variations?
par: Selvakumar, Ramaneswaran, et autres
Publié: (2024)
par: Selvakumar, Ramaneswaran, et autres
Publié: (2024)
MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
par: Sakshi, S, et autres
Publié: (2024)
par: Sakshi, S, et autres
Publié: (2024)
CompA: Addressing the Gap in Compositional Reasoning in Audio-Language Models
par: Ghosh, Sreyan, et autres
Publié: (2023)
par: Ghosh, Sreyan, et autres
Publié: (2023)
GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities
par: Ghosh, Sreyan, et autres
Publié: (2024)
par: Ghosh, Sreyan, et autres
Publié: (2024)
ReCLAP: Improving Zero Shot Audio Classification by Describing Sounds
par: Ghosh, Sreyan, et autres
Publié: (2024)
par: Ghosh, Sreyan, et autres
Publié: (2024)
RECAP: Retrieval-Augmented Audio Captioning
par: Ghosh, Sreyan, et autres
Publié: (2023)
par: Ghosh, Sreyan, et autres
Publié: (2023)
Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities
par: Ghosh, Sreyan, et autres
Publié: (2025)
par: Ghosh, Sreyan, et autres
Publié: (2025)
SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation
par: Kumar, Sonal, et autres
Publié: (2024)
par: Kumar, Sonal, et autres
Publié: (2024)
LipGER: Visually-Conditioned Generative Error Correction for Robust Automatic Speech Recognition
par: Ghosh, Sreyan, et autres
Publié: (2024)
par: Ghosh, Sreyan, et autres
Publié: (2024)
Failing Forward: Improving Generative Error Correction for ASR with Synthetic Data and Retrieval Augmentation
par: Ghosh, Sreyan, et autres
Publié: (2024)
par: Ghosh, Sreyan, et autres
Publié: (2024)
VQTTS: High-Fidelity Text-to-Speech Synthesis with Self-Supervised VQ Acoustic Feature
par: Du, Chenpeng, et autres
Publié: (2022)
par: Du, Chenpeng, et autres
Publié: (2022)
Refining Self-Supervised Learnt Speech Representation using Brain Activations
par: Li, Hengyu, et autres
Publié: (2024)
par: Li, Hengyu, et autres
Publié: (2024)
Emotion-Aware Speech Self-Supervised Representation Learning with Intensity Knowledge
par: Liu, Rui, et autres
Publié: (2024)
par: Liu, Rui, et autres
Publié: (2024)
ProSE: Diffusion Priors for Speech Enhancement
par: Kumar, Sonal, et autres
Publié: (2025)
par: Kumar, Sonal, et autres
Publié: (2025)
Unsupervised Accent Adaptation Through Masked Language Model Correction Of Discrete Self-Supervised Speech Units
par: Poncelet, Jakob, et autres
Publié: (2023)
par: Poncelet, Jakob, et autres
Publié: (2023)
Multi-resolution HuBERT: Multi-resolution Speech Self-Supervised Learning with Masked Unit Prediction
par: Shi, Jiatong, et autres
Publié: (2023)
par: Shi, Jiatong, et autres
Publié: (2023)
Distillation and Pruning for Scalable Self-Supervised Representation-Based Speech Quality Assessment
par: Stahl, Benjamin, et autres
Publié: (2025)
par: Stahl, Benjamin, et autres
Publié: (2025)
Multi-Domain Audio Question Answering Benchmark Toward Acoustic Content Reasoning
par: Yang, Chao-Han Huck, et autres
Publié: (2025)
par: Yang, Chao-Han Huck, et autres
Publié: (2025)
Self-Supervised Learning of Spatial Acoustic Representation with Cross-Channel Signal Reconstruction and Multi-Channel Conformer
par: Yang, Bing, et autres
Publié: (2023)
par: Yang, Bing, et autres
Publié: (2023)
A Large-Scale Probing Analysis of Speaker-Specific Attributes in Self-Supervised Speech Representations
par: Chiu, Aemon Yat Fei, et autres
Publié: (2025)
par: Chiu, Aemon Yat Fei, et autres
Publié: (2025)
Generic Speech Enhancement with Self-Supervised Representation Space Loss
par: Sato, Hiroshi, et autres
Publié: (2025)
par: Sato, Hiroshi, et autres
Publié: (2025)
Analysis of Self-Supervised Speech Models on Children's Speech and Infant Vocalizations
par: Li, Jialu, et autres
Publié: (2024)
par: Li, Jialu, et autres
Publié: (2024)
Listen through the Sound: Generative Speech Restoration Leveraging Acoustic Context Representation
par: Chung, Soo-Whan, et autres
Publié: (2025)
par: Chung, Soo-Whan, et autres
Publié: (2025)
TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification
par: Anand, Nishit, et autres
Publié: (2024)
par: Anand, Nishit, et autres
Publié: (2024)
Textless Acoustic Model with Self-Supervised Distillation for Noise-Robust Expressive Speech-to-Speech Translation
par: Hwang, Min-Jae, et autres
Publié: (2024)
par: Hwang, Min-Jae, et autres
Publié: (2024)
Generative Data Augmentation Challenge: Synthesis of Room Acoustics for Speaker Distance Estimation
par: Lin, Jackie, et autres
Publié: (2025)
par: Lin, Jackie, et autres
Publié: (2025)
Leveraging Self-Supervised Models for Automatic Whispered Speech Recognition
par: Farhadipour, Aref, et autres
Publié: (2024)
par: Farhadipour, Aref, et autres
Publié: (2024)
GigaAM: Efficient Self-Supervised Learner for Speech Recognition
par: Kutsakov, Aleksandr, et autres
Publié: (2025)
par: Kutsakov, Aleksandr, et autres
Publié: (2025)
Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models
par: Goel, Arushi, et autres
Publié: (2025)
par: Goel, Arushi, et autres
Publié: (2025)
Analysing the Masked predictive coding training criterion for pre-training a Speech Representation Model
par: Yadav, Hemant, et autres
Publié: (2023)
par: Yadav, Hemant, et autres
Publié: (2023)
Leveraging Self-supervised Audio Representations for Data-Efficient Acoustic Scene Classification
par: Cai, Yiqiang, et autres
Publié: (2024)
par: Cai, Yiqiang, et autres
Publié: (2024)
Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition
par: Wang, Shih-heng, et autres
Publié: (2024)
par: Wang, Shih-heng, et autres
Publié: (2024)
Comparison of Self-Supervised Speech Pre-Training Methods on Flemish Dutch
par: Poncelet, Jakob, et autres
Publié: (2021)
par: Poncelet, Jakob, et autres
Publié: (2021)
Towards Automatic Assessment of Self-Supervised Speech Models using Rank
par: Aldeneh, Zakaria, et autres
Publié: (2024)
par: Aldeneh, Zakaria, et autres
Publié: (2024)
Evaluating Self-Supervised Speech Models via Text-Based LLMS
par: Maekaku, Takashi, et autres
Publié: (2025)
par: Maekaku, Takashi, et autres
Publié: (2025)
Self-Supervised Speech Quality Assessment (S3QA): Leveraging Speech Foundation Models for a Scalable Speech Quality Metric
par: Ogg, Mattson, et autres
Publié: (2025)
par: Ogg, Mattson, et autres
Publié: (2025)
Acoustic BPE for Speech Generation with Discrete Tokens
par: Shen, Feiyu, et autres
Publié: (2023)
par: Shen, Feiyu, et autres
Publié: (2023)
Self-Supervised Singing Voice Pre-Training towards Speech-to-Singing Conversion
par: Li, Ruiqi, et autres
Publié: (2024)
par: Li, Ruiqi, et autres
Publié: (2024)
Edit Content, Preserve Acoustics: Imperceptible Text-Based Speech Editing via Self-Consistency Rewards
par: Ren, Yong, et autres
Publié: (2026)
par: Ren, Yong, et autres
Publié: (2026)
Documents similaires
-
PAT: Parameter-Free Audio-Text Aligner to Boost Zero-Shot Audio Classification
par: Seth, Ashish, et autres
Publié: (2024) -
Do Audio-Language Models Understand Linguistic Variations?
par: Selvakumar, Ramaneswaran, et autres
Publié: (2024) -
MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
par: Sakshi, S, et autres
Publié: (2024) -
CompA: Addressing the Gap in Compositional Reasoning in Audio-Language Models
par: Ghosh, Sreyan, et autres
Publié: (2023) -
GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities
par: Ghosh, Sreyan, et autres
Publié: (2024)