EnCLAP++: Analyzing the EnCLAP Framework for Optimizing Automated Audio Captioning Performance
Fuente:
arXiv
Guardado en:
| Autores principales: | Kim, Jaeyeon, Jeon, Minjeon, Jung, Jaeyoon, Woo, Sang Hoon, Lee, Jinjoo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Expanding on EnCLAP with Auxiliary Retrieval Model for Automated Audio Captioning
por: Kim, Jaeyeon, et al.
Publicado: (2024)
por: Kim, Jaeyeon, et al.
Publicado: (2024)
EnCLAP: Combining Neural Audio Codec and Audio-Text Joint Embedding for Automated Audio Captioning
por: Kim, Jaeyeon, et al.
Publicado: (2024)
por: Kim, Jaeyeon, et al.
Publicado: (2024)
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer
por: Takeuchi, Daiki, et al.
Publicado: (2025)
por: Takeuchi, Daiki, et al.
Publicado: (2025)
SLAM-AAC: Enhancing Audio Captioning with Paraphrasing Augmentation and CLAP-Refine through LLMs
por: Chen, Wenxi, et al.
Publicado: (2024)
por: Chen, Wenxi, et al.
Publicado: (2024)
Masked Audio Modeling with CLAP and Multi-Objective Learning
por: Xin, Yifei, et al.
Publicado: (2024)
por: Xin, Yifei, et al.
Publicado: (2024)
DRCap: Decoding CLAP Latents with Retrieval-Augmented Generation for Zero-shot Audio Captioning
por: Li, Xiquan, et al.
Publicado: (2024)
por: Li, Xiquan, et al.
Publicado: (2024)
Text2FX: Harnessing CLAP Embeddings for Text-Guided Audio Effects
por: Chu, Annie, et al.
Publicado: (2024)
por: Chu, Annie, et al.
Publicado: (2024)
SmoothCLAP: Soft-Target Enhanced Contrastive Language\--Audio Pretraining for Affective Computing
por: Jing, Xin, et al.
Publicado: (2026)
por: Jing, Xin, et al.
Publicado: (2026)
The TMU System for the XACLE Challenge: Training Large Audio Language Models with CLAP Pseudo-Labels
por: Tsutsumi, Ayuto, et al.
Publicado: (2026)
por: Tsutsumi, Ayuto, et al.
Publicado: (2026)
M2D-CLAP: Masked Modeling Duo Meets CLAP for Learning General-purpose Audio-Language Representation
por: Niizumi, Daisuke, et al.
Publicado: (2024)
por: Niizumi, Daisuke, et al.
Publicado: (2024)
RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval
por: Sun, Haoqin, et al.
Publicado: (2025)
por: Sun, Haoqin, et al.
Publicado: (2025)
ReCLAP: Improving Zero Shot Audio Classification by Describing Sounds
por: Ghosh, Sreyan, et al.
Publicado: (2024)
por: Ghosh, Sreyan, et al.
Publicado: (2024)
CLAP-Based Automatic Word Naming Recognition in Post-Stroke Aphasia
por: Kaloga, Yacouba, et al.
Publicado: (2026)
por: Kaloga, Yacouba, et al.
Publicado: (2026)
Human-CLAP: Human-perception-based contrastive language-audio pretraining
por: Takano, Taisei, et al.
Publicado: (2025)
por: Takano, Taisei, et al.
Publicado: (2025)
T-CLAP: Temporal-Enhanced Contrastive Language-Audio Pretraining
por: Yuan, Yi, et al.
Publicado: (2024)
por: Yuan, Yi, et al.
Publicado: (2024)
SPO-CLAPScore: Enhancing CLAP-based alignment prediction system with Standardize Preference Optimization, for the first XACLE Challenge
por: Takano, Taisei, et al.
Publicado: (2026)
por: Takano, Taisei, et al.
Publicado: (2026)
ParaCLAP -- Towards a general language-audio model for computational paralinguistic tasks
por: Jing, Xin, et al.
Publicado: (2024)
por: Jing, Xin, et al.
Publicado: (2024)
tinyCLAP: Distilling Constrastive Language-Audio Pretrained Models
por: Paissan, Francesco, et al.
Publicado: (2023)
por: Paissan, Francesco, et al.
Publicado: (2023)
WoW-Bench: Evaluating Fine-Grained Acoustic Perception in Audio-Language Models via Marine Mammal Vocalizations
por: Kim, Jaeyeon, et al.
Publicado: (2025)
por: Kim, Jaeyeon, et al.
Publicado: (2025)
Discrete Audio Representations for Automated Audio Captioning
por: Tian, Jingguang, et al.
Publicado: (2025)
por: Tian, Jingguang, et al.
Publicado: (2025)
Zero-Shot Crate Digging: DJ Tool Retrieval Using Speech Activity, Music Structure And CLAP Embeddings
por: Orife, Iroro
Publicado: (2024)
por: Orife, Iroro
Publicado: (2024)
Diverse Audio Embeddings -- Bringing Features Back Outperforms CLAP!
por: Verma, Prateek
Publicado: (2023)
por: Verma, Prateek
Publicado: (2023)
DeFT-Mamba: Universal Multichannel Sound Separation and Polyphonic Audio Classification
por: Lee, Dongheon, et al.
Publicado: (2024)
por: Lee, Dongheon, et al.
Publicado: (2024)
ACES: Evaluating Automated Audio Captioning Models on the Semantics of Sounds
por: Wijngaard, Gijs, et al.
Publicado: (2024)
por: Wijngaard, Gijs, et al.
Publicado: (2024)
FlowAVSE: Efficient Audio-Visual Speech Enhancement with Conditional Flow Matching
por: Jung, Chaeyoung, et al.
Publicado: (2024)
por: Jung, Chaeyoung, et al.
Publicado: (2024)
Learning Semantic Information from Raw Audio Signal Using Both Contextual and Phonetic Representations
por: Kim, Jaeyeon, et al.
Publicado: (2024)
por: Kim, Jaeyeon, et al.
Publicado: (2024)
ParaSpeechCLAP: A Dual-Encoder Speech-Text Model for Rich Stylistic Language-Audio Pretraining
por: Diwan, Anuj, et al.
Publicado: (2026)
por: Diwan, Anuj, et al.
Publicado: (2026)
ViSAGe: Video-to-Spatial Audio Generation
por: Kim, Jaeyeon, et al.
Publicado: (2025)
por: Kim, Jaeyeon, et al.
Publicado: (2025)
MACE: Leveraging Audio for Evaluating Audio Captioning Systems
por: Dixit, Satvik, et al.
Publicado: (2024)
por: Dixit, Satvik, et al.
Publicado: (2024)
Enhancing Automated Audio Captioning via Large Language Models with Optimized Audio Encoding
por: Liu, Jizhong, et al.
Publicado: (2024)
por: Liu, Jizhong, et al.
Publicado: (2024)
SEMamba++: A General Speech Restoration Framework Leveraging Global, Local, and Periodic Spectral Patterns
por: Lee, Yongjoon, et al.
Publicado: (2026)
por: Lee, Yongjoon, et al.
Publicado: (2026)
Audio-Based Linguistic Feature Extraction for Enhancing Multi-lingual and Low-Resource Text-to-Speech
por: Kim, Youngjae, et al.
Publicado: (2024)
por: Kim, Youngjae, et al.
Publicado: (2024)
Improving Audio Captioning Models with Fine-grained Audio Features, Text Embedding Supervision, and LLM Mix-up Augmentation
por: Wu, Shih-Lun, et al.
Publicado: (2023)
por: Wu, Shih-Lun, et al.
Publicado: (2023)
MiDashengLM: Efficient Audio Understanding with General Audio Captions
por: Dinkel, Heinrich, et al.
Publicado: (2025)
por: Dinkel, Heinrich, et al.
Publicado: (2025)
CosyAudio: Improving Audio Generation with Confidence Scores and Synthetic Captions
por: Zhu, Xinfa, et al.
Publicado: (2025)
por: Zhu, Xinfa, et al.
Publicado: (2025)
Adversarial Multi-Task Learning for Disentangling Timbre and Pitch in Singing Voice Synthesis
por: Kim, Tae-Woo, et al.
Publicado: (2022)
por: Kim, Tae-Woo, et al.
Publicado: (2022)
Efficient Audio Captioning with Encoder-Level Knowledge Distillation
por: Xu, Xuenan, et al.
Publicado: (2024)
por: Xu, Xuenan, et al.
Publicado: (2024)
Inter-channel Conv-TasNet for multichannel speech enhancement
por: Lee, Dongheon, et al.
Publicado: (2021)
por: Lee, Dongheon, et al.
Publicado: (2021)
Zero-Shot Audio Captioning Using Soft and Hard Prompts
por: Zhang, Yiming, et al.
Publicado: (2024)
por: Zhang, Yiming, et al.
Publicado: (2024)
Enhance Temporal Relations in Audio Captioning with Sound Event Detection
por: Xie, Zeyu, et al.
Publicado: (2023)
por: Xie, Zeyu, et al.
Publicado: (2023)
Ejemplares similares
-
Expanding on EnCLAP with Auxiliary Retrieval Model for Automated Audio Captioning
por: Kim, Jaeyeon, et al.
Publicado: (2024) -
EnCLAP: Combining Neural Audio Codec and Audio-Text Joint Embedding for Automated Audio Captioning
por: Kim, Jaeyeon, et al.
Publicado: (2024) -
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer
por: Takeuchi, Daiki, et al.
Publicado: (2025) -
SLAM-AAC: Enhancing Audio Captioning with Paraphrasing Augmentation and CLAP-Refine through LLMs
por: Chen, Wenxi, et al.
Publicado: (2024) -
Masked Audio Modeling with CLAP and Multi-Objective Learning
por: Xin, Yifei, et al.
Publicado: (2024)