The TMU System for the XACLE Challenge: Training Large Audio Language Models with CLAP Pseudo-Labels
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Tsutsumi, Ayuto, Tanaka, Kohei, Shiota, Sayaka |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
SPO-CLAPScore: Enhancing CLAP-based alignment prediction system with Standardize Preference Optimization, for the first XACLE Challenge
par: Takano, Taisei, et autres
Publié: (2026)
par: Takano, Taisei, et autres
Publié: (2026)
Masked Audio Modeling with CLAP and Multi-Objective Learning
par: Xin, Yifei, et autres
Publié: (2024)
par: Xin, Yifei, et autres
Publié: (2024)
SmoothCLAP: Soft-Target Enhanced Contrastive Language\--Audio Pretraining for Affective Computing
par: Jing, Xin, et autres
Publié: (2026)
par: Jing, Xin, et autres
Publié: (2026)
RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval
par: Sun, Haoqin, et autres
Publié: (2025)
par: Sun, Haoqin, et autres
Publié: (2025)
Text2FX: Harnessing CLAP Embeddings for Text-Guided Audio Effects
par: Chu, Annie, et autres
Publié: (2024)
par: Chu, Annie, et autres
Publié: (2024)
The Interspeech 2026 Audio Encoder Capability Challenge for Large Audio Language Models
par: Dinkel, Heinrich, et autres
Publié: (2026)
par: Dinkel, Heinrich, et autres
Publié: (2026)
YODAS: Youtube-Oriented Dataset for Audio and Speech
par: Li, Xinjian, et autres
Publié: (2024)
par: Li, Xinjian, et autres
Publié: (2024)
Few-Shot and Pseudo-Label Guided Speech Quality Evaluation with Large Language Models
par: Zezario, Ryandhimas E., et autres
Publié: (2026)
par: Zezario, Ryandhimas E., et autres
Publié: (2026)
SLAM-AAC: Enhancing Audio Captioning with Paraphrasing Augmentation and CLAP-Refine through LLMs
par: Chen, Wenxi, et autres
Publié: (2024)
par: Chen, Wenxi, et autres
Publié: (2024)
Speech privacy-preserving methods using secret key for convolutional neural network models and their robustness evaluation
par: Niwa, Shoko, et autres
Publié: (2024)
par: Niwa, Shoko, et autres
Publié: (2024)
EnCLAP++: Analyzing the EnCLAP Framework for Optimizing Automated Audio Captioning Performance
par: Kim, Jaeyeon, et autres
Publié: (2024)
par: Kim, Jaeyeon, et autres
Publié: (2024)
Voice Privacy Preservation with Multiple Random Orthogonal Secret Keys: Attack Resistance Analysis
par: Tanaka, Kohei, et autres
Publié: (2025)
par: Tanaka, Kohei, et autres
Publié: (2025)
M2D-CLAP: Masked Modeling Duo Meets CLAP for Learning General-purpose Audio-Language Representation
par: Niizumi, Daisuke, et autres
Publié: (2024)
par: Niizumi, Daisuke, et autres
Publié: (2024)
Expanding on EnCLAP with Auxiliary Retrieval Model for Automated Audio Captioning
par: Kim, Jaeyeon, et autres
Publié: (2024)
par: Kim, Jaeyeon, et autres
Publié: (2024)
Pseudo Labels-based Neural Speech Enhancement for the AVSR Task in the MISP-Meeting Challenge
par: Luo, Longjie, et autres
Publié: (2025)
par: Luo, Longjie, et autres
Publié: (2025)
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer
par: Takeuchi, Daiki, et autres
Publié: (2025)
par: Takeuchi, Daiki, et autres
Publié: (2025)
tinyCLAP: Distilling Constrastive Language-Audio Pretrained Models
par: Paissan, Francesco, et autres
Publié: (2023)
par: Paissan, Francesco, et autres
Publié: (2023)
EnCLAP: Combining Neural Audio Codec and Audio-Text Joint Embedding for Automated Audio Captioning
par: Kim, Jaeyeon, et autres
Publié: (2024)
par: Kim, Jaeyeon, et autres
Publié: (2024)
Unlocking Large Audio-Language Models for Interactive Language Learning
par: Liu, Hongfu, et autres
Publié: (2026)
par: Liu, Hongfu, et autres
Publié: (2026)
ReCLAP: Improving Zero Shot Audio Classification by Describing Sounds
par: Ghosh, Sreyan, et autres
Publié: (2024)
par: Ghosh, Sreyan, et autres
Publié: (2024)
T-CLAP: Temporal-Enhanced Contrastive Language-Audio Pretraining
par: Yuan, Yi, et autres
Publié: (2024)
par: Yuan, Yi, et autres
Publié: (2024)
Can Large Language Models Understand Spatial Audio?
par: Tang, Changli, et autres
Publié: (2024)
par: Tang, Changli, et autres
Publié: (2024)
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training
par: Udupa, Sathvik, et autres
Publié: (2025)
par: Udupa, Sathvik, et autres
Publié: (2025)
The T12 System for AudioMOS Challenge 2025: Audio Aesthetics Score Prediction System Using KAN- and VERSA-based Models
par: Yamamoto, Katsuhiko, et autres
Publié: (2025)
par: Yamamoto, Katsuhiko, et autres
Publié: (2025)
Audio Array-Based 3D UAV Trajectory Estimation with LiDAR Pseudo-Labeling
par: Lei, Allen, et autres
Publié: (2024)
par: Lei, Allen, et autres
Publié: (2024)
CLAP-Based Automatic Word Naming Recognition in Post-Stroke Aphasia
par: Kaloga, Yacouba, et autres
Publié: (2026)
par: Kaloga, Yacouba, et autres
Publié: (2026)
Human-CLAP: Human-perception-based contrastive language-audio pretraining
par: Takano, Taisei, et autres
Publié: (2025)
par: Takano, Taisei, et autres
Publié: (2025)
Can Audio Large Language Models Verify Speaker Identity?
par: Ren, Yiming, et autres
Publié: (2025)
par: Ren, Yiming, et autres
Publié: (2025)
ProLAP: Probabilistic Language-Audio Pre-Training
par: Manabe, Toranosuke, et autres
Publié: (2025)
par: Manabe, Toranosuke, et autres
Publié: (2025)
UniAudio 1.5: Large Language Model-driven Audio Codec is A Few-shot Audio Task Learner
par: Yang, Dongchao, et autres
Publié: (2024)
par: Yang, Dongchao, et autres
Publié: (2024)
Pengi: An Audio Language Model for Audio Tasks
par: Deshmukh, Soham, et autres
Publié: (2023)
par: Deshmukh, Soham, et autres
Publié: (2023)
LongCat-Audio-Codec: An Audio Tokenizer and Detokenizer Solution Designed for Speech Large Language Models
par: Zhao, Xiaohan, et autres
Publié: (2025)
par: Zhao, Xiaohan, et autres
Publié: (2025)
DRCap: Decoding CLAP Latents with Retrieval-Augmented Generation for Zero-shot Audio Captioning
par: Li, Xiquan, et autres
Publié: (2024)
par: Li, Xiquan, et autres
Publié: (2024)
SALT: Standardized Audio event Label Taxonomy
par: Stamatiadis, Paraskevas, et autres
Publié: (2024)
par: Stamatiadis, Paraskevas, et autres
Publié: (2024)
Continuous Audio Language Models
par: Rouard, Simon, et autres
Publié: (2025)
par: Rouard, Simon, et autres
Publié: (2025)
Channel Adaptation for Speaker Verification Using Optimal Transport with Pseudo Label
par: Yang, Wenhao, et autres
Publié: (2024)
par: Yang, Wenhao, et autres
Publié: (2024)
ParaCLAP -- Towards a general language-audio model for computational paralinguistic tasks
par: Jing, Xin, et autres
Publié: (2024)
par: Jing, Xin, et autres
Publié: (2024)
MINT: Boosting Audio-Language Model via Multi-Target Pre-Training and Instruction Tuning
par: Zhao, Hang, et autres
Publié: (2024)
par: Zhao, Hang, et autres
Publié: (2024)
Enhancing Speech Large Language Models with Prompt-Aware Mixture of Audio Encoders
par: Shan, Weiqiao, et autres
Publié: (2025)
par: Shan, Weiqiao, et autres
Publié: (2025)
PAM: Prompting Audio-Language Models for Audio Quality Assessment
par: Deshmukh, Soham, et autres
Publié: (2024)
par: Deshmukh, Soham, et autres
Publié: (2024)
Documents similaires
-
SPO-CLAPScore: Enhancing CLAP-based alignment prediction system with Standardize Preference Optimization, for the first XACLE Challenge
par: Takano, Taisei, et autres
Publié: (2026) -
Masked Audio Modeling with CLAP and Multi-Objective Learning
par: Xin, Yifei, et autres
Publié: (2024) -
SmoothCLAP: Soft-Target Enhanced Contrastive Language\--Audio Pretraining for Affective Computing
par: Jing, Xin, et autres
Publié: (2026) -
RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval
par: Sun, Haoqin, et autres
Publié: (2025) -
Text2FX: Harnessing CLAP Embeddings for Text-Guided Audio Effects
par: Chu, Annie, et autres
Publié: (2024)