The CMU-AIST submission for the ICME 2025 Audio Encoder Challenge

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bharadwaj, Shikhar, Cornell, Samuele, Choi, Kwanghee, Shim, Hye-jin, Deshmukh, Soham, Fukayama, Satoru, Watanabe, Shinji
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918301230170112
author Bharadwaj, Shikhar
Cornell, Samuele
Choi, Kwanghee
Shim, Hye-jin
Deshmukh, Soham
Fukayama, Satoru
Watanabe, Shinji
author_facet Bharadwaj, Shikhar
Cornell, Samuele
Choi, Kwanghee
Shim, Hye-jin
Deshmukh, Soham
Fukayama, Satoru
Watanabe, Shinji
contents This technical report describes our submission to the ICME 2025 audio encoder challenge. Our submitted system is built on BEATs, a masked speech token prediction based audio encoder. We extend the BEATs model using 74,000 hours of data derived from various speech, music, and sound corpora and scale its architecture upto 300 million parameters. We experiment with speech-heavy and balanced pre-training mixtures to study the impact of different domains on final performance. Our submitted system consists of an ensemble of the Dasheng 1.2 billion model with two custom scaled-up BEATs models trained on the aforementioned pre-training data mixtures. We also propose a simple ensembling technique that retains the best capabilities of constituent models and surpasses both the baseline and Dasheng 1.2B. For open science, we publicly release our trained checkpoints via huggingface at https://huggingface.co/shikhar7ssu/OpenBEATs-ICME-SOUND and https://huggingface.co/shikhar7ssu/OpenBEATs-ICME.
format Preprint
id arxiv_https___arxiv_org_abs_2601_16273
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle The CMU-AIST submission for the ICME 2025 Audio Encoder Challenge
Bharadwaj, Shikhar
Cornell, Samuele
Choi, Kwanghee
Shim, Hye-jin
Deshmukh, Soham
Fukayama, Satoru
Watanabe, Shinji
Sound
Audio and Speech Processing
This technical report describes our submission to the ICME 2025 audio encoder challenge. Our submitted system is built on BEATs, a masked speech token prediction based audio encoder. We extend the BEATs model using 74,000 hours of data derived from various speech, music, and sound corpora and scale its architecture upto 300 million parameters. We experiment with speech-heavy and balanced pre-training mixtures to study the impact of different domains on final performance. Our submitted system consists of an ensemble of the Dasheng 1.2 billion model with two custom scaled-up BEATs models trained on the aforementioned pre-training data mixtures. We also propose a simple ensembling technique that retains the best capabilities of constituent models and surpasses both the baseline and Dasheng 1.2B. For open science, we publicly release our trained checkpoints via huggingface at https://huggingface.co/shikhar7ssu/OpenBEATs-ICME-SOUND and https://huggingface.co/shikhar7ssu/OpenBEATs-ICME.
title The CMU-AIST submission for the ICME 2025 Audio Encoder Challenge
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2601.16273