MERaLiON-AudioLLM: Bridging Audio and Language with Large Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: He, Yingxu, Liu, Zhuohan, Sun, Shuo, Wang, Bin, Zhang, Wenyu, Zou, Xunlong, Chen, Nancy F., Aw, Ai Ti
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917893495586816
author He, Yingxu
Liu, Zhuohan
Sun, Shuo
Wang, Bin
Zhang, Wenyu
Zou, Xunlong
Chen, Nancy F.
Aw, Ai Ti
author_facet He, Yingxu
Liu, Zhuohan
Sun, Shuo
Wang, Bin
Zhang, Wenyu
Zou, Xunlong
Chen, Nancy F.
Aw, Ai Ti
contents We introduce MERaLiON-AudioLLM (Multimodal Empathetic Reasoning and Learning in One Network), the first speech-text model tailored for Singapore's multilingual and multicultural landscape. Developed under the National Large Language Models Funding Initiative, Singapore, MERaLiON-AudioLLM integrates advanced speech and text processing to address the diverse linguistic nuances of local accents and dialects, enhancing accessibility and usability in complex, multilingual environments. Our results demonstrate improvements in both speech recognition and task-specific understanding, positioning MERaLiON-AudioLLM as a pioneering solution for region specific AI applications. We envision this release to set a precedent for future models designed to address localised linguistic and cultural contexts in a global framework.
format Preprint
id arxiv_https___arxiv_org_abs_2412_09818
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MERaLiON-AudioLLM: Bridging Audio and Language with Large Language Models
He, Yingxu
Liu, Zhuohan
Sun, Shuo
Wang, Bin
Zhang, Wenyu
Zou, Xunlong
Chen, Nancy F.
Aw, Ai Ti
Computation and Language
Artificial Intelligence
We introduce MERaLiON-AudioLLM (Multimodal Empathetic Reasoning and Learning in One Network), the first speech-text model tailored for Singapore's multilingual and multicultural landscape. Developed under the National Large Language Models Funding Initiative, Singapore, MERaLiON-AudioLLM integrates advanced speech and text processing to address the diverse linguistic nuances of local accents and dialects, enhancing accessibility and usability in complex, multilingual environments. Our results demonstrate improvements in both speech recognition and task-specific understanding, positioning MERaLiON-AudioLLM as a pioneering solution for region specific AI applications. We envision this release to set a precedent for future models designed to address localised linguistic and cultural contexts in a global framework.
title MERaLiON-AudioLLM: Bridging Audio and Language with Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2412.09818