Chitranuvad: Adapting Multi-Lingual LLMs for Multimodal Translation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Khan, Shaharukh, Tarun, Ayush, Faraz, Ali, Kamble, Palash, Dahiya, Vivek, Pokala, Praveen, Kulkarni, Ashish, Khatri, Chandra, Ravi, Abhinav, Agarwal, Shubham
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912251553775616
author Khan, Shaharukh
Tarun, Ayush
Faraz, Ali
Kamble, Palash
Dahiya, Vivek
Pokala, Praveen
Kulkarni, Ashish
Khatri, Chandra
Ravi, Abhinav
Agarwal, Shubham
author_facet Khan, Shaharukh
Tarun, Ayush
Faraz, Ali
Kamble, Palash
Dahiya, Vivek
Pokala, Praveen
Kulkarni, Ashish
Khatri, Chandra
Ravi, Abhinav
Agarwal, Shubham
contents In this work, we provide the system description of our submission as part of the English to Lowres Multimodal Translation Task at the Workshop on Asian Translation (WAT2024). We introduce Chitranuvad, a multimodal model that effectively integrates Multilingual LLM and a vision module for Multimodal Translation. Our method uses a ViT image encoder to extract visual representations as visual token embeddings which are projected to the LLM space by an adapter layer and generates translation in an autoregressive fashion. We participated in all the three tracks (Image Captioning, Text only and Multimodal translation tasks) for Indic languages (ie. English translation to Hindi, Bengali and Malyalam) and achieved SOTA results for Hindi in all of them on the Challenge set while remaining competitive for the other languages in the shared task.
format Preprint
id arxiv_https___arxiv_org_abs_2502_20420
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Chitranuvad: Adapting Multi-Lingual LLMs for Multimodal Translation
Khan, Shaharukh
Tarun, Ayush
Faraz, Ali
Kamble, Palash
Dahiya, Vivek
Pokala, Praveen
Kulkarni, Ashish
Khatri, Chandra
Ravi, Abhinav
Agarwal, Shubham
Computation and Language
Computer Vision and Pattern Recognition
In this work, we provide the system description of our submission as part of the English to Lowres Multimodal Translation Task at the Workshop on Asian Translation (WAT2024). We introduce Chitranuvad, a multimodal model that effectively integrates Multilingual LLM and a vision module for Multimodal Translation. Our method uses a ViT image encoder to extract visual representations as visual token embeddings which are projected to the LLM space by an adapter layer and generates translation in an autoregressive fashion. We participated in all the three tracks (Image Captioning, Text only and Multimodal translation tasks) for Indic languages (ie. English translation to Hindi, Bengali and Malyalam) and achieved SOTA results for Hindi in all of them on the Challenge set while remaining competitive for the other languages in the shared task.
title Chitranuvad: Adapting Multi-Lingual LLMs for Multimodal Translation
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2502.20420