Chitranuvad: Adapting Multi-Lingual LLMs for Multimodal Translation
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866912251553775616 |
|---|---|
| author | Khan, Shaharukh Tarun, Ayush Faraz, Ali Kamble, Palash Dahiya, Vivek Pokala, Praveen Kulkarni, Ashish Khatri, Chandra Ravi, Abhinav Agarwal, Shubham |
| author_facet | Khan, Shaharukh Tarun, Ayush Faraz, Ali Kamble, Palash Dahiya, Vivek Pokala, Praveen Kulkarni, Ashish Khatri, Chandra Ravi, Abhinav Agarwal, Shubham |
| contents | In this work, we provide the system description of our submission as part of the English to Lowres Multimodal Translation Task at the Workshop on Asian Translation (WAT2024). We introduce Chitranuvad, a multimodal model that effectively integrates Multilingual LLM and a vision module for Multimodal Translation. Our method uses a ViT image encoder to extract visual representations as visual token embeddings which are projected to the LLM space by an adapter layer and generates translation in an autoregressive fashion. We participated in all the three tracks (Image Captioning, Text only and Multimodal translation tasks) for Indic languages (ie. English translation to Hindi, Bengali and Malyalam) and achieved SOTA results for Hindi in all of them on the Challenge set while remaining competitive for the other languages in the shared task. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2502_20420 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Chitranuvad: Adapting Multi-Lingual LLMs for Multimodal Translation Khan, Shaharukh Tarun, Ayush Faraz, Ali Kamble, Palash Dahiya, Vivek Pokala, Praveen Kulkarni, Ashish Khatri, Chandra Ravi, Abhinav Agarwal, Shubham Computation and Language Computer Vision and Pattern Recognition In this work, we provide the system description of our submission as part of the English to Lowres Multimodal Translation Task at the Workshop on Asian Translation (WAT2024). We introduce Chitranuvad, a multimodal model that effectively integrates Multilingual LLM and a vision module for Multimodal Translation. Our method uses a ViT image encoder to extract visual representations as visual token embeddings which are projected to the LLM space by an adapter layer and generates translation in an autoregressive fashion. We participated in all the three tracks (Image Captioning, Text only and Multimodal translation tasks) for Indic languages (ie. English translation to Hindi, Bengali and Malyalam) and achieved SOTA results for Hindi in all of them on the Challenge set while remaining competitive for the other languages in the shared task. |
| title | Chitranuvad: Adapting Multi-Lingual LLMs for Multimodal Translation |
| topic | Computation and Language Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2502.20420 |