LLMs can see and hear without any training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ashutosh, Kumar, Gandelsman, Yossi, Chen, Xinlei, Misra, Ishan, Girdhar, Rohit
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910805769846784
author Ashutosh, Kumar
Gandelsman, Yossi
Chen, Xinlei
Misra, Ishan
Girdhar, Rohit
author_facet Ashutosh, Kumar
Gandelsman, Yossi
Chen, Xinlei
Misra, Ishan
Girdhar, Rohit
contents We present MILS: Multimodal Iterative LLM Solver, a surprisingly simple, training-free approach, to imbue multimodal capabilities into your favorite LLM. Leveraging their innate ability to perform multi-step reasoning, MILS prompts the LLM to generate candidate outputs, each of which are scored and fed back iteratively, eventually generating a solution to the task. This enables various applications that typically require training specialized models on task-specific data. In particular, we establish a new state-of-the-art on emergent zero-shot image, video and audio captioning. MILS seamlessly applies to media generation as well, discovering prompt rewrites to improve text-to-image generation, and even edit prompts for style transfer! Finally, being a gradient-free optimization approach, MILS can invert multimodal embeddings into text, enabling applications like cross-modal arithmetic.
format Preprint
id arxiv_https___arxiv_org_abs_2501_18096
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LLMs can see and hear without any training
Ashutosh, Kumar
Gandelsman, Yossi
Chen, Xinlei
Misra, Ishan
Girdhar, Rohit
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
We present MILS: Multimodal Iterative LLM Solver, a surprisingly simple, training-free approach, to imbue multimodal capabilities into your favorite LLM. Leveraging their innate ability to perform multi-step reasoning, MILS prompts the LLM to generate candidate outputs, each of which are scored and fed back iteratively, eventually generating a solution to the task. This enables various applications that typically require training specialized models on task-specific data. In particular, we establish a new state-of-the-art on emergent zero-shot image, video and audio captioning. MILS seamlessly applies to media generation as well, discovering prompt rewrites to improve text-to-image generation, and even edit prompts for style transfer! Finally, being a gradient-free optimization approach, MILS can invert multimodal embeddings into text, enabling applications like cross-modal arithmetic.
title LLMs can see and hear without any training
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2501.18096