Large Language Models Implicitly Learn to See and Hear Just By Reading

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Verma, Prateek, Pilanci, Mert
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909801201532928
author Verma, Prateek
Pilanci, Mert
author_facet Verma, Prateek
Pilanci, Mert
contents This paper presents a fascinating find: By training an auto-regressive LLM model on text tokens, the text model inherently develops internally an ability to understand images and audio, thereby developing the ability to see and hear just by reading. Popular audio and visual LLM models fine-tune text LLM models to give text output conditioned on images and audio embeddings. On the other hand, our architecture takes in patches of images, audio waveforms or tokens as input. It gives us the embeddings or category labels typical of a classification pipeline. We show the generality of text weights in aiding audio classification for datasets FSD-50K and GTZAN. Further, we show this working for image classification on CIFAR-10 and Fashion-MNIST, as well on image patches. This pushes the notion of text-LLMs learning powerful internal circuits that can be utilized by activating necessary connections for various applications rather than training models from scratch every single time.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17091
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Large Language Models Implicitly Learn to See and Hear Just By Reading
Verma, Prateek
Pilanci, Mert
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
Sound
Audio and Speech Processing
This paper presents a fascinating find: By training an auto-regressive LLM model on text tokens, the text model inherently develops internally an ability to understand images and audio, thereby developing the ability to see and hear just by reading. Popular audio and visual LLM models fine-tune text LLM models to give text output conditioned on images and audio embeddings. On the other hand, our architecture takes in patches of images, audio waveforms or tokens as input. It gives us the embeddings or category labels typical of a classification pipeline. We show the generality of text weights in aiding audio classification for datasets FSD-50K and GTZAN. Further, we show this working for image classification on CIFAR-10 and Fashion-MNIST, as well on image patches. This pushes the notion of text-LLMs learning powerful internal circuits that can be utilized by activating necessary connections for various applications rather than training models from scratch every single time.
title Large Language Models Implicitly Learn to See and Hear Just By Reading
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2505.17091