TextTeacher: What Can Language Teach About Images?

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Nauen, Tobias Christian, Frolov, Stanislav, Moser, Brian Bernhard, Raue, Federico, Anwar, Ahmed, Dengel, Andreas
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913152111738880
author Nauen, Tobias Christian
Frolov, Stanislav
Moser, Brian Bernhard
Raue, Federico
Anwar, Ahmed
Dengel, Andreas
author_facet Nauen, Tobias Christian
Frolov, Stanislav
Moser, Brian Bernhard
Raue, Federico
Anwar, Ahmed
Dengel, Andreas
contents The platonic representation hypothesis suggests that sufficiently large models converge to a shared representation geometry, even across modalities. Motivated by this, we ask: Can the semantic knowledge of a language model efficiently improve a vision model? As an answer, we introduce TextTeacher, a simple auxiliary objective that injects text embeddings as additional information into image classification training. TextTeacher uses readily available image captions, a pre-trained and frozen text encoder, and a lightweight projection to produce semantic anchors that efficiently guide representations during training while leaving the inference-time model unchanged. On ImageNet with standard ViT backbones, TextTeacher improves accuracy by up to +2.7 percentage points (p.p.) and yields consistent transfer gains (on average +1.0 p.p.) under the same recipe and compute. It outperforms vision knowledge distillation, yielding more accuracy at a constant compute budget or similar accuracy, but 33% faster. Our analysis indicates that TextTeacher acts as a feature-space preconditioner, shaping deeper layers in the first stages of training, and aiding generalization by supplying complementary semantic cues. TextTeacher adds negligible overhead, requires no costly multimodal training of the target model and preserves the simplicity and latency of pure vision models. Project page with code and captions: https://nauen-it.de/publications/text-teacher
format Preprint
id arxiv_https___arxiv_org_abs_2605_22098
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TextTeacher: What Can Language Teach About Images?
Nauen, Tobias Christian
Frolov, Stanislav
Moser, Brian Bernhard
Raue, Federico
Anwar, Ahmed
Dengel, Andreas
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
68T05 (Primary), 68T45 (Secondary)
I.2.6; I.2.10
The platonic representation hypothesis suggests that sufficiently large models converge to a shared representation geometry, even across modalities. Motivated by this, we ask: Can the semantic knowledge of a language model efficiently improve a vision model? As an answer, we introduce TextTeacher, a simple auxiliary objective that injects text embeddings as additional information into image classification training. TextTeacher uses readily available image captions, a pre-trained and frozen text encoder, and a lightweight projection to produce semantic anchors that efficiently guide representations during training while leaving the inference-time model unchanged. On ImageNet with standard ViT backbones, TextTeacher improves accuracy by up to +2.7 percentage points (p.p.) and yields consistent transfer gains (on average +1.0 p.p.) under the same recipe and compute. It outperforms vision knowledge distillation, yielding more accuracy at a constant compute budget or similar accuracy, but 33% faster. Our analysis indicates that TextTeacher acts as a feature-space preconditioner, shaping deeper layers in the first stages of training, and aiding generalization by supplying complementary semantic cues. TextTeacher adds negligible overhead, requires no costly multimodal training of the target model and preserves the simplicity and latency of pure vision models. Project page with code and captions: https://nauen-it.de/publications/text-teacher
title TextTeacher: What Can Language Teach About Images?
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
68T05 (Primary), 68T45 (Secondary)
I.2.6; I.2.10
url https://arxiv.org/abs/2605.22098