Quantifying the human visual exposome with vision language models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Rominger, Christian, Schwerdtfeger, Andreas R., Singh, Malay Gaherwar, Khudyakow, Dimitri, Michels, Elizabeth A. M., Wolf, Fabian, Kather, Jakob Nikolas, Wekenborg, Magdalena Katharina
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909014597566464
author Rominger, Christian
Schwerdtfeger, Andreas R.
Singh, Malay Gaherwar
Khudyakow, Dimitri
Michels, Elizabeth A. M.
Wolf, Fabian
Kather, Jakob Nikolas
Wekenborg, Magdalena Katharina
author_facet Rominger, Christian
Schwerdtfeger, Andreas R.
Singh, Malay Gaherwar
Khudyakow, Dimitri
Michels, Elizabeth A. M.
Wolf, Fabian
Kather, Jakob Nikolas
Wekenborg, Magdalena Katharina
contents The visual environment is a fundamental yet unquantified determinant of mental health. While the concept of the environmental exposome is well established, current methods rely on coarse geospatial proxies or biased self reports, failing to capture the first person visual context of daily life. We addressed this gap by coupling ecological momentary assessment with vision language models (VLMs) to quantify the semantic richness of human visual experience. Across 2674 participant generated photographs, VLM derived estimates of greenness robustly predicted momentary affect and chronic stress, consistent with established benchmarks. We then developed a semi autonomous large language model (LLM) based pipeline that mined over seven million scientific publications to extract nearly 1000 environmental features empirically linked to mental health. When applied to real world imagery, up to 33 percent of VLM extracted context ratings significantly correlated with affect and stress. These findings establish a scalable objective paradigm for visual exposomics, enabling high throughput decoding of how the visible world is associated with mental health.
format Preprint
id arxiv_https___arxiv_org_abs_2605_03863
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Quantifying the human visual exposome with vision language models
Rominger, Christian
Schwerdtfeger, Andreas R.
Singh, Malay Gaherwar
Khudyakow, Dimitri
Michels, Elizabeth A. M.
Wolf, Fabian
Kather, Jakob Nikolas
Wekenborg, Magdalena Katharina
Artificial Intelligence
Computer Vision and Pattern Recognition
The visual environment is a fundamental yet unquantified determinant of mental health. While the concept of the environmental exposome is well established, current methods rely on coarse geospatial proxies or biased self reports, failing to capture the first person visual context of daily life. We addressed this gap by coupling ecological momentary assessment with vision language models (VLMs) to quantify the semantic richness of human visual experience. Across 2674 participant generated photographs, VLM derived estimates of greenness robustly predicted momentary affect and chronic stress, consistent with established benchmarks. We then developed a semi autonomous large language model (LLM) based pipeline that mined over seven million scientific publications to extract nearly 1000 environmental features empirically linked to mental health. When applied to real world imagery, up to 33 percent of VLM extracted context ratings significantly correlated with affect and stress. These findings establish a scalable objective paradigm for visual exposomics, enabling high throughput decoding of how the visible world is associated with mental health.
title Quantifying the human visual exposome with vision language models
topic Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.03863