ChartGaze: Enhancing Chart Understanding in LVLMs with Eye-Tracking Guided Attention Refinement

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Salamatian, Ali, Abaskohi, Amirhossein, Fan, Wan-Cyuan, Hossain, Mir Rayat Imtiaz, Sigal, Leonid, Carenini, Giuseppe
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915497524592640
author Salamatian, Ali
Abaskohi, Amirhossein
Fan, Wan-Cyuan
Hossain, Mir Rayat Imtiaz
Sigal, Leonid
Carenini, Giuseppe
author_facet Salamatian, Ali
Abaskohi, Amirhossein
Fan, Wan-Cyuan
Hossain, Mir Rayat Imtiaz
Sigal, Leonid
Carenini, Giuseppe
contents Charts are a crucial visual medium for communicating and representing information. While Large Vision-Language Models (LVLMs) have made progress on chart question answering (CQA), the task remains challenging, particularly when models attend to irrelevant regions of the chart. In this work, we present ChartGaze, a new eye-tracking dataset that captures human gaze patterns during chart reasoning tasks. Through a systematic comparison of human and model attention, we find that LVLMs often diverge from human gaze, leading to reduced interpretability and accuracy. To address this, we propose a gaze-guided attention refinement that aligns image-text attention with human fixations. Our approach improves both answer accuracy and attention alignment, yielding gains of up to 2.56 percentage points across multiple models. These results demonstrate the promise of incorporating human gaze to enhance both the reasoning quality and interpretability of chart-focused LVLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2509_13282
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ChartGaze: Enhancing Chart Understanding in LVLMs with Eye-Tracking Guided Attention Refinement
Salamatian, Ali
Abaskohi, Amirhossein
Fan, Wan-Cyuan
Hossain, Mir Rayat Imtiaz
Sigal, Leonid
Carenini, Giuseppe
Computation and Language
Computer Vision and Pattern Recognition
Machine Learning
Charts are a crucial visual medium for communicating and representing information. While Large Vision-Language Models (LVLMs) have made progress on chart question answering (CQA), the task remains challenging, particularly when models attend to irrelevant regions of the chart. In this work, we present ChartGaze, a new eye-tracking dataset that captures human gaze patterns during chart reasoning tasks. Through a systematic comparison of human and model attention, we find that LVLMs often diverge from human gaze, leading to reduced interpretability and accuracy. To address this, we propose a gaze-guided attention refinement that aligns image-text attention with human fixations. Our approach improves both answer accuracy and attention alignment, yielding gains of up to 2.56 percentage points across multiple models. These results demonstrate the promise of incorporating human gaze to enhance both the reasoning quality and interpretability of chart-focused LVLMs.
title ChartGaze: Enhancing Chart Understanding in LVLMs with Eye-Tracking Guided Attention Refinement
topic Computation and Language
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2509.13282