GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Siingh, Shikhhar, Rawat, Abhinav, Baral, Chitta, Gupta, Vivek
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916774020120576
author Siingh, Shikhhar
Rawat, Abhinav
Baral, Chitta
Gupta, Vivek
author_facet Siingh, Shikhhar
Rawat, Abhinav
Baral, Chitta
Gupta, Vivek
contents Publicly significant images from events hold valuable contextual information, crucial for journalism and education. However, existing methods often struggle to extract this relevance accurately. To address this, we introduce GETReason (Geospatial Event Temporal Reasoning), a framework that moves beyond surface-level image descriptions to infer deeper contextual meaning. We propose that extracting global event, temporal, and geospatial information enhances understanding of an image's significance. Additionally, we introduce GREAT (Geospatial Reasoning and Event Accuracy with Temporal Alignment), a new metric for evaluating reasoning-based image understanding. Our layered multi-agent approach, assessed using a reasoning-weighted metric, demonstrates that meaningful insights can be inferred, effectively linking images to their broader event context.
format Preprint
id arxiv_https___arxiv_org_abs_2505_21863
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning
Siingh, Shikhhar
Rawat, Abhinav
Baral, Chitta
Gupta, Vivek
Computer Vision and Pattern Recognition
Computation and Language
Publicly significant images from events hold valuable contextual information, crucial for journalism and education. However, existing methods often struggle to extract this relevance accurately. To address this, we introduce GETReason (Geospatial Event Temporal Reasoning), a framework that moves beyond surface-level image descriptions to infer deeper contextual meaning. We propose that extracting global event, temporal, and geospatial information enhances understanding of an image's significance. Additionally, we introduce GREAT (Geospatial Reasoning and Event Accuracy with Temporal Alignment), a new metric for evaluating reasoning-based image understanding. Our layered multi-agent approach, assessed using a reasoning-weighted metric, demonstrates that meaningful insights can be inferred, effectively linking images to their broader event context.
title GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2505.21863