GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916774020120576 |
|---|---|
| author | Siingh, Shikhhar Rawat, Abhinav Baral, Chitta Gupta, Vivek |
| author_facet | Siingh, Shikhhar Rawat, Abhinav Baral, Chitta Gupta, Vivek |
| contents | Publicly significant images from events hold valuable contextual information, crucial for journalism and education. However, existing methods often struggle to extract this relevance accurately. To address this, we introduce GETReason (Geospatial Event Temporal Reasoning), a framework that moves beyond surface-level image descriptions to infer deeper contextual meaning. We propose that extracting global event, temporal, and geospatial information enhances understanding of an image's significance. Additionally, we introduce GREAT (Geospatial Reasoning and Event Accuracy with Temporal Alignment), a new metric for evaluating reasoning-based image understanding. Our layered multi-agent approach, assessed using a reasoning-weighted metric, demonstrates that meaningful insights can be inferred, effectively linking images to their broader event context. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_21863 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning Siingh, Shikhhar Rawat, Abhinav Baral, Chitta Gupta, Vivek Computer Vision and Pattern Recognition Computation and Language Publicly significant images from events hold valuable contextual information, crucial for journalism and education. However, existing methods often struggle to extract this relevance accurately. To address this, we introduce GETReason (Geospatial Event Temporal Reasoning), a framework that moves beyond surface-level image descriptions to infer deeper contextual meaning. We propose that extracting global event, temporal, and geospatial information enhances understanding of an image's significance. Additionally, we introduce GREAT (Geospatial Reasoning and Event Accuracy with Temporal Alignment), a new metric for evaluating reasoning-based image understanding. Our layered multi-agent approach, assessed using a reasoning-weighted metric, demonstrates that meaningful insights can be inferred, effectively linking images to their broader event context. |
| title | GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning |
| topic | Computer Vision and Pattern Recognition Computation and Language |
| url | https://arxiv.org/abs/2505.21863 |