ARGOS: Who, Where, and When in Agentic Multi-Camera Person Search

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Myungchul, Park, Kwanyong, Kim, Junmo, Kweon, In So
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908963100950528
author Kim, Myungchul
Park, Kwanyong
Kim, Junmo
Kweon, In So
author_facet Kim, Myungchul
Park, Kwanyong
Kim, Junmo
Kweon, In So
contents We introduce ARGOS, the first benchmark and framework that reformulates multi-camera person search as an interactive reasoning problem requiring an agent to plan, question, and eliminate candidates under information asymmetry. An ARGOS agent receives a vague witness statement and must decide what to ask, when to invoke spatial or temporal tools, and how to interpret ambiguous responses, all within a limited turn budget. Reasoning is grounded in a Spatio-Temporal Topology Graph (STTG) encoding camera connectivity and empirically validated transition times. The benchmark comprises 2,691 tasks across 14 real-world scenarios in three progressive tracks: semantic perception (Who), spatial reasoning (Where), and temporal reasoning (When). Experiments with four LLM backbones show the benchmark is far from solved (best TWS: 0.383 on Track 2, 0.590 on Track 3), and ablations confirm that removing domain-specific tools drops accuracy by up to 49.6 percentage points.
format Preprint
id arxiv_https___arxiv_org_abs_2604_12762
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ARGOS: Who, Where, and When in Agentic Multi-Camera Person Search
Kim, Myungchul
Park, Kwanyong
Kim, Junmo
Kweon, In So
Computer Vision and Pattern Recognition
Artificial Intelligence
Multiagent Systems
We introduce ARGOS, the first benchmark and framework that reformulates multi-camera person search as an interactive reasoning problem requiring an agent to plan, question, and eliminate candidates under information asymmetry. An ARGOS agent receives a vague witness statement and must decide what to ask, when to invoke spatial or temporal tools, and how to interpret ambiguous responses, all within a limited turn budget. Reasoning is grounded in a Spatio-Temporal Topology Graph (STTG) encoding camera connectivity and empirically validated transition times. The benchmark comprises 2,691 tasks across 14 real-world scenarios in three progressive tracks: semantic perception (Who), spatial reasoning (Where), and temporal reasoning (When). Experiments with four LLM backbones show the benchmark is far from solved (best TWS: 0.383 on Track 2, 0.590 on Track 3), and ablations confirm that removing domain-specific tools drops accuracy by up to 49.6 percentage points.
title ARGOS: Who, Where, and When in Agentic Multi-Camera Person Search
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Multiagent Systems
url https://arxiv.org/abs/2604.12762