Tracking without Seeing: Geospatial Inference using Encrypted Traffic from Distributed Nodes

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yetim, Sadik Yagiz, Dong, Gaofeng, Zanoria, Isaac-Neil, Barman, Ronit, Wigness, Maggie, Abdelzaher, Tarek, Srivastava, Mani, Diggavi, Suhas
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911552224886784
author Yetim, Sadik Yagiz
Dong, Gaofeng
Zanoria, Isaac-Neil
Barman, Ronit
Wigness, Maggie
Abdelzaher, Tarek
Srivastava, Mani
Diggavi, Suhas
author_facet Yetim, Sadik Yagiz
Dong, Gaofeng
Zanoria, Isaac-Neil
Barman, Ronit
Wigness, Maggie
Abdelzaher, Tarek
Srivastava, Mani
Diggavi, Suhas
contents Accurate observation of dynamic environments traditionally relies on synthesizing raw, signal-level information from multiple distributed sensors. This work investigates an alternative approach: performing geospatial inference using only encrypted packet-level information, without access to the raw sensory data. We further explore how this indirect information can be fused with directly available sensory data to extend overall inference capabilities. We introduce GraySense, a learning-based framework that performs geospatial object tracking by analyzing encrypted wireless video transmission traffic, such as packet sizes, from cameras with inaccessible streams. GraySense leverages the inherent relationship between scene dynamics and transmitted packet sizes to infer object motion. The framework consists of two stages: (1) a Packet Grouping module that identifies frame boundaries and estimates frame sizes from encrypted network traffic, and (2) a Tracker module, based on a Transformer encoder with a recurrent state, which fuses indirect packet-based inputs with optional direct camera-based inputs to estimate the object's position. Extensive experiments with realistic videos from the CARLA simulator and emulated networks under varying conditions show that GraySense achieves 2.33 meters tracking error (Euclidean distance) without raw signal access, within the dimensions of tracked objects (4.61m x 1.93m). To our knowledge, this capability has not been previously demonstrated, expanding the use of latent signals for sensing.
format Preprint
id arxiv_https___arxiv_org_abs_2603_27811
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Tracking without Seeing: Geospatial Inference using Encrypted Traffic from Distributed Nodes
Yetim, Sadik Yagiz
Dong, Gaofeng
Zanoria, Isaac-Neil
Barman, Ronit
Wigness, Maggie
Abdelzaher, Tarek
Srivastava, Mani
Diggavi, Suhas
Computer Vision and Pattern Recognition
Machine Learning
Networking and Internet Architecture
Accurate observation of dynamic environments traditionally relies on synthesizing raw, signal-level information from multiple distributed sensors. This work investigates an alternative approach: performing geospatial inference using only encrypted packet-level information, without access to the raw sensory data. We further explore how this indirect information can be fused with directly available sensory data to extend overall inference capabilities. We introduce GraySense, a learning-based framework that performs geospatial object tracking by analyzing encrypted wireless video transmission traffic, such as packet sizes, from cameras with inaccessible streams. GraySense leverages the inherent relationship between scene dynamics and transmitted packet sizes to infer object motion. The framework consists of two stages: (1) a Packet Grouping module that identifies frame boundaries and estimates frame sizes from encrypted network traffic, and (2) a Tracker module, based on a Transformer encoder with a recurrent state, which fuses indirect packet-based inputs with optional direct camera-based inputs to estimate the object's position. Extensive experiments with realistic videos from the CARLA simulator and emulated networks under varying conditions show that GraySense achieves 2.33 meters tracking error (Euclidean distance) without raw signal access, within the dimensions of tracked objects (4.61m x 1.93m). To our knowledge, this capability has not been previously demonstrated, expanding the use of latent signals for sensing.
title Tracking without Seeing: Geospatial Inference using Encrypted Traffic from Distributed Nodes
topic Computer Vision and Pattern Recognition
Machine Learning
Networking and Internet Architecture
url https://arxiv.org/abs/2603.27811