Look Around and Pay Attention: Multi-camera Point Tracking Reimagined with Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Galoaa, Bishoy, Bai, Xiangyu, Moezzi, Shayda, Nandi, Utsav, Rangoju, Sai Siddhartha Vivek Dhir, Amraee, Somaieh, Ostadabbas, Sarah
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918230861283328
author Galoaa, Bishoy
Bai, Xiangyu
Moezzi, Shayda
Nandi, Utsav
Rangoju, Sai Siddhartha Vivek Dhir
Amraee, Somaieh
Ostadabbas, Sarah
author_facet Galoaa, Bishoy
Bai, Xiangyu
Moezzi, Shayda
Nandi, Utsav
Rangoju, Sai Siddhartha Vivek Dhir
Amraee, Somaieh
Ostadabbas, Sarah
contents This paper presents LAPA (Look Around and Pay Attention), a novel end-to-end transformer-based architecture for multi-camera point tracking that integrates appearance-based matching with geometric constraints. Traditional pipelines decouple detection, association, and tracking, leading to error propagation and temporal inconsistency in challenging scenarios. LAPA addresses these limitations by leveraging attention mechanisms to jointly reason across views and time, establishing soft correspondences through a cross-view attention mechanism enhanced with geometric priors. Instead of relying on classical triangulation, we construct 3D point representations via attention-weighted aggregation, inherently accommodating uncertainty and partial observations. Temporal consistency is further maintained through a transformer decoder that models long-range dependencies, preserving identities through extended occlusions. Extensive experiments on challenging datasets, including our newly created multi-camera (MC) versions of TAPVid-3D panoptic and PointOdyssey, demonstrate that our unified approach significantly outperforms existing methods, achieving 37.5% APD on TAPVid-3D-MC and 90.3% APD on PointOdyssey-MC, particularly excelling in scenarios with complex motions and occlusions. Code is available at https://github.com/ostadabbas/Look-Around-and-Pay-Attention-LAPA-
format Preprint
id arxiv_https___arxiv_org_abs_2512_04213
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Look Around and Pay Attention: Multi-camera Point Tracking Reimagined with Transformers
Galoaa, Bishoy
Bai, Xiangyu
Moezzi, Shayda
Nandi, Utsav
Rangoju, Sai Siddhartha Vivek Dhir
Amraee, Somaieh
Ostadabbas, Sarah
Computer Vision and Pattern Recognition
This paper presents LAPA (Look Around and Pay Attention), a novel end-to-end transformer-based architecture for multi-camera point tracking that integrates appearance-based matching with geometric constraints. Traditional pipelines decouple detection, association, and tracking, leading to error propagation and temporal inconsistency in challenging scenarios. LAPA addresses these limitations by leveraging attention mechanisms to jointly reason across views and time, establishing soft correspondences through a cross-view attention mechanism enhanced with geometric priors. Instead of relying on classical triangulation, we construct 3D point representations via attention-weighted aggregation, inherently accommodating uncertainty and partial observations. Temporal consistency is further maintained through a transformer decoder that models long-range dependencies, preserving identities through extended occlusions. Extensive experiments on challenging datasets, including our newly created multi-camera (MC) versions of TAPVid-3D panoptic and PointOdyssey, demonstrate that our unified approach significantly outperforms existing methods, achieving 37.5% APD on TAPVid-3D-MC and 90.3% APD on PointOdyssey-MC, particularly excelling in scenarios with complex motions and occlusions. Code is available at https://github.com/ostadabbas/Look-Around-and-Pay-Attention-LAPA-
title Look Around and Pay Attention: Multi-camera Point Tracking Reimagined with Transformers
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.04213