Seeing without Pixels: Perception from Camera Trajectories

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xue, Zihui, Grauman, Kristen, Damen, Dima, Zisserman, Andrew, Han, Tengda
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911561641099264
author Xue, Zihui
Grauman, Kristen
Damen, Dima
Zisserman, Andrew
Han, Tengda
author_facet Xue, Zihui
Grauman, Kristen
Damen, Dima
Zisserman, Andrew
Han, Tengda
contents Can one perceive a video's content without seeing its pixels, just from the camera trajectory-the path it carves through space? This paper is the first to systematically investigate this seemingly implausible question. Towards this end, we propose a contrastive learning framework to train CamFormer, a dedicated encoder that projects camera pose trajectories into a joint embedding space, aligning them with natural language. We find that, contrary to its apparent simplicity, the camera trajectory is a remarkably informative signal to uncover video content. In other words, "how you move" can indeed provide valuable cues about "what you are doing" (egocentric) or "observing" (exocentric). We demonstrate the versatility of our learned CamFormer embeddings on a diverse suite of downstream tasks, ranging from cross-modal alignment to classification and temporal analysis. Importantly, our representations are robust across diverse camera pose estimation methods, including both high-fidelity multi-sensored and standard RGB-only estimators. Our findings establish camera trajectory as a lightweight, robust, and versatile modality for perceiving video content.
format Preprint
id arxiv_https___arxiv_org_abs_2511_21681
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Seeing without Pixels: Perception from Camera Trajectories
Xue, Zihui
Grauman, Kristen
Damen, Dima
Zisserman, Andrew
Han, Tengda
Computer Vision and Pattern Recognition
Can one perceive a video's content without seeing its pixels, just from the camera trajectory-the path it carves through space? This paper is the first to systematically investigate this seemingly implausible question. Towards this end, we propose a contrastive learning framework to train CamFormer, a dedicated encoder that projects camera pose trajectories into a joint embedding space, aligning them with natural language. We find that, contrary to its apparent simplicity, the camera trajectory is a remarkably informative signal to uncover video content. In other words, "how you move" can indeed provide valuable cues about "what you are doing" (egocentric) or "observing" (exocentric). We demonstrate the versatility of our learned CamFormer embeddings on a diverse suite of downstream tasks, ranging from cross-modal alignment to classification and temporal analysis. Importantly, our representations are robust across diverse camera pose estimation methods, including both high-fidelity multi-sensored and standard RGB-only estimators. Our findings establish camera trajectory as a lightweight, robust, and versatile modality for perceiving video content.
title Seeing without Pixels: Perception from Camera Trajectories
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.21681