Swift Sampling: Selecting Temporal Surprises via Taylor Series

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Dahye, Sachdeva, Bhuvan, Uppal, Karan, Gupta, Naman, Balasubramanian, Vineeth N., Ghadiyaram, Deepti
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911704992972800
author Kim, Dahye
Sachdeva, Bhuvan
Uppal, Karan
Gupta, Naman
Balasubramanian, Vineeth N.
Ghadiyaram, Deepti
author_facet Kim, Dahye
Sachdeva, Bhuvan
Uppal, Karan
Gupta, Naman
Balasubramanian, Vineeth N.
Ghadiyaram, Deepti
contents While most frames in long-form video are redundant, the critical information resides in temporal surprises: moments where the actual visual features deviate from their predicted evolution. Inspired by the human brain's predictive coding, we introduce Swift Sampling, an elegant, training-free frame selection algorithm that automatically identifies high-information moments in a video. Specifically, we model a video as a differentiable trajectory in the visual latent space and compute the velocity and acceleration of its features. Then, we apply Taylor expansion to project the expected path of subsequent frames. Frames that diverge sharply from this predicted manifold are identified as temporally surprising frames and selected for sampling. Unlike prior training-free methods that rely on auxiliary networks or video-specific hyperparameter tuning, Swift Sampling is incredibly lightweight, adding only 0.02x additional computational cost over baseline making it 30x cheaper overhead than leading baselines. Across three long-video question answering benchmarks and 10 different downstream tasks, Swift Sampling outperforms uniform sampling and prior query-agnostic baselines. It is especially powerful for long videos with limited frame budgets improving accuracy by up to +12.5 points.
format Preprint
id arxiv_https___arxiv_org_abs_2605_22678
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Swift Sampling: Selecting Temporal Surprises via Taylor Series
Kim, Dahye
Sachdeva, Bhuvan
Uppal, Karan
Gupta, Naman
Balasubramanian, Vineeth N.
Ghadiyaram, Deepti
Computer Vision and Pattern Recognition
Artificial Intelligence
While most frames in long-form video are redundant, the critical information resides in temporal surprises: moments where the actual visual features deviate from their predicted evolution. Inspired by the human brain's predictive coding, we introduce Swift Sampling, an elegant, training-free frame selection algorithm that automatically identifies high-information moments in a video. Specifically, we model a video as a differentiable trajectory in the visual latent space and compute the velocity and acceleration of its features. Then, we apply Taylor expansion to project the expected path of subsequent frames. Frames that diverge sharply from this predicted manifold are identified as temporally surprising frames and selected for sampling. Unlike prior training-free methods that rely on auxiliary networks or video-specific hyperparameter tuning, Swift Sampling is incredibly lightweight, adding only 0.02x additional computational cost over baseline making it 30x cheaper overhead than leading baselines. Across three long-video question answering benchmarks and 10 different downstream tasks, Swift Sampling outperforms uniform sampling and prior query-agnostic baselines. It is especially powerful for long videos with limited frame budgets improving accuracy by up to +12.5 points.
title Swift Sampling: Selecting Temporal Surprises via Taylor Series
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2605.22678