Déjà Vu: Efficient Video-Language Query Engine with Learning-based Inter-Frame Computation Reuse

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hwang, Jinwoo, Kim, Daeun, Lee, Sangyeop, Kim, Yoonsung, Heo, Guseul, Kim, Hojoon, Jeong, Yunseok, Meaza, Tadiwos, Park, Eunhyeok, Ahn, Jeongseob, Park, Jongse
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909779526418432
author Hwang, Jinwoo
Kim, Daeun
Lee, Sangyeop
Kim, Yoonsung
Heo, Guseul
Kim, Hojoon
Jeong, Yunseok
Meaza, Tadiwos
Park, Eunhyeok
Ahn, Jeongseob
Park, Jongse
author_facet Hwang, Jinwoo
Kim, Daeun
Lee, Sangyeop
Kim, Yoonsung
Heo, Guseul
Kim, Hojoon
Jeong, Yunseok
Meaza, Tadiwos
Park, Eunhyeok
Ahn, Jeongseob
Park, Jongse
contents Recently, Video-Language Models (VideoLMs) have demonstrated remarkable capabilities, offering significant potential for flexible and powerful video query systems. These models typically rely on Vision Transformers (ViTs), which process video frames individually to extract visual embeddings. However, generating embeddings for large-scale videos requires ViT inferencing across numerous frames, posing a major hurdle to real-world deployment and necessitating solutions for integration into scalable video data management systems. This paper introduces Déjà Vu, a video-language query engine that accelerates ViT-based VideoLMs by reusing computations across consecutive frames. At its core is ReuseViT, a modified ViT model specifically designed for VideoLM tasks, which learns to detect inter-frame reuse opportunities, striking an effective balance between accuracy and reuse. Although ReuseViT significantly reduces computation, these savings do not directly translate into performance gains on GPUs. To overcome this, Déjà Vu integrates memory-compute joint compaction techniques that convert the FLOP savings into tangible performance gains. Evaluations on three VideoLM tasks show that Déjà Vu accelerates embedding generation by up to a 2.64x within a 2% error bound, dramatically enhancing the practicality of VideoLMs for large-scale video analytics.
format Preprint
id arxiv_https___arxiv_org_abs_2506_14107
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Déjà Vu: Efficient Video-Language Query Engine with Learning-based Inter-Frame Computation Reuse
Hwang, Jinwoo
Kim, Daeun
Lee, Sangyeop
Kim, Yoonsung
Heo, Guseul
Kim, Hojoon
Jeong, Yunseok
Meaza, Tadiwos
Park, Eunhyeok
Ahn, Jeongseob
Park, Jongse
Distributed, Parallel, and Cluster Computing
Computer Vision and Pattern Recognition
Recently, Video-Language Models (VideoLMs) have demonstrated remarkable capabilities, offering significant potential for flexible and powerful video query systems. These models typically rely on Vision Transformers (ViTs), which process video frames individually to extract visual embeddings. However, generating embeddings for large-scale videos requires ViT inferencing across numerous frames, posing a major hurdle to real-world deployment and necessitating solutions for integration into scalable video data management systems. This paper introduces Déjà Vu, a video-language query engine that accelerates ViT-based VideoLMs by reusing computations across consecutive frames. At its core is ReuseViT, a modified ViT model specifically designed for VideoLM tasks, which learns to detect inter-frame reuse opportunities, striking an effective balance between accuracy and reuse. Although ReuseViT significantly reduces computation, these savings do not directly translate into performance gains on GPUs. To overcome this, Déjà Vu integrates memory-compute joint compaction techniques that convert the FLOP savings into tangible performance gains. Evaluations on three VideoLM tasks show that Déjà Vu accelerates embedding generation by up to a 2.64x within a 2% error bound, dramatically enhancing the practicality of VideoLMs for large-scale video analytics.
title Déjà Vu: Efficient Video-Language Query Engine with Learning-based Inter-Frame Computation Reuse
topic Distributed, Parallel, and Cluster Computing
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.14107