Free-VSC: Free Semantics from Visual Foundation Models for Unsupervised Video Semantic Compression

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tian, Yuan, Lu, Guo, Zhai, Guangtao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910615675600896
author Tian, Yuan
Lu, Guo
Zhai, Guangtao
author_facet Tian, Yuan
Lu, Guo
Zhai, Guangtao
contents Unsupervised video semantic compression (UVSC), i.e., compressing videos to better support various analysis tasks, has recently garnered attention. However, the semantic richness of previous methods remains limited, due to the single semantic learning objective, limited training data, etc. To address this, we propose to boost the UVSC task by absorbing the off-the-shelf rich semantics from VFMs. Specifically, we introduce a VFMs-shared semantic alignment layer, complemented by VFM-specific prompts, to flexibly align semantics between the compressed video and various VFMs. This allows different VFMs to collaboratively build a mutually-enhanced semantic space, guiding the learning of the compression model. Moreover, we introduce a dynamic trajectory-based inter-frame compression scheme, which first estimates the semantic trajectory based on the historical content, and then traverses along the trajectory to predict the future semantics as the coding context. This reduces the overall bitcost of the system, further improving the compression efficiency. Our approach outperforms previous coding methods on three mainstream tasks and six datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2409_11718
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Free-VSC: Free Semantics from Visual Foundation Models for Unsupervised Video Semantic Compression
Tian, Yuan
Lu, Guo
Zhai, Guangtao
Computer Vision and Pattern Recognition
Unsupervised video semantic compression (UVSC), i.e., compressing videos to better support various analysis tasks, has recently garnered attention. However, the semantic richness of previous methods remains limited, due to the single semantic learning objective, limited training data, etc. To address this, we propose to boost the UVSC task by absorbing the off-the-shelf rich semantics from VFMs. Specifically, we introduce a VFMs-shared semantic alignment layer, complemented by VFM-specific prompts, to flexibly align semantics between the compressed video and various VFMs. This allows different VFMs to collaboratively build a mutually-enhanced semantic space, guiding the learning of the compression model. Moreover, we introduce a dynamic trajectory-based inter-frame compression scheme, which first estimates the semantic trajectory based on the historical content, and then traverses along the trajectory to predict the future semantics as the coding context. This reduces the overall bitcost of the system, further improving the compression efficiency. Our approach outperforms previous coding methods on three mainstream tasks and six datasets.
title Free-VSC: Free Semantics from Visual Foundation Models for Unsupervised Video Semantic Compression
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2409.11718