RL-RC-DoT: A Block-level RL agent for Task-Aware Video Compression

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gadot, Uri, Shocher, Assaf, Mannor, Shie, Chechik, Gal, Hallak, Assaf
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916661584461824
author Gadot, Uri
Shocher, Assaf
Mannor, Shie
Chechik, Gal
Hallak, Assaf
author_facet Gadot, Uri
Shocher, Assaf
Mannor, Shie
Chechik, Gal
Hallak, Assaf
contents Video encoders optimize compression for human perception by minimizing reconstruction error under bit-rate constraints. In many modern applications such as autonomous driving, an overwhelming majority of videos serve as input for AI systems performing tasks like object recognition or segmentation, rather than being watched by humans. It is therefore useful to optimize the encoder for a downstream task instead of for perceptual image quality. However, a major challenge is how to combine such downstream optimization with existing standard video encoders, which are highly efficient and popular. Here, we address this challenge by controlling the Quantization Parameters (QPs) at the macro-block level to optimize the downstream task. This granular control allows us to prioritize encoding for task-relevant regions within each frame. We formulate this optimization problem as a Reinforcement Learning (RL) task, where the agent learns to balance long-term implications of choosing QPs on both task performance and bit-rate constraints. Notably, our policy does not require the downstream task as an input during inference, making it suitable for streaming applications and edge devices such as vehicles. We demonstrate significant improvements in two tasks, car detection, and ROI (saliency) encoding. Our approach improves task performance for a given bit rate compared to traditional task agnostic encoding methods, paving the way for more efficient task-aware video compression.
format Preprint
id arxiv_https___arxiv_org_abs_2501_12216
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RL-RC-DoT: A Block-level RL agent for Task-Aware Video Compression
Gadot, Uri
Shocher, Assaf
Mannor, Shie
Chechik, Gal
Hallak, Assaf
Machine Learning
Computer Vision and Pattern Recognition
Image and Video Processing
Video encoders optimize compression for human perception by minimizing reconstruction error under bit-rate constraints. In many modern applications such as autonomous driving, an overwhelming majority of videos serve as input for AI systems performing tasks like object recognition or segmentation, rather than being watched by humans. It is therefore useful to optimize the encoder for a downstream task instead of for perceptual image quality. However, a major challenge is how to combine such downstream optimization with existing standard video encoders, which are highly efficient and popular. Here, we address this challenge by controlling the Quantization Parameters (QPs) at the macro-block level to optimize the downstream task. This granular control allows us to prioritize encoding for task-relevant regions within each frame. We formulate this optimization problem as a Reinforcement Learning (RL) task, where the agent learns to balance long-term implications of choosing QPs on both task performance and bit-rate constraints. Notably, our policy does not require the downstream task as an input during inference, making it suitable for streaming applications and edge devices such as vehicles. We demonstrate significant improvements in two tasks, car detection, and ROI (saliency) encoding. Our approach improves task performance for a given bit rate compared to traditional task agnostic encoding methods, paving the way for more efficient task-aware video compression.
title RL-RC-DoT: A Block-level RL agent for Task-Aware Video Compression
topic Machine Learning
Computer Vision and Pattern Recognition
Image and Video Processing
url https://arxiv.org/abs/2501.12216