New VVC profiles targeting Feature Coding for Machines

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Eimon, Md Eimran Hossain, Perera, Ashan, Merlos, Juan, Adzic, Velibor, Kalva, Hari
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915932779053056
author Eimon, Md Eimran Hossain
Perera, Ashan
Merlos, Juan
Adzic, Velibor
Kalva, Hari
author_facet Eimon, Md Eimran Hossain
Perera, Ashan
Merlos, Juan
Adzic, Velibor
Kalva, Hari
contents Modern video codecs have been extensively optimized to preserve perceptual quality, leveraging models of the human visual system. However, in split inference systems-where intermediate features from neural network are transmitted instead of pixel data-these assumptions no longer apply. Intermediate features are abstract, sparse, and task-specific, making perceptual fidelity irrelevant. In this paper, we investigate the use of Versatile Video Coding (VVC) for compressing such features under the MPEG-AI Feature Coding for Machines (FCM) standard. We perform a tool-level analysis to understand the impact of individual coding components on compression efficiency and downstream vision task accuracy. Based on these insights, we propose three lightweight essential VVC profiles-Fast, Faster, and Fastest. The Fast profile provides 2.96% BD-Rate gain while reducing encoding time by 21.8%. Faster achieves a 1.85% BD-Rate gain with a 51.5% speedup. Fastest reduces encoding time by 95.6% with only a 1.71% loss in BD-Rate.
format Preprint
id arxiv_https___arxiv_org_abs_2512_08227
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle New VVC profiles targeting Feature Coding for Machines
Eimon, Md Eimran Hossain
Perera, Ashan
Merlos, Juan
Adzic, Velibor
Kalva, Hari
Computer Vision and Pattern Recognition
Modern video codecs have been extensively optimized to preserve perceptual quality, leveraging models of the human visual system. However, in split inference systems-where intermediate features from neural network are transmitted instead of pixel data-these assumptions no longer apply. Intermediate features are abstract, sparse, and task-specific, making perceptual fidelity irrelevant. In this paper, we investigate the use of Versatile Video Coding (VVC) for compressing such features under the MPEG-AI Feature Coding for Machines (FCM) standard. We perform a tool-level analysis to understand the impact of individual coding components on compression efficiency and downstream vision task accuracy. Based on these insights, we propose three lightweight essential VVC profiles-Fast, Faster, and Fastest. The Fast profile provides 2.96% BD-Rate gain while reducing encoding time by 21.8%. Faster achieves a 1.85% BD-Rate gain with a 51.5% speedup. Fastest reduces encoding time by 95.6% with only a 1.71% loss in BD-Rate.
title New VVC profiles targeting Feature Coding for Machines
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.08227