Machines Serve Human: A Novel Variable Human-machine Collaborative Compression Framework

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Zifu, Li, Shengxi, Sun, Xiancheng, Xu, Mai, Liu, Zhengyuan, Xia, Jingyuan
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911261289086976
author Zhang, Zifu
Li, Shengxi
Sun, Xiancheng
Xu, Mai
Liu, Zhengyuan
Xia, Jingyuan
author_facet Zhang, Zifu
Li, Shengxi
Sun, Xiancheng
Xu, Mai
Liu, Zhengyuan
Xia, Jingyuan
contents Human-machine collaborative compression has been receiving increasing research efforts for reducing image/video data, serving as the basis for both human perception and machine intelligence. Existing collaborative methods are dominantly built upon the de facto human-vision compression pipeline, witnessing deficiency on complexity and bit-rates when aggregating the machine-vision compression. Indeed, machine vision solely focuses on the core regions within the image/video, requiring much less information compared with the compressed information for human vision. In this paper, we thus set out the first successful attempt by a novel collaborative compression method based on the machine-vision-oriented compression, instead of human-vision pipeline. In other words, machine vision serves as the basis for human vision within collaborative compression. A plug-and-play variable bit-rate strategy is also developed for machine vision tasks. Then, we propose to progressively aggregate the semantics from the machine-vision compression, whilst seamlessly tailing the diffusion prior to restore high-fidelity details for human vision, thus named as diffusion-prior based feature compression for human and machine visions (Diff-FCHM). Experimental results verify the consistently superior performances of our Diff-FCHM, on both machine-vision and human-vision compression with remarkable margins. Our code will be released upon acceptance.
format Preprint
id arxiv_https___arxiv_org_abs_2511_08915
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Machines Serve Human: A Novel Variable Human-machine Collaborative Compression Framework
Zhang, Zifu
Li, Shengxi
Sun, Xiancheng
Xu, Mai
Liu, Zhengyuan
Xia, Jingyuan
Computer Vision and Pattern Recognition
Human-machine collaborative compression has been receiving increasing research efforts for reducing image/video data, serving as the basis for both human perception and machine intelligence. Existing collaborative methods are dominantly built upon the de facto human-vision compression pipeline, witnessing deficiency on complexity and bit-rates when aggregating the machine-vision compression. Indeed, machine vision solely focuses on the core regions within the image/video, requiring much less information compared with the compressed information for human vision. In this paper, we thus set out the first successful attempt by a novel collaborative compression method based on the machine-vision-oriented compression, instead of human-vision pipeline. In other words, machine vision serves as the basis for human vision within collaborative compression. A plug-and-play variable bit-rate strategy is also developed for machine vision tasks. Then, we propose to progressively aggregate the semantics from the machine-vision compression, whilst seamlessly tailing the diffusion prior to restore high-fidelity details for human vision, thus named as diffusion-prior based feature compression for human and machine visions (Diff-FCHM). Experimental results verify the consistently superior performances of our Diff-FCHM, on both machine-vision and human-vision compression with remarkable margins. Our code will be released upon acceptance.
title Machines Serve Human: A Novel Variable Human-machine Collaborative Compression Framework
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.08915