Comp-X: On Defining an Interactive Learned Image Compression Paradigm With Expert-driven LLM Agent

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gao, Yixin, Li, Xin, Pan, Xiaohan, Feng, Runsen, Li, Bingchen, Qi, Yunpeng, Lu, Yiting, Cheng, Zhengxue, Chen, Zhibo, Ostermann, Jörn
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916910968340480
author Gao, Yixin
Li, Xin
Pan, Xiaohan
Feng, Runsen
Li, Bingchen
Qi, Yunpeng
Lu, Yiting
Cheng, Zhengxue
Chen, Zhibo
Ostermann, Jörn
author_facet Gao, Yixin
Li, Xin
Pan, Xiaohan
Feng, Runsen
Li, Bingchen
Qi, Yunpeng
Lu, Yiting
Cheng, Zhengxue
Chen, Zhibo
Ostermann, Jörn
contents We present Comp-X, the first intelligently interactive image compression paradigm empowered by the impressive reasoning capability of large language model (LLM) agent. Notably, commonly used image codecs usually suffer from limited coding modes and rely on manual mode selection by engineers, making them unfriendly for unprofessional users. To overcome this, we advance the evolution of image coding paradigm by introducing three key innovations: (i) multi-functional coding framework, which unifies different coding modes of various objective/requirements, including human-machine perception, variable coding, and spatial bit allocation, into one framework. (ii) interactive coding agent, where we propose an augmented in-context learning method with coding expert feedback to teach the LLM agent how to understand the coding request, mode selection, and the use of the coding tools. (iii) IIC-bench, the first dedicated benchmark comprising diverse user requests and the corresponding annotations from coding experts, which is systematically designed for intelligently interactive image compression evaluation. Extensive experimental results demonstrate that our proposed Comp-X can understand the coding requests efficiently and achieve impressive textual interaction capability. Meanwhile, it can maintain comparable compression performance even with a single coding framework, providing a promising avenue for artificial general intelligence (AGI) in image compression.
format Preprint
id arxiv_https___arxiv_org_abs_2508_15243
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Comp-X: On Defining an Interactive Learned Image Compression Paradigm With Expert-driven LLM Agent
Gao, Yixin
Li, Xin
Pan, Xiaohan
Feng, Runsen
Li, Bingchen
Qi, Yunpeng
Lu, Yiting
Cheng, Zhengxue
Chen, Zhibo
Ostermann, Jörn
Computer Vision and Pattern Recognition
We present Comp-X, the first intelligently interactive image compression paradigm empowered by the impressive reasoning capability of large language model (LLM) agent. Notably, commonly used image codecs usually suffer from limited coding modes and rely on manual mode selection by engineers, making them unfriendly for unprofessional users. To overcome this, we advance the evolution of image coding paradigm by introducing three key innovations: (i) multi-functional coding framework, which unifies different coding modes of various objective/requirements, including human-machine perception, variable coding, and spatial bit allocation, into one framework. (ii) interactive coding agent, where we propose an augmented in-context learning method with coding expert feedback to teach the LLM agent how to understand the coding request, mode selection, and the use of the coding tools. (iii) IIC-bench, the first dedicated benchmark comprising diverse user requests and the corresponding annotations from coding experts, which is systematically designed for intelligently interactive image compression evaluation. Extensive experimental results demonstrate that our proposed Comp-X can understand the coding requests efficiently and achieve impressive textual interaction capability. Meanwhile, it can maintain comparable compression performance even with a single coding framework, providing a promising avenue for artificial general intelligence (AGI) in image compression.
title Comp-X: On Defining an Interactive Learned Image Compression Paradigm With Expert-driven LLM Agent
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.15243