ML-CLIPSim: Multi-Layer CLIP Similarity for Machine-Oriented Image Quality
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913108174307328 |
|---|---|
| author | Ding, Feng Fu, Haisheng Liang, Jie Xu, Qihan Zhu, Siyu Han, Jingning |
| author_facet | Ding, Feng Fu, Haisheng Liang, Jie Xu, Qihan Zhu, Siyu Han, Jingning |
| contents | We study full-reference image quality assessment from a machine-centric perspective, where images are evaluated by how well they preserve information for downstream models. We formulate machine-oriented quality as a latent machine utility and approximate it through pairwise predictive-consistency comparisons. To this end, we construct PCMP, a dataset of PSNR-matched distortion pairs labeled by consistency votes from multiple pretrained models. We further propose ML-CLIPSim, a differentiable quality metric built on a frozen CLIP visual encoder, which aggregates intermediate patch-token similarities and global image embeddings. Experiments on machine-preference benchmarks, human-IQA datasets, and learned image compression show that ML-CLIPSim better aligns with machine-oriented preferences than conventional fidelity and perceptual metrics, while remaining competitive for human quality prediction. Used as a compression distortion term, it improves rate--task trade-offs across multiple downstream tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_09479 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | ML-CLIPSim: Multi-Layer CLIP Similarity for Machine-Oriented Image Quality Ding, Feng Fu, Haisheng Liang, Jie Xu, Qihan Zhu, Siyu Han, Jingning Image and Video Processing Computer Vision and Pattern Recognition Multimedia We study full-reference image quality assessment from a machine-centric perspective, where images are evaluated by how well they preserve information for downstream models. We formulate machine-oriented quality as a latent machine utility and approximate it through pairwise predictive-consistency comparisons. To this end, we construct PCMP, a dataset of PSNR-matched distortion pairs labeled by consistency votes from multiple pretrained models. We further propose ML-CLIPSim, a differentiable quality metric built on a frozen CLIP visual encoder, which aggregates intermediate patch-token similarities and global image embeddings. Experiments on machine-preference benchmarks, human-IQA datasets, and learned image compression show that ML-CLIPSim better aligns with machine-oriented preferences than conventional fidelity and perceptual metrics, while remaining competitive for human quality prediction. Used as a compression distortion term, it improves rate--task trade-offs across multiple downstream tasks. |
| title | ML-CLIPSim: Multi-Layer CLIP Similarity for Machine-Oriented Image Quality |
| topic | Image and Video Processing Computer Vision and Pattern Recognition Multimedia |
| url | https://arxiv.org/abs/2605.09479 |