VLIC: Vision-Language Models As Perceptual Judges for Human-Aligned Image Compression
Fuente:
arXiv
Saved in:
| Main Authors: | Sargent, Kyle, Gao, Ruiqi, Henzler, Philipp, Herrmann, Charles, Holynski, Aleksander, Fei-Fei, Li, Wu, Jiajun, Zhang, Jason |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Flow to the Mode: Mode-Seeking Diffusion Autoencoders for State-of-the-Art Image Tokenization
by: Sargent, Kyle, et al.
Published: (2025)
by: Sargent, Kyle, et al.
Published: (2025)
CAT3D: Create Anything in 3D with Multi-View Diffusion Models
by: Gao, Ruiqi, et al.
Published: (2024)
by: Gao, Ruiqi, et al.
Published: (2024)
Bolt3D: Generating 3D Scenes in Seconds
by: Szymanowicz, Stanislaw, et al.
Published: (2025)
by: Szymanowicz, Stanislaw, et al.
Published: (2025)
ZeroNVS: Zero-Shot 360-Degree View Synthesis from a Single Image
by: Sargent, Kyle, et al.
Published: (2023)
by: Sargent, Kyle, et al.
Published: (2023)
Generative Image Dynamics
by: Li, Zhengqi, et al.
Published: (2023)
by: Li, Zhengqi, et al.
Published: (2023)
CAT4D: Create Anything in 4D with Multi-View Video Diffusion Models
by: Wu, Rundi, et al.
Published: (2024)
by: Wu, Rundi, et al.
Published: (2024)
UFO-4D: Unposed Feedforward 4D Reconstruction from Two Images
by: Hur, Junhwa, et al.
Published: (2026)
by: Hur, Junhwa, et al.
Published: (2026)
CityRAG: Stepping Into a City via Spatially-Grounded Video Generation
by: Chou, Gene, et al.
Published: (2026)
by: Chou, Gene, et al.
Published: (2026)
SimVS: Simulating World Inconsistencies for Robust View Synthesis
by: Trevithick, Alex, et al.
Published: (2024)
by: Trevithick, Alex, et al.
Published: (2024)
GPIC: A Giant Permissive Image Corpus for Visual Generation
by: Chandrasegaran, Keshigeyan, et al.
Published: (2026)
by: Chandrasegaran, Keshigeyan, et al.
Published: (2026)
ZipMap: Linear-Time Stateful 3D Reconstruction via Test-Time Training
by: Jin, Haian, et al.
Published: (2026)
by: Jin, Haian, et al.
Published: (2026)
Dual-Process Image Generation
by: Luo, Grace, et al.
Published: (2025)
by: Luo, Grace, et al.
Published: (2025)
TIQA: Human-Aligned Perceptual Text Quality Assessment in Generated Images
by: Koltsov, Kirill, et al.
Published: (2026)
by: Koltsov, Kirill, et al.
Published: (2026)
GPS as a Control Signal for Image Generation
by: Feng, Chao, et al.
Published: (2025)
by: Feng, Chao, et al.
Published: (2025)
Latent Forcing: Reordering the Diffusion Trajectory for Pixel-Space Image Generation
by: Baade, Alan, et al.
Published: (2026)
by: Baade, Alan, et al.
Published: (2026)
Vision-Language Models vs Human: Perceptual Image Quality Assessment
by: Mehmood, Imran, et al.
Published: (2026)
by: Mehmood, Imran, et al.
Published: (2026)
Idempotence and Perceptual Image Compression
by: Xu, Tongda, et al.
Published: (2024)
by: Xu, Tongda, et al.
Published: (2024)
One-Step Diffusion for Perceptual Image Compression
by: Jia, Yiwen, et al.
Published: (2026)
by: Jia, Yiwen, et al.
Published: (2026)
Omni-Judge: Can Omni-LLMs Serve as Human-Aligned Judges for Text-Conditioned Audio-Video Generation?
by: Liang, Susan, et al.
Published: (2026)
by: Liang, Susan, et al.
Published: (2026)
DiT-IC: Aligned Diffusion Transformer for Efficient Image Compression
by: Shi, Junqi, et al.
Published: (2026)
by: Shi, Junqi, et al.
Published: (2026)
Generative Inbetweening: Adapting Image-to-Video Models for Keyframe Interpolation
by: Wang, Xiaojuan, et al.
Published: (2024)
by: Wang, Xiaojuan, et al.
Published: (2024)
Infinite Texture: Text-guided High Resolution Diffusion Texture Synthesis
by: Wang, Yifan, et al.
Published: (2024)
by: Wang, Yifan, et al.
Published: (2024)
Derm1M: A Million-scale Vision-Language Dataset Aligned with Clinical Ontology Knowledge for Dermatology
by: Yan, Siyuan, et al.
Published: (2025)
by: Yan, Siyuan, et al.
Published: (2025)
PICD: Versatile Perceptual Image Compression with Diffusion Rendering
by: Xu, Tongda, et al.
Published: (2025)
by: Xu, Tongda, et al.
Published: (2025)
Continual Retinal Vision-Language Pre-training upon Incremental Imaging Modalities
by: Yao, Yuang, et al.
Published: (2025)
by: Yao, Yuang, et al.
Published: (2025)
A Multimodal Recaptioning Framework to Account for Perceptual Diversity Across Languages in Vision-Language Modeling
by: Buettner, Kyle, et al.
Published: (2025)
by: Buettner, Kyle, et al.
Published: (2025)
How Animals Dance (When You're Not Looking)
by: Wang, Xiaojuan, et al.
Published: (2025)
by: Wang, Xiaojuan, et al.
Published: (2025)
Physically Grounded Vision-Language Models for Robotic Manipulation
by: Gao, Jensen, et al.
Published: (2023)
by: Gao, Jensen, et al.
Published: (2023)
WonderJourney: Going from Anywhere to Everywhere
by: Yu, Hong-Xing, et al.
Published: (2023)
by: Yu, Hong-Xing, et al.
Published: (2023)
Readout Guidance: Learning Control from Diffusion Features
by: Luo, Grace, et al.
Published: (2023)
by: Luo, Grace, et al.
Published: (2023)
Continuous 3D Perception Model with Persistent State
by: Wang, Qianqian, et al.
Published: (2025)
by: Wang, Qianqian, et al.
Published: (2025)
Diffusion Hyperfeatures: Searching Through Time and Space for Semantic Correspondence
by: Luo, Grace, et al.
Published: (2023)
by: Luo, Grace, et al.
Published: (2023)
WonderWorld: Interactive 3D Scene Generation from a Single Image
by: Yu, Hong-Xing, et al.
Published: (2024)
by: Yu, Hong-Xing, et al.
Published: (2024)
Categorical Knowledge Fused Recognition: Fusing Hierarchical Knowledge with Image Classification through Aligning and Deep Metric Learning
by: Zhao, Yunfeng, et al.
Published: (2024)
by: Zhao, Yunfeng, et al.
Published: (2024)
Ultra-Low Bitrate Perceptual Image Compression with Shallow Encoder
by: Zhang, Tianyu, et al.
Published: (2025)
by: Zhang, Tianyu, et al.
Published: (2025)
Fast Training-free Perceptual Image Compression
by: Zhu, Ziran, et al.
Published: (2025)
by: Zhu, Ziran, et al.
Published: (2025)
Diff-ICMH: Harmonizing Machine and Human Vision in Image Compression with Generative Prior
by: Feng, Ruoyu, et al.
Published: (2025)
by: Feng, Ruoyu, et al.
Published: (2025)
Aligning What EEG Can See: Structural Representations for Brain-Vision Matching
by: Tang, Jingyi, et al.
Published: (2026)
by: Tang, Jingyi, et al.
Published: (2026)
GFix: Perceptually Enhanced Gaussian Splatting Video Compression
by: Teng, Siyue, et al.
Published: (2025)
by: Teng, Siyue, et al.
Published: (2025)
UniMIC: Towards Universal Multi-modality Perceptual Image Compression
by: Gao, Yixin, et al.
Published: (2024)
by: Gao, Yixin, et al.
Published: (2024)
Similar Items
-
Flow to the Mode: Mode-Seeking Diffusion Autoencoders for State-of-the-Art Image Tokenization
by: Sargent, Kyle, et al.
Published: (2025) -
CAT3D: Create Anything in 3D with Multi-View Diffusion Models
by: Gao, Ruiqi, et al.
Published: (2024) -
Bolt3D: Generating 3D Scenes in Seconds
by: Szymanowicz, Stanislaw, et al.
Published: (2025) -
ZeroNVS: Zero-Shot 360-Degree View Synthesis from a Single Image
by: Sargent, Kyle, et al.
Published: (2023) -
Generative Image Dynamics
by: Li, Zhengqi, et al.
Published: (2023)