Benchmarking and Enhancing VLM for Compressed Image Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Zifu, Xu, Tongda, Li, Siqi, Li, Shengxi, Zhang, Yue, Xu, Mai, Wang, Yan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Continuous Patch Stitching for Block-wise Image Compression
by: Zhang, Zifu, et al.
Published: (2025)
by: Zhang, Zifu, et al.
Published: (2025)
Hierarchical Semantic Compression for Consistent Image Semantic Restoration
by: Li, Shengxi, et al.
Published: (2025)
by: Li, Shengxi, et al.
Published: (2025)
Machines Serve Human: A Novel Variable Human-machine Collaborative Compression Framework
by: Zhang, Zifu, et al.
Published: (2025)
by: Zhang, Zifu, et al.
Published: (2025)
Noise Dimension of GAN: An Image Compression Perspective
by: Zhu, Ziran, et al.
Published: (2024)
by: Zhu, Ziran, et al.
Published: (2024)
PICD: Versatile Perceptual Image Compression with Diffusion Rendering
by: Xu, Tongda, et al.
Published: (2025)
by: Xu, Tongda, et al.
Published: (2025)
REMAC: Reference-Based Martian Asymmetrical Image Compression
by: Ding, Qing, et al.
Published: (2026)
by: Ding, Qing, et al.
Published: (2026)
RAWIC: Bit-Depth Adaptive Lossless Raw Image Compression
by: Zheng, Chunhang, et al.
Published: (2026)
by: Zheng, Chunhang, et al.
Published: (2026)
GaussianImage++: Boosted Image Representation and Compression with 2D Gaussian Splatting
by: Li, Tiantian, et al.
Published: (2025)
by: Li, Tiantian, et al.
Published: (2025)
Enhancing Quality of Compressed Images by Mitigating Enhancement Bias Towards Compression Domain
by: Xing, Qunliang, et al.
Published: (2024)
by: Xing, Qunliang, et al.
Published: (2024)
Burst Image Quality Assessment: A New Benchmark and Unified Framework for Multiple Downstream Tasks
by: Liang, Xiaoye, et al.
Published: (2025)
by: Liang, Xiaoye, et al.
Published: (2025)
Fast Training-free Perceptual Image Compression
by: Zhu, Ziran, et al.
Published: (2025)
by: Zhu, Ziran, et al.
Published: (2025)
Causal Context Adjustment Loss for Learned Image Compression
by: Han, Minghao, et al.
Published: (2024)
by: Han, Minghao, et al.
Published: (2024)
EMIFF: Enhanced Multi-scale Image Feature Fusion for Vehicle-Infrastructure Cooperative 3D Object Detection
by: Wang, Zhe, et al.
Published: (2024)
by: Wang, Zhe, et al.
Published: (2024)
Optimizing Input of Denoising Score Matching is Biased Towards Higher Score Norm
by: Xu, Tongda
Published: (2025)
by: Xu, Tongda
Published: (2025)
SenseFlow: Scaling Distribution Matching for Flow-based Text-to-Image Distillation
by: Ge, Xingtong, et al.
Published: (2025)
by: Ge, Xingtong, et al.
Published: (2025)
Parallax to Align Them All: An OmniParallax Attention Mechanism for Distributed Multi-View Image Compression
by: Zhang, Haotian, et al.
Published: (2026)
by: Zhang, Haotian, et al.
Published: (2026)
Idempotence and Perceptual Image Compression
by: Xu, Tongda, et al.
Published: (2024)
by: Xu, Tongda, et al.
Published: (2024)
A Visually Impaired Assistance Benchmark for VLM-as-a-Judge Evaluation
by: Zhao, Yi, et al.
Published: (2026)
by: Zhao, Yi, et al.
Published: (2026)
CogVLM2: Visual Language Models for Image and Video Understanding
by: Hong, Wenyi, et al.
Published: (2024)
by: Hong, Wenyi, et al.
Published: (2024)
Versatile Recompression-Aware Perceptual Image Super-Resolution
by: He, Mingwei, et al.
Published: (2025)
by: He, Mingwei, et al.
Published: (2025)
Advancing Applications of Satellite Photogrammetry: Novel Approaches for Built-up Area Modeling and Natural Environment Monitoring using Stereo/Multi-view Satellite Image-derived 3D Data
by: Gui, Shengxi
Published: (2024)
by: Gui, Shengxi
Published: (2024)
MedSG-Bench: A Benchmark for Medical Image Sequences Grounding
by: Yue, Jingkun, et al.
Published: (2025)
by: Yue, Jingkun, et al.
Published: (2025)
CAMSIC: Content-aware Masked Image Modeling Transformer for Stereo Image Compression
by: Zhang, Xinjie, et al.
Published: (2024)
by: Zhang, Xinjie, et al.
Published: (2024)
GaussianImage: 1000 FPS Image Representation and Compression by 2D Gaussian Splatting
by: Zhang, Xinjie, et al.
Published: (2024)
by: Zhang, Xinjie, et al.
Published: (2024)
V-Shuffle: Zero-Shot Style Transfer via Value Shuffle
by: Tang, Haojun, et al.
Published: (2025)
by: Tang, Haojun, et al.
Published: (2025)
KFFocus: Highlighting Keyframes for Enhanced Video Understanding
by: Nie, Ming, et al.
Published: (2025)
by: Nie, Ming, et al.
Published: (2025)
VILTA: A VLM-in-the-Loop Adversary for Enhancing Driving Policy Robustness
by: Chen, Qimao, et al.
Published: (2026)
by: Chen, Qimao, et al.
Published: (2026)
Training-Free Image Editing with Visual Context Integration and Concept Alignment
by: Song, Rui, et al.
Published: (2026)
by: Song, Rui, et al.
Published: (2026)
CoopDETR: A Unified Cooperative Perception Framework for 3D Detection via Object Query
by: Wang, Zhe, et al.
Published: (2025)
by: Wang, Zhe, et al.
Published: (2025)
LVC: A Lightweight Compression Framework for Enhancing VLMs in Long Video Understanding
by: Wang, Ziyi, et al.
Published: (2025)
by: Wang, Ziyi, et al.
Published: (2025)
ROOT: VLM based System for Indoor Scene Understanding and Beyond
by: Wang, Yonghui, et al.
Published: (2024)
by: Wang, Yonghui, et al.
Published: (2024)
UI2V-Bench: An Understanding-based Image-to-video Generation Benchmark
by: Zhang, Ailing, et al.
Published: (2025)
by: Zhang, Ailing, et al.
Published: (2025)
Amber-Image: Efficient Compression of Large-Scale Diffusion Transformers
by: Yang, Chaojie, et al.
Published: (2026)
by: Yang, Chaojie, et al.
Published: (2026)
Unleashing the Multi-View Fusion Potential: Noise Correction in VLM for Open-Vocabulary 3D Scene Understanding
by: Yin, Xingyilang, et al.
Published: (2025)
by: Yin, Xingyilang, et al.
Published: (2025)
Concept-based Explainable Data Mining with VLM for 3D Detection
by: Tsujimoto, Mai
Published: (2025)
by: Tsujimoto, Mai
Published: (2025)
RoboStressBench: Benchmarking VLM Robustness to Physical Visual Stress in Embodied Scenes
by: Wu, Leyi, et al.
Published: (2026)
by: Wu, Leyi, et al.
Published: (2026)
RelationVLM: Making Large Vision-Language Models Understand Visual Relations
by: Huang, Zhipeng, et al.
Published: (2024)
by: Huang, Zhipeng, et al.
Published: (2024)
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM
by: Wang, Han, et al.
Published: (2024)
by: Wang, Han, et al.
Published: (2024)
BackdoorVLM: A Benchmark for Backdoor Attacks on Vision-Language Models
by: Li, Juncheng, et al.
Published: (2025)
by: Li, Juncheng, et al.
Published: (2025)
Enhancing Video Transformers for Action Understanding with VLM-aided Training
by: Lu, Hui, et al.
Published: (2024)
by: Lu, Hui, et al.
Published: (2024)
Similar Items
-
Continuous Patch Stitching for Block-wise Image Compression
by: Zhang, Zifu, et al.
Published: (2025) -
Hierarchical Semantic Compression for Consistent Image Semantic Restoration
by: Li, Shengxi, et al.
Published: (2025) -
Machines Serve Human: A Novel Variable Human-machine Collaborative Compression Framework
by: Zhang, Zifu, et al.
Published: (2025) -
Noise Dimension of GAN: An Image Compression Perspective
by: Zhu, Ziran, et al.
Published: (2024) -
PICD: Versatile Perceptual Image Compression with Diffusion Rendering
by: Xu, Tongda, et al.
Published: (2025)