Enhancing Descriptive Image Quality Assessment with A Large-scale Multi-modal Dataset
Fuente:
arXiv
Saved in:
| Main Authors: | You, Zhiyuan, Gu, Jinjin, Cai, Xin, Li, Zheyuan, Zhu, Kaiwen, Dong, Chao, Xue, Tianfan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Depicting Beyond Scores: Advancing Image Quality Assessment through Multi-modal Language Models
by: You, Zhiyuan, et al.
Published: (2023)
by: You, Zhiyuan, et al.
Published: (2023)
Teaching Large Language Models to Regress Accurate Image Quality Scores using Score Distribution
by: You, Zhiyuan, et al.
Published: (2025)
by: You, Zhiyuan, et al.
Published: (2025)
PhotoFramer: Multi-modal Image Composition Instruction
by: You, Zhiyuan, et al.
Published: (2025)
by: You, Zhiyuan, et al.
Published: (2025)
Revisiting the Generalization Problem of Low-level Vision Models Through the Lens of Image Deraining
by: Hu, Jinfan, et al.
Published: (2025)
by: Hu, Jinfan, et al.
Published: (2025)
An Intelligent Agentic System for Complex Image Restoration Problems
by: Zhu, Kaiwen, et al.
Published: (2024)
by: Zhu, Kaiwen, et al.
Published: (2024)
DA-VAE: Plug-in Latent Compression for Diffusion via Detail Alignment
by: Cai, Xin, et al.
Published: (2026)
by: Cai, Xin, et al.
Published: (2026)
PhoCoLens: Photorealistic and Consistent Reconstruction in Lensless Imaging
by: Cai, Xin, et al.
Published: (2024)
by: Cai, Xin, et al.
Published: (2024)
UniCon: Unidirectional Information Flow for Effective Control of Large-Scale Diffusion Models
by: Yu, Fanghua, et al.
Published: (2025)
by: Yu, Fanghua, et al.
Published: (2025)
Interpreting Low-level Vision Models with Causal Effect Maps
by: Hu, Jinfan, et al.
Published: (2024)
by: Hu, Jinfan, et al.
Published: (2024)
UltraFusion: Ultra High Dynamic Imaging using Exposure Fusion
by: Chen, Zixuan, et al.
Published: (2025)
by: Chen, Zixuan, et al.
Published: (2025)
How far have we gone in Generative Image Restoration? A study on its capability, limitations and evaluation practices
by: Yin, Xiang, et al.
Published: (2026)
by: Yin, Xiang, et al.
Published: (2026)
Q-Doc: Benchmarking Document Image Quality Assessment Capabilities in Multi-modal Large Language Models
by: Huang, Jiaxi, et al.
Published: (2025)
by: Huang, Jiaxi, et al.
Published: (2025)
Harnessing Diffusion-Yielded Score Priors for Image Restoration
by: Lin, Xinqi, et al.
Published: (2025)
by: Lin, Xinqi, et al.
Published: (2025)
Large Multi-modality Model Assisted AI-Generated Image Quality Assessment
by: Wang, Puyi, et al.
Published: (2024)
by: Wang, Puyi, et al.
Published: (2024)
Reasoning as Representation: Rethinking Visual Reinforcement Learning in Image Quality Assessment
by: Zhao, Shijie, et al.
Published: (2025)
by: Zhao, Shijie, et al.
Published: (2025)
HDRFlow: Real-Time HDR Video Reconstruction with Large Motions
by: Xu, Gangwei, et al.
Published: (2024)
by: Xu, Gangwei, et al.
Published: (2024)
LM4LV: A Frozen Large Language Model for Low-level Vision Tasks
by: Zheng, Boyang, et al.
Published: (2024)
by: Zheng, Boyang, et al.
Published: (2024)
OmniScience: A Large-scale Multi-modal Dataset for Scientific Image Understanding
by: Tao, Haoyi, et al.
Published: (2026)
by: Tao, Haoyi, et al.
Published: (2026)
Position: Agentic Systems Constitute a Key Component of Next-Generation Intelligent Image Processing
by: Gu, Jinjin
Published: (2025)
by: Gu, Jinjin
Published: (2025)
Scaling Up to Excellence: Practicing Model Scaling for Photo-Realistic Image Restoration In the Wild
by: Yu, Fanghua, et al.
Published: (2024)
by: Yu, Fanghua, et al.
Published: (2024)
GAIA: A Global, Multi-modal, Multi-scale Vision-Language Dataset for Remote Sensing Image Analysis
by: Zavras, Angelos, et al.
Published: (2025)
by: Zavras, Angelos, et al.
Published: (2025)
Position: Evaluation of Visual Processing Should Be Human-Centered, Not Metric-Centered
by: Hu, Jinfan, et al.
Published: (2026)
by: Hu, Jinfan, et al.
Published: (2026)
AutoDIR: Automatic All-in-One Image Restoration with Latent Diffusion
by: Jiang, Yitong, et al.
Published: (2023)
by: Jiang, Yitong, et al.
Published: (2023)
PhotoAgent: Agentic Photo Editing with Exploratory Visual Aesthetic Planning
by: Yao, Mingde, et al.
Published: (2026)
by: Yao, Mingde, et al.
Published: (2026)
A Multi-annotated and Multi-modal Dataset for Wide-angle Video Quality Assessment
by: Hu, Bo, et al.
Published: (2025)
by: Hu, Bo, et al.
Published: (2025)
S2R-HDR: A Large-Scale Rendered Dataset for HDR Fusion
by: Wang, Yujin, et al.
Published: (2025)
by: Wang, Yujin, et al.
Published: (2025)
LenslessFace: An End-to-End Optimized Lensless System for Privacy-Preserving Face Verification
by: Cai, Xin, et al.
Published: (2024)
by: Cai, Xin, et al.
Published: (2024)
Q-Ground: Image Quality Grounding with Large Multi-modality Models
by: Chen, Chaofeng, et al.
Published: (2024)
by: Chen, Chaofeng, et al.
Published: (2024)
Image Quality Assessment for Machines: Paradigm, Large-scale Database, and Models
by: Wang, Xiaoqi, et al.
Published: (2025)
by: Wang, Xiaoqi, et al.
Published: (2025)
A Physics-Informed Blur Learning Framework for Imaging Systems
by: Chen, Liqun, et al.
Published: (2025)
by: Chen, Liqun, et al.
Published: (2025)
AdaptiveISP: Learning an Adaptive Image Signal Processor for Object Detection
by: Wang, Yujin, et al.
Published: (2024)
by: Wang, Yujin, et al.
Published: (2024)
Multi-V2X: A Large Scale Multi-modal Multi-penetration-rate Dataset for Cooperative Perception
by: Li, Rongsong, et al.
Published: (2024)
by: Li, Rongsong, et al.
Published: (2024)
DeQA-Doc: Adapting DeQA-Score to Document Image Quality Assessment
by: Gao, Junjie, et al.
Published: (2025)
by: Gao, Junjie, et al.
Published: (2025)
Segmentation Quality and Volumetric Accuracy in Medical Imaging
by: Zhang, Zheyuan, et al.
Published: (2024)
by: Zhang, Zheyuan, et al.
Published: (2024)
Candidate Set Re-ranking for Composed Image Retrieval with Dual Multi-modal Encoder
by: Liu, Zheyuan, et al.
Published: (2023)
by: Liu, Zheyuan, et al.
Published: (2023)
Multi-modal Learnable Queries for Image Aesthetics Assessment
by: Xiong, Zhiwei, et al.
Published: (2024)
by: Xiong, Zhiwei, et al.
Published: (2024)
Enhanced Masked Image Modeling to Avoid Model Collapse on Multi-modal MRI Datasets
by: Han, Linxuan, et al.
Published: (2024)
by: Han, Linxuan, et al.
Published: (2024)
Accelerating Masked Image Generation by Learning Latent Controlled Dynamics
by: Zhu, Kaiwen, et al.
Published: (2026)
by: Zhu, Kaiwen, et al.
Published: (2026)
MoniRefer: A Real-world Large-scale Multi-modal Dataset based on Roadside Infrastructure for 3D Visual Grounding
by: Yang, Panquan, et al.
Published: (2025)
by: Yang, Panquan, et al.
Published: (2025)
A Preliminary Exploration Towards General Image Restoration
by: Kong, Xiangtao, et al.
Published: (2024)
by: Kong, Xiangtao, et al.
Published: (2024)
Similar Items
-
Depicting Beyond Scores: Advancing Image Quality Assessment through Multi-modal Language Models
by: You, Zhiyuan, et al.
Published: (2023) -
Teaching Large Language Models to Regress Accurate Image Quality Scores using Score Distribution
by: You, Zhiyuan, et al.
Published: (2025) -
PhotoFramer: Multi-modal Image Composition Instruction
by: You, Zhiyuan, et al.
Published: (2025) -
Revisiting the Generalization Problem of Low-level Vision Models Through the Lens of Image Deraining
by: Hu, Jinfan, et al.
Published: (2025) -
An Intelligent Agentic System for Complex Image Restoration Problems
by: Zhu, Kaiwen, et al.
Published: (2024)