Unleashing Vision Transformer Potential In Image Quality Assessment via Global-Local Adaptive Interaction
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Yu, Zhou, Puchao, Mi, Yachun, Wu, Yanfeng, Wang, Xiaoming, Liu, Shaohui |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Q-CLIP: Unleashing the Power of Vision-Language Models for Video Quality Assessment through Unified Cross-Modal Adaptation
by: Mi, Yachun, et al.
Published: (2025)
by: Mi, Yachun, et al.
Published: (2025)
MVQA: Mamba with Unified Sampling for Efficient Video Quality Assessment
by: Mi, Yachun, et al.
Published: (2025)
by: Mi, Yachun, et al.
Published: (2025)
Segmenting and Understanding: Region-aware Semantic Attention for Fine-grained Image Quality Assessment with Large Language Models
by: Song, Chenyue, et al.
Published: (2025)
by: Song, Chenyue, et al.
Published: (2025)
UGD-IML: A Unified Generative Diffusion-based Framework for Constrained and Unconstrained Image Manipulation Localization
by: Mi, Yachun, et al.
Published: (2025)
by: Mi, Yachun, et al.
Published: (2025)
Youtu-VL: Unleashing Visual Potential via Unified Vision-Language Supervision
by: Wei, Zhixiang, et al.
Published: (2026)
by: Wei, Zhixiang, et al.
Published: (2026)
MAP: Unleashing Hybrid Mamba-Transformer Vision Backbone's Potential with Masked Autoregressive Pretraining
by: Liu, Yunze, et al.
Published: (2024)
by: Liu, Yunze, et al.
Published: (2024)
Progressive Retinal Image Registration via Global and Local Deformable Transformations
by: Liu, Yepeng, et al.
Published: (2024)
by: Liu, Yepeng, et al.
Published: (2024)
EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editing
by: Chow, Wei, et al.
Published: (2025)
by: Chow, Wei, et al.
Published: (2025)
BPCLIP: A Bottom-up Image Quality Assessment from Distortion to Semantics Based on CLIP
by: Song, Chenyue, et al.
Published: (2025)
by: Song, Chenyue, et al.
Published: (2025)
TwinBrainVLA: Unleashing the Potential of Generalist VLMs for Embodied Tasks via Asymmetric Mixture-of-Transformers
by: Yu, Bin, et al.
Published: (2026)
by: Yu, Bin, et al.
Published: (2026)
Unleashing the Potential of SAM for Medical Adaptation via Hierarchical Decoding
by: Cheng, Zhiheng, et al.
Published: (2024)
by: Cheng, Zhiheng, et al.
Published: (2024)
Unleashing the Potential of the Semantic Latent Space in Diffusion Models for Image Dehazing
by: Yang, Zizheng, et al.
Published: (2025)
by: Yang, Zizheng, et al.
Published: (2025)
Revisiting Vision Language Foundations for No-Reference Image Quality Assessment
by: Yadav, Ankit, et al.
Published: (2025)
by: Yadav, Ankit, et al.
Published: (2025)
Vision Language Modeling of Content, Distortion and Appearance for Image Quality Assessment
by: Zhou, Fei, et al.
Published: (2024)
by: Zhou, Fei, et al.
Published: (2024)
ALGM: Adaptive Local-then-Global Token Merging for Efficient Semantic Segmentation with Plain Vision Transformers
by: Norouzi, Narges, et al.
Published: (2024)
by: Norouzi, Narges, et al.
Published: (2024)
ChangeViT: Unleashing Plain Vision Transformers for Change Detection
by: Zhu, Duowang, et al.
Published: (2024)
by: Zhu, Duowang, et al.
Published: (2024)
Unleashing the Potential of Synthetic Images: A Study on Histopathology Image Classification
by: Benito-Del-Valle, Leire, et al.
Published: (2024)
by: Benito-Del-Valle, Leire, et al.
Published: (2024)
Few-Shot Image Quality Assessment via Adaptation of Vision-Language Models
by: Li, Xudong, et al.
Published: (2024)
by: Li, Xudong, et al.
Published: (2024)
Curriculum Group Policy Optimization: Adaptive Sampling for Unleashing the Potential of Text-to-Image Generation
by: Li, Baoteng, et al.
Published: (2026)
by: Li, Baoteng, et al.
Published: (2026)
AgenticIQA: An Agentic Framework for Adaptive and Interpretable Image Quality Assessment
by: Zhu, Hanwei, et al.
Published: (2025)
by: Zhu, Hanwei, et al.
Published: (2025)
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads
by: Li, Yifan, et al.
Published: (2025)
by: Li, Yifan, et al.
Published: (2025)
IML-ViT: Benchmarking Image Manipulation Localization by Vision Transformer
by: Ma, Xiaochen, et al.
Published: (2023)
by: Ma, Xiaochen, et al.
Published: (2023)
Interpretable Image Classification with Adaptive Prototype-based Vision Transformers
by: Ma, Chiyu, et al.
Published: (2024)
by: Ma, Chiyu, et al.
Published: (2024)
Federated Vision Transformer with Adaptive Focal Loss for Medical Image Classification
by: Zhao, Xinyuan, et al.
Published: (2026)
by: Zhao, Xinyuan, et al.
Published: (2026)
Self-Evolving Vision-Language Models for Image Quality Assessment via Voting and Ranking
by: Wen, Wen, et al.
Published: (2025)
by: Wen, Wen, et al.
Published: (2025)
MedQ-UNI: Toward Unified Medical Image Quality Assessment and Restoration via Vision-Language Modeling
by: Liu, Jiyao, et al.
Published: (2026)
by: Liu, Jiyao, et al.
Published: (2026)
Locatability-Guided Adaptive Reasoning for Image Geo-Localization with Vision-Language Models
by: Yu, Bo, et al.
Published: (2026)
by: Yu, Bo, et al.
Published: (2026)
Content-Distortion High-Order Interaction for Blind Image Quality Assessment
by: Liu, Shuai, et al.
Published: (2025)
by: Liu, Shuai, et al.
Published: (2025)
Unleashing the Potential of SAM2 for Biomedical Images and Videos: A Survey
by: Zhang, Yichi, et al.
Published: (2024)
by: Zhang, Yichi, et al.
Published: (2024)
Employing Vision-Language Models for Face Image Quality Assessment
by: Sarıtaş, Erdi, et al.
Published: (2026)
by: Sarıtaş, Erdi, et al.
Published: (2026)
No-Reference Image Quality Assessment with Global-Local Progressive Integration and Semantic-Aligned Quality Transfer
by: Wang, Xiaoqi, et al.
Published: (2024)
by: Wang, Xiaoqi, et al.
Published: (2024)
DA-Mamba: Learning Domain-Aware State Space Model for Global-Local Alignment in Domain Adaptive Object Detection
by: Li, Haochen, et al.
Published: (2026)
by: Li, Haochen, et al.
Published: (2026)
AdaQual-Diff: Diffusion-Based Image Restoration via Adaptive Quality Prompting
by: Su, Xin, et al.
Published: (2025)
by: Su, Xin, et al.
Published: (2025)
ViTNT-FIQA: Training-Free Face Image Quality Assessment with Vision Transformers
by: Ozgur, Guray, et al.
Published: (2026)
by: Ozgur, Guray, et al.
Published: (2026)
Contrastive Local Manifold Learning for No-Reference Image Quality Assessment
by: Huang, Zihao, et al.
Published: (2024)
by: Huang, Zihao, et al.
Published: (2024)
Unleashing VLA Potentials in Autonomous Driving via Explicit Learning from Failures
by: Luo, Yuechen, et al.
Published: (2026)
by: Luo, Yuechen, et al.
Published: (2026)
Lightweight Vision Transformer with Bidirectional Interaction
by: Fan, Qihang, et al.
Published: (2023)
by: Fan, Qihang, et al.
Published: (2023)
Unleashing the Power of Chain-of-Prediction for Monocular 3D Object Detection
by: Zhang, Zhihao, et al.
Published: (2025)
by: Zhang, Zhihao, et al.
Published: (2025)
SEAGULL: No-reference Image Quality Assessment for Regions of Interest via Vision-Language Instruction Tuning
by: Chen, Zewen, et al.
Published: (2024)
by: Chen, Zewen, et al.
Published: (2024)
Unleashing Guidance Without Classifiers for Human-Object Interaction Animation
by: Wang, Ziyin, et al.
Published: (2026)
by: Wang, Ziyin, et al.
Published: (2026)
Similar Items
-
Q-CLIP: Unleashing the Power of Vision-Language Models for Video Quality Assessment through Unified Cross-Modal Adaptation
by: Mi, Yachun, et al.
Published: (2025) -
MVQA: Mamba with Unified Sampling for Efficient Video Quality Assessment
by: Mi, Yachun, et al.
Published: (2025) -
Segmenting and Understanding: Region-aware Semantic Attention for Fine-grained Image Quality Assessment with Large Language Models
by: Song, Chenyue, et al.
Published: (2025) -
UGD-IML: A Unified Generative Diffusion-based Framework for Constrained and Unconstrained Image Manipulation Localization
by: Mi, Yachun, et al.
Published: (2025) -
Youtu-VL: Unleashing Visual Potential via Unified Vision-Language Supervision
by: Wei, Zhixiang, et al.
Published: (2026)