Q-CLIP: Unleashing the Power of Vision-Language Models for Video Quality Assessment through Unified Cross-Modal Adaptation
Fuente:
arXiv
Saved in:
| Main Authors: | Mi, Yachun, Li, Yu, Li, Yanting, Hui, Chen, Zhang, Tong, Li, Zhixuan, Song, Chenyue, Lim, Wei Yang Bryan, Liu, Shaohui |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MVQA: Mamba with Unified Sampling for Efficient Video Quality Assessment
by: Mi, Yachun, et al.
Published: (2025)
by: Mi, Yachun, et al.
Published: (2025)
Unleashing Vision Transformer Potential In Image Quality Assessment via Global-Local Adaptive Interaction
by: Li, Yu, et al.
Published: (2026)
by: Li, Yu, et al.
Published: (2026)
Segmenting and Understanding: Region-aware Semantic Attention for Fine-grained Image Quality Assessment with Large Language Models
by: Song, Chenyue, et al.
Published: (2025)
by: Song, Chenyue, et al.
Published: (2025)
UGD-IML: A Unified Generative Diffusion-based Framework for Constrained and Unconstrained Image Manipulation Localization
by: Mi, Yachun, et al.
Published: (2025)
by: Mi, Yachun, et al.
Published: (2025)
BPCLIP: A Bottom-up Image Quality Assessment from Distortion to Semantics Based on CLIP
by: Song, Chenyue, et al.
Published: (2025)
by: Song, Chenyue, et al.
Published: (2025)
MS-IQA: A Multi-Scale Feature Fusion Network for PET/CT Image Quality Assessment
by: Li, Siqiao, et al.
Published: (2025)
by: Li, Siqiao, et al.
Published: (2025)
CLIPVQA:Video Quality Assessment via CLIP
by: Xing, Fengchuang, et al.
Published: (2024)
by: Xing, Fengchuang, et al.
Published: (2024)
NeuroLIP: Interpretable and Fair Cross-Modal Alignment of fMRI and Phenotypic Text
by: Yang, Yanting, et al.
Published: (2025)
by: Yang, Yanting, et al.
Published: (2025)
LVPNet: A Latent-variable-based Prediction-driven End-to-end Framework for Lossless Compression of Medical Images
by: Song, Chenyue, et al.
Published: (2025)
by: Song, Chenyue, et al.
Published: (2025)
Unleash the Potential of CLIP for Video Highlight Detection
by: Han, Donghoon, et al.
Published: (2024)
by: Han, Donghoon, et al.
Published: (2024)
Unleashing Diverse Thinking Modes in LLMs through Multi-Agent Collaboration
by: He, Zhixuan, et al.
Published: (2025)
by: He, Zhixuan, et al.
Published: (2025)
BriMA: Bridged Modality Adaptation for Multi-Modal Continual Action Quality Assessment
by: Zhou, Kanglei, et al.
Published: (2026)
by: Zhou, Kanglei, et al.
Published: (2026)
KAnoCLIP: Zero-Shot Anomaly Detection through Knowledge-Driven Prompt Learning and Enhanced Cross-Modal Integration
by: Li, Chengyuan, et al.
Published: (2025)
by: Li, Chengyuan, et al.
Published: (2025)
DialCLIP: Empowering CLIP as Multi-Modal Dialog Retriever
by: Yin, Zhichao, et al.
Published: (2024)
by: Yin, Zhichao, et al.
Published: (2024)
DVLTA-VQA: Decoupled Vision-Language Modeling with Text-Guided Adaptation for Blind Video Quality Assessment
by: Yu, Li, et al.
Published: (2025)
by: Yu, Li, et al.
Published: (2025)
CardiacCLIP: Video-based CLIP Adaptation for LVEF Prediction in a Few-shot Manner
by: Du, Yao, et al.
Published: (2025)
by: Du, Yao, et al.
Published: (2025)
Unleashing MLLMs on the Edge: A Unified Framework for Cross-Modal ReID via Adaptive SVD Distillation
by: Jiang, Hongbo, et al.
Published: (2026)
by: Jiang, Hongbo, et al.
Published: (2026)
Cross-Domain Attribute Alignment with CLIP: A Rehearsal-Free Approach for Class-Incremental Unsupervised Domain Adaptation
by: Mi, Kerun, et al.
Published: (2025)
by: Mi, Kerun, et al.
Published: (2025)
Split to Merge: Unifying Separated Modalities for Unsupervised Domain Adaptation
by: Li, Xinyao, et al.
Published: (2024)
by: Li, Xinyao, et al.
Published: (2024)
CLIP-Powered Domain Generalization and Domain Adaptation: A Comprehensive Survey
by: Li, Jindong, et al.
Published: (2025)
by: Li, Jindong, et al.
Published: (2025)
The Global Existence of Martingale Solutions to Stochastic Compressible Navier-Stokes Equations with Density-dependent Viscosity
by: Li, Yachun, et al.
Published: (2022)
by: Li, Yachun, et al.
Published: (2022)
Initial boundary value problem for one-dimensional hyperbolic compressible Navier-Stokes equations
by: Hu, Yuxi, et al.
Published: (2025)
by: Hu, Yuxi, et al.
Published: (2025)
Local classical solutions to Navier-Stokes equations with degenerate viscosities and vacuum
by: Li, Yachun, et al.
Published: (2024)
by: Li, Yachun, et al.
Published: (2024)
Proximity QA: Unleashing the Power of Multi-Modal Large Language Models for Spatial Proximity Analysis
by: Li, Jianing, et al.
Published: (2024)
by: Li, Jianing, et al.
Published: (2024)
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads
by: Li, Yifan, et al.
Published: (2025)
by: Li, Yifan, et al.
Published: (2025)
Few-Shot Image Quality Assessment via Adaptation of Vision-Language Models
by: Li, Xudong, et al.
Published: (2024)
by: Li, Xudong, et al.
Published: (2024)
Image Compression for Machine and Human Vision with Spatial-Frequency Adaptation
by: Li, Han, et al.
Published: (2024)
by: Li, Han, et al.
Published: (2024)
GaLore$+$: Boosting Low-Rank Adaptation for LLMs with Cross-Head Projection
by: Liao, Xutao, et al.
Published: (2024)
by: Liao, Xutao, et al.
Published: (2024)
SynGR: Unleashing the Potential of Cross-Modal Synergy for Generative Recommendation
by: Chen, Wei, et al.
Published: (2026)
by: Chen, Wei, et al.
Published: (2026)
Forensics Adapter: Unleashing CLIP for Generalizable Face Forgery Detection
by: Cui, Xinjie, et al.
Published: (2024)
by: Cui, Xinjie, et al.
Published: (2024)
SEGA: A Transferable Signed Ensemble Gaussian Black-Box Attack against No-Reference Image Quality Assessment Models
by: Liu, Yujia, et al.
Published: (2025)
by: Liu, Yujia, et al.
Published: (2025)
UARE: A Unified Vision-Language Model for Image Quality Assessment, Restoration, and Enhancement
by: Li, Weiqi, et al.
Published: (2025)
by: Li, Weiqi, et al.
Published: (2025)
Prediction and Reference Quality Adaptation for Learned Video Compression
by: Sheng, Xihua, et al.
Published: (2024)
by: Sheng, Xihua, et al.
Published: (2024)
Towards Unified Video Quality Assessment
by: Feng, Chen, et al.
Published: (2025)
by: Feng, Chen, et al.
Published: (2025)
Youtu-VL: Unleashing Visual Potential via Unified Vision-Language Supervision
by: Wei, Zhixiang, et al.
Published: (2026)
by: Wei, Zhixiang, et al.
Published: (2026)
Data-Efficient CLIP-Powered Dual-Branch Networks for Source-Free Unsupervised Domain Adaptation
by: Li, Yongguang, et al.
Published: (2024)
by: Li, Yongguang, et al.
Published: (2024)
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers
by: Liu, Hongbo
Published: (2024)
by: Liu, Hongbo
Published: (2024)
VarParser: Unleashing the Neglected Power of Variables for LLM-based Log Parsing
by: Sun, Jinrui, et al.
Published: (2026)
by: Sun, Jinrui, et al.
Published: (2026)
Mosaic: Cross-Modal Clustering for Efficient Video Understanding
by: Wang, Tuowei, et al.
Published: (2026)
by: Wang, Tuowei, et al.
Published: (2026)
Enhancing Multimodal Sentiment Analysis for Missing Modality through Self-Distillation and Unified Modality Cross-Attention
by: Weng, Yuzhe, et al.
Published: (2024)
by: Weng, Yuzhe, et al.
Published: (2024)
Similar Items
-
MVQA: Mamba with Unified Sampling for Efficient Video Quality Assessment
by: Mi, Yachun, et al.
Published: (2025) -
Unleashing Vision Transformer Potential In Image Quality Assessment via Global-Local Adaptive Interaction
by: Li, Yu, et al.
Published: (2026) -
Segmenting and Understanding: Region-aware Semantic Attention for Fine-grained Image Quality Assessment with Large Language Models
by: Song, Chenyue, et al.
Published: (2025) -
UGD-IML: A Unified Generative Diffusion-based Framework for Constrained and Unconstrained Image Manipulation Localization
by: Mi, Yachun, et al.
Published: (2025) -
BPCLIP: A Bottom-up Image Quality Assessment from Distortion to Semantics Based on CLIP
by: Song, Chenyue, et al.
Published: (2025)