Parameter-Efficient Adaptation of mPLUG-Owl2 via Pixel-Level Visual Prompts for NR-IQA
Fuente:
arXiv
Salvato in:
| Autori principali: | Benmahane, Yahya, Hassouni, Mohammed El |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
di: Ye, Qinghao, et al.
Pubblicazione: (2023)
di: Ye, Qinghao, et al.
Pubblicazione: (2023)
mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding
di: Hu, Anwen, et al.
Pubblicazione: (2024)
di: Hu, Anwen, et al.
Pubblicazione: (2024)
mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding
di: Hu, Anwen, et al.
Pubblicazione: (2024)
di: Hu, Anwen, et al.
Pubblicazione: (2024)
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
di: Ye, Jiabo, et al.
Pubblicazione: (2024)
di: Ye, Jiabo, et al.
Pubblicazione: (2024)
DynSTG-Mamba: Dynamic Spatio-Temporal Graph Mamba with Cross-Graph Knowledge Distillation for Gait Disorders Recognition
di: Zrimek, Zakariae, et al.
Pubblicazione: (2025)
di: Zrimek, Zakariae, et al.
Pubblicazione: (2025)
Point Cloud Quality Assessment Using the Perceptual Clustering Weighted Graph (PCW-Graph) and Attention Fusion Network
di: Laazoufi, Abdelouahed, et al.
Pubblicazione: (2025)
di: Laazoufi, Abdelouahed, et al.
Pubblicazione: (2025)
SafePLUG: Empowering Multimodal LLMs with Pixel-Level Insight and Temporal Grounding for Traffic Accident Understanding
di: Sheng, Zihao, et al.
Pubblicazione: (2025)
di: Sheng, Zihao, et al.
Pubblicazione: (2025)
mPLUG-PaperOwl: Scientific Diagram Analysis with the Multimodal Large Language Model
di: Hu, Anwen, et al.
Pubblicazione: (2023)
di: Hu, Anwen, et al.
Pubblicazione: (2023)
AdaptPrompt: Parameter-Efficient Adaptation of VLMs for Generalizable Deepfake Detection
di: Jiang, Yichen, et al.
Pubblicazione: (2025)
di: Jiang, Yichen, et al.
Pubblicazione: (2025)
PromptIQA: Boosting the Performance and Generalization for No-Reference Image Quality Assessment via Prompts
di: Chen, Zewen, et al.
Pubblicazione: (2024)
di: Chen, Zewen, et al.
Pubblicazione: (2024)
CAP-IQA: Context-Aware Prompt-Guided CT Image Quality Assessment
di: Rifa, Kazi Ramisa, et al.
Pubblicazione: (2026)
di: Rifa, Kazi Ramisa, et al.
Pubblicazione: (2026)
PLUG: Revisiting Amodal Segmentation with Foundation Model and Hierarchical Focus
di: Liu, Zhaochen, et al.
Pubblicazione: (2024)
di: Liu, Zhaochen, et al.
Pubblicazione: (2024)
Parameter choices in HaarPSI for IQA with medical images
di: Karner, Clemens, et al.
Pubblicazione: (2024)
di: Karner, Clemens, et al.
Pubblicazione: (2024)
Time-, Memory- and Parameter-Efficient Visual Adaptation
di: Mercea, Otniel-Bogdan, et al.
Pubblicazione: (2024)
di: Mercea, Otniel-Bogdan, et al.
Pubblicazione: (2024)
Banana100: Breaking NR-IQA Metrics by 100 Iterative Image Replications with Nano Banana Pro
di: Tang, Kenan, et al.
Pubblicazione: (2026)
di: Tang, Kenan, et al.
Pubblicazione: (2026)
PixelPrune: Pixel-Level Adaptive Visual Token Reduction via Predictive Coding
di: Wang, Nan, et al.
Pubblicazione: (2026)
di: Wang, Nan, et al.
Pubblicazione: (2026)
OwlSight: A Robust Illumination Adaptation Framework for Dark Video Human Action Recognition
di: Cheng, Shihao, et al.
Pubblicazione: (2025)
di: Cheng, Shihao, et al.
Pubblicazione: (2025)
LoR-VP: Low-Rank Visual Prompting for Efficient Vision Model Adaptation
di: Jin, Can, et al.
Pubblicazione: (2025)
di: Jin, Can, et al.
Pubblicazione: (2025)
GeoPixel: Pixel Grounding Large Multimodal Model in Remote Sensing
di: Shabbir, Akashah, et al.
Pubblicazione: (2025)
di: Shabbir, Akashah, et al.
Pubblicazione: (2025)
Aquila-plus: Prompt-Driven Visual-Language Models for Pixel-Level Remote Sensing Image Understanding
di: Lu, Kaixuan
Pubblicazione: (2024)
di: Lu, Kaixuan
Pubblicazione: (2024)
UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning
di: Liu, Ye, et al.
Pubblicazione: (2025)
di: Liu, Ye, et al.
Pubblicazione: (2025)
ME-IQA: Memory-Enhanced Image Quality Assessment via Re-Ranking
di: Fan, Kanglong, et al.
Pubblicazione: (2026)
di: Fan, Kanglong, et al.
Pubblicazione: (2026)
Pixel Intensity Tracking for Remote Respiratory Monitoring: A Study on Indonesian Subject
di: Mujahidan, Muhammad Yahya Ayyashy, et al.
Pubblicazione: (2024)
di: Mujahidan, Muhammad Yahya Ayyashy, et al.
Pubblicazione: (2024)
Pixel-Level Domain Adaptation: A New Perspective for Enhancing Weakly Supervised Semantic Segmentation
di: Du, Ye, et al.
Pubblicazione: (2024)
di: Du, Ye, et al.
Pubblicazione: (2024)
Prompt-Free and Efficient SAM2 Adaptation for Biomedical Semantic Segmentation via Dual Adapters
di: Mitsuoka, Hinako, et al.
Pubblicazione: (2026)
di: Mitsuoka, Hinako, et al.
Pubblicazione: (2026)
Lightweight Pixel Difference Networks for Efficient Visual Representation Learning
di: Su, Zhuo, et al.
Pubblicazione: (2024)
di: Su, Zhuo, et al.
Pubblicazione: (2024)
Learning A Low-Level Vision Generalist via Visual Task Prompt
di: Chen, Xiangyu, et al.
Pubblicazione: (2024)
di: Chen, Xiangyu, et al.
Pubblicazione: (2024)
Weakly Supervised Pixel-Level Annotation with Visual Interpretability
di: Nasir, Basma, et al.
Pubblicazione: (2025)
di: Nasir, Basma, et al.
Pubblicazione: (2025)
Parameter-Efficient Multi-Task Learning via Progressive Task-Specific Adaptation
di: Gangwar, Neeraj, et al.
Pubblicazione: (2025)
di: Gangwar, Neeraj, et al.
Pubblicazione: (2025)
MedIQA: A Scalable Foundation Model for Prompt-Driven Medical Image Quality Assessment
di: Xun, Siyi, et al.
Pubblicazione: (2025)
di: Xun, Siyi, et al.
Pubblicazione: (2025)
Granulon: Awakening Pixel-Level Visual Encoders with Adaptive Multi-Granularity Semantics for MLLM
di: Mao, Junyuan, et al.
Pubblicazione: (2026)
di: Mao, Junyuan, et al.
Pubblicazione: (2026)
iDETEX: Empowering MLLMs for Intelligent DETailed EXplainable IQA
di: Zhao, Zhaoran, et al.
Pubblicazione: (2025)
di: Zhao, Zhaoran, et al.
Pubblicazione: (2025)
GroPrompt: Efficient Grounded Prompting and Adaptation for Referring Video Object Segmentation
di: Lin, Ci-Siang, et al.
Pubblicazione: (2024)
di: Lin, Ci-Siang, et al.
Pubblicazione: (2024)
Enhancing Diffusion-based Restoration Models via Difficulty-Adaptive Reinforcement Learning with IQA Reward
di: Xu, Xiaogang, et al.
Pubblicazione: (2025)
di: Xu, Xiaogang, et al.
Pubblicazione: (2025)
Depth Completion as Parameter-Efficient Test-Time Adaptation
di: Ke, Bingxin, et al.
Pubblicazione: (2026)
di: Ke, Bingxin, et al.
Pubblicazione: (2026)
Poetry in Pixels: Prompt Tuning for Poem Image Generation via Diffusion Models
di: Jamil, Sofia, et al.
Pubblicazione: (2025)
di: Jamil, Sofia, et al.
Pubblicazione: (2025)
Towards Pixel-Level VLM Perception via Simple Points Prediction
di: Song, Tianhui, et al.
Pubblicazione: (2026)
di: Song, Tianhui, et al.
Pubblicazione: (2026)
SALT: Parameter-Efficient Fine-Tuning via Singular Value Adaptation with Low-Rank Transformation
di: Elsayed, Abdelrahman, et al.
Pubblicazione: (2025)
di: Elsayed, Abdelrahman, et al.
Pubblicazione: (2025)
VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos
di: Munasinghe, Shehan, et al.
Pubblicazione: (2024)
di: Munasinghe, Shehan, et al.
Pubblicazione: (2024)
Attention Prompt Tuning: Parameter-efficient Adaptation of Pre-trained Models for Spatiotemporal Modeling
di: Bandara, Wele Gedara Chaminda, et al.
Pubblicazione: (2024)
di: Bandara, Wele Gedara Chaminda, et al.
Pubblicazione: (2024)
Documenti analoghi
-
mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
di: Ye, Qinghao, et al.
Pubblicazione: (2023) -
mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding
di: Hu, Anwen, et al.
Pubblicazione: (2024) -
mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding
di: Hu, Anwen, et al.
Pubblicazione: (2024) -
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
di: Ye, Jiabo, et al.
Pubblicazione: (2024) -
DynSTG-Mamba: Dynamic Spatio-Temporal Graph Mamba with Cross-Graph Knowledge Distillation for Gait Disorders Recognition
di: Zrimek, Zakariae, et al.
Pubblicazione: (2025)