NEARL-CLIP: Interacted Query Adaptation with Orthogonal Regularization for Medical Vision-Language Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Peng, Zelin, Zhao, Yichen, Huang, Yu, Yang, Piao, Tang, Feilong, Xu, Zhengqin, Yang, Xiaokang, Shen, Wei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MedSeg-R: Reasoning Segmentation in Medical Images with Multimodal Large Language Models
by: Huang, Yu, et al.
Published: (2025)
by: Huang, Yu, et al.
Published: (2025)
HyperET: Efficient Training in Hyperbolic Space for Multi-modal Large Language Models
by: Peng, Zelin, et al.
Published: (2025)
by: Peng, Zelin, et al.
Published: (2025)
MMRad-22K: A Structured Multimodal Evidence Dataset for Chest X-ray Report Generation
by: Zhao, Yichen, et al.
Published: (2026)
by: Zhao, Yichen, et al.
Published: (2026)
Parameter-efficient Fine-tuning in Hyperspherical Space for Open-vocabulary Semantic Segmentation
by: Peng, Zelin, et al.
Published: (2024)
by: Peng, Zelin, et al.
Published: (2024)
Vision-Informed Flow Image Super-Resolution with Quaternion Spatial Modeling and Dynamic Flow Convolution
by: Cao, Qinglong, et al.
Published: (2024)
by: Cao, Qinglong, et al.
Published: (2024)
Unlocking 3D Affordance Segmentation with 2D Semantic Knowledge
by: Huang, Yu, et al.
Published: (2025)
by: Huang, Yu, et al.
Published: (2025)
Parameter Efficient Fine-tuning via Cross Block Orchestration for Segment Anything Model
by: Peng, Zelin, et al.
Published: (2023)
by: Peng, Zelin, et al.
Published: (2023)
Group Orthogonalization Regularization For Vision Models Adaptation and Robustness
by: Kurtz, Yoav, et al.
Published: (2023)
by: Kurtz, Yoav, et al.
Published: (2023)
Test-Time Adaptation with CLIP Reward for Zero-Shot Generalization in Vision-Language Models
by: Zhao, Shuai, et al.
Published: (2023)
by: Zhao, Shuai, et al.
Published: (2023)
Efficient Adaptation of Pre-trained Vision Transformer underpinned by Approximately Orthogonal Fine-Tuning Strategy
by: Yang, Yiting, et al.
Published: (2025)
by: Yang, Yiting, et al.
Published: (2025)
Toward Modality Gap: Vision Prototype Learning for Weakly-supervised Semantic Segmentation with CLIP
by: Xu, Zhongxing, et al.
Published: (2024)
by: Xu, Zhongxing, et al.
Published: (2024)
Maintaining Structural Integrity in Parameter Spaces for Parameter Efficient Fine-tuning
by: Si, Chongjie, et al.
Published: (2024)
by: Si, Chongjie, et al.
Published: (2024)
Few-shot Adaptation of Medical Vision-Language Models
by: Shakeri, Fereshteh, et al.
Published: (2024)
by: Shakeri, Fereshteh, et al.
Published: (2024)
OphCLIP: Hierarchical Retrieval-Augmented Learning for Ophthalmic Surgical Video-Language Pretraining
by: Hu, Ming, et al.
Published: (2024)
by: Hu, Ming, et al.
Published: (2024)
HarmoCLIP: Harmonizing Global and Regional Representations in Contrastive Vision-Language Models
by: Zeng, Haoxi, et al.
Published: (2025)
by: Zeng, Haoxi, et al.
Published: (2025)
Exploring CLIP's Dense Knowledge for Weakly Supervised Semantic Segmentation
by: Yang, Zhiwei, et al.
Published: (2025)
by: Yang, Zhiwei, et al.
Published: (2025)
Enhancing Robustness of Vision-Language Models through Orthogonality Learning and Self-Regularization
by: Li, Jinlong, et al.
Published: (2024)
by: Li, Jinlong, et al.
Published: (2024)
A Medical Multimodal Diagnostic Framework Integrating Vision-Language Models and Logic Tree Reasoning
by: Zang, Zelin, et al.
Published: (2025)
by: Zang, Zelin, et al.
Published: (2025)
The Generation of All Regular Rational Orthogonal Matrices
by: Tang, Quanyu, et al.
Published: (2024)
by: Tang, Quanyu, et al.
Published: (2024)
CLIP-Mamba: CLIP Pretrained Mamba Models with OOD and Hessian Evaluation
by: Huang, Weiquan, et al.
Published: (2024)
by: Huang, Weiquan, et al.
Published: (2024)
QueryAdapter: Rapid Adaptation of Vision-Language Models in Response to Natural Language Queries
by: Chapman, Nicolas Harvey, et al.
Published: (2025)
by: Chapman, Nicolas Harvey, et al.
Published: (2025)
QwenCLIP: Boosting Medical Vision-Language Pretraining via LLM Embeddings and Prompt tuning
by: Wei, Xiaoyang, et al.
Published: (2025)
by: Wei, Xiaoyang, et al.
Published: (2025)
MedProbCLIP: Probabilistic Adaptation of Vision-Language Foundation Model for Reliable Radiograph-Report Retrieval
by: Elallaf, Ahmad, et al.
Published: (2026)
by: Elallaf, Ahmad, et al.
Published: (2026)
MaskedCLIP: Bridging the Masked and CLIP Space for Semi-Supervised Medical Vision-Language Pre-training
by: Zhu, Lei, et al.
Published: (2025)
by: Zhu, Lei, et al.
Published: (2025)
SORSA: Singular Values and Orthonormal Regularized Singular Vectors Adaptation of Large Language Models
by: Cao, Yang, et al.
Published: (2024)
by: Cao, Yang, et al.
Published: (2024)
QAgent: A modular Search Agent with Interactive Query Understanding
by: Jiang, Yi, et al.
Published: (2025)
by: Jiang, Yi, et al.
Published: (2025)
Q-CLIP: Unleashing the Power of Vision-Language Models for Video Quality Assessment through Unified Cross-Modal Adaptation
by: Mi, Yachun, et al.
Published: (2025)
by: Mi, Yachun, et al.
Published: (2025)
MAP: Revisiting Weight Decomposition for Low-Rank Adaptation
by: Si, Chongjie, et al.
Published: (2025)
by: Si, Chongjie, et al.
Published: (2025)
Enhancing Graph Query Generation for Medical Text by Aligning Domain Knowledge Graph with Large Language Models
by: Zhen Xu, et al.
Published: (2026)
by: Zhen Xu, et al.
Published: (2026)
Generalized Tensor-based Parameter-Efficient Fine-Tuning via Lie Group Transformations
by: Si, Chongjie, et al.
Published: (2025)
by: Si, Chongjie, et al.
Published: (2025)
CLIP-SVD: Efficient and Interpretable Vision-Language Adaptation via Singular Values
by: Koleilat, Taha, et al.
Published: (2025)
by: Koleilat, Taha, et al.
Published: (2025)
Bridging The Gap between Low-rank and Orthogonal Adaptation via Householder Reflection Adaptation
by: Yuan, Shen, et al.
Published: (2024)
by: Yuan, Shen, et al.
Published: (2024)
PG-SAM: Prior-Guided SAM with Medical for Multi-organ Segmentation
by: Zhong, Yiheng, et al.
Published: (2025)
by: Zhong, Yiheng, et al.
Published: (2025)
Mean velocity profiles and total shear stress profiles in adverse-pressure-gradient turbulent boundary layers considering history effect
by: Shu, Zhengqin, et al.
Published: (2025)
by: Shu, Zhengqin, et al.
Published: (2025)
Category Query Learning for Human-Object Interaction Classification
by: Xie, Chi, et al.
Published: (2023)
by: Xie, Chi, et al.
Published: (2023)
ODCR: Orthogonal Decoupling Contrastive Regularization for Unpaired Image Dehazing
by: Wang, Zhongze, et al.
Published: (2024)
by: Wang, Zhongze, et al.
Published: (2024)
TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives
by: Patel, Maitreya, et al.
Published: (2024)
by: Patel, Maitreya, et al.
Published: (2024)
Unifying 3D Vision-Language Understanding via Promptable Queries
by: Zhu, Ziyu, et al.
Published: (2024)
by: Zhu, Ziyu, et al.
Published: (2024)
Weight Spectra Induced Efficient Model Adaptation
by: Si, Chongjie, et al.
Published: (2025)
by: Si, Chongjie, et al.
Published: (2025)
See Further for Parameter Efficient Fine-tuning by Standing on the Shoulders of Decomposition
by: Si, Chongjie, et al.
Published: (2024)
by: Si, Chongjie, et al.
Published: (2024)
Similar Items
-
MedSeg-R: Reasoning Segmentation in Medical Images with Multimodal Large Language Models
by: Huang, Yu, et al.
Published: (2025) -
HyperET: Efficient Training in Hyperbolic Space for Multi-modal Large Language Models
by: Peng, Zelin, et al.
Published: (2025) -
MMRad-22K: A Structured Multimodal Evidence Dataset for Chest X-ray Report Generation
by: Zhao, Yichen, et al.
Published: (2026) -
Parameter-efficient Fine-tuning in Hyperspherical Space for Open-vocabulary Semantic Segmentation
by: Peng, Zelin, et al.
Published: (2024) -
Vision-Informed Flow Image Super-Resolution with Quaternion Spatial Modeling and Dynamic Flow Convolution
by: Cao, Qinglong, et al.
Published: (2024)