PERL: Parameter Efficient Reasoning in CLIP Latent Space
Fuente:
arXiv
Saved in:
| Main Authors: | Carnemolla, Simone, Calcagno, Salvatore, Giordano, Daniela, Spampinato, Concetto, Pennisi, Matteo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SeeingSounds: Learning Audio-to-Visual Alignment via Text
by: Carnemolla, Simone, et al.
Published: (2025)
by: Carnemolla, Simone, et al.
Published: (2025)
DEXTER: Diffusion-Guided EXplanations with TExtual Reasoning for Vision Models
by: Carnemolla, Simone, et al.
Published: (2025)
by: Carnemolla, Simone, et al.
Published: (2025)
UNBOX: Unveiling Black-box visual models with Natural-language
by: Carnemolla, Simone, et al.
Published: (2026)
by: Carnemolla, Simone, et al.
Published: (2026)
Diffexplainer: Towards Cross-modal Global Explanations with Diffusion Models
by: Pennisi, Matteo, et al.
Published: (2024)
by: Pennisi, Matteo, et al.
Published: (2024)
Selective Attention-based Modulation for Continual Learning
by: Bellitto, Giovanni, et al.
Published: (2024)
by: Bellitto, Giovanni, et al.
Published: (2024)
Zero-Shot Decentralized Federated Learning
by: Masano, Alessio, et al.
Published: (2025)
by: Masano, Alessio, et al.
Published: (2025)
Transformer-based Video Saliency Prediction with High Temporal Dimension Decoding
by: Moradi, Morteza, et al.
Published: (2024)
by: Moradi, Morteza, et al.
Published: (2024)
Wake-Sleep Consolidated Learning
by: Sorrenti, Amelia, et al.
Published: (2023)
by: Sorrenti, Amelia, et al.
Published: (2023)
MatFuse: Controllable Material Generation with Diffusion Models
by: Vecchio, Giuseppe, et al.
Published: (2023)
by: Vecchio, Giuseppe, et al.
Published: (2023)
Global-Local Feature Decoding with Adapter-Guided SAMv2 for Salient Object Detection
by: Moradi, Morteza, et al.
Published: (2026)
by: Moradi, Morteza, et al.
Published: (2026)
OCCAM: Open-set Causal Concept explAnation and Ontology induction for black-box vision Models
by: Russo, Chiara Maria, et al.
Published: (2026)
by: Russo, Chiara Maria, et al.
Published: (2026)
QuantFormer: Learning to Quantize for Neural Activity Forecasting in Mouse Visual Cortex
by: Calcagno, Salvatore, et al.
Published: (2024)
by: Calcagno, Salvatore, et al.
Published: (2024)
SalFoM: Dynamic Saliency Prediction with Video Foundation Models
by: Moradi, Morteza, et al.
Published: (2024)
by: Moradi, Morteza, et al.
Published: (2024)
Evidential Federated Learning for Skin Lesion Image Classification
by: Hendrix, Rutger, et al.
Published: (2024)
by: Hendrix, Rutger, et al.
Published: (2024)
CLIP-Map: Structured Matrix Mapping for Parameter-Efficient CLIP Compression
by: Zhang, Kangjie, et al.
Published: (2026)
by: Zhang, Kangjie, et al.
Published: (2026)
Self-supervised learning for radio-astronomy source classification: a benchmark
by: Cecconello, Thomas, et al.
Published: (2024)
by: Cecconello, Thomas, et al.
Published: (2024)
CLIP-SLA: Parameter-Efficient CLIP Adaptation for Continuous Sign Language Recognition
by: Alyami, Sarah, et al.
Published: (2025)
by: Alyami, Sarah, et al.
Published: (2025)
Back to Supervision: Boosting Word Boundary Detection through Frame Classification
by: Carnemolla, Simone, et al.
Published: (2024)
by: Carnemolla, Simone, et al.
Published: (2024)
HAC: Parameter-Efficient Hyperbolic Adaptation of CLIP for Zero-Shot VQA
by: Dibitonto, Francesco, et al.
Published: (2026)
by: Dibitonto, Francesco, et al.
Published: (2026)
CLIP-Guided SAM: Parameter-Efficient Semantic Conditioning for Promptable Segmentation
by: Jalilian, Shayan, et al.
Published: (2026)
by: Jalilian, Shayan, et al.
Published: (2026)
EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters
by: Sun, Quan, et al.
Published: (2024)
by: Sun, Quan, et al.
Published: (2024)
AD-CLIP: Adapting Domains in Prompt Space Using CLIP
by: Singha, Mainak, et al.
Published: (2023)
by: Singha, Mainak, et al.
Published: (2023)
DetailCLIP: Injecting Image Details into CLIP's Feature Space
by: Zhang, Zilun, et al.
Published: (2022)
by: Zhang, Zilun, et al.
Published: (2022)
LatentGeo: Learnable Auxiliary Constructions in Latent Space for Multimodal Geometric Reasoning
by: Xu, Haiying, et al.
Published: (2026)
by: Xu, Haiying, et al.
Published: (2026)
uCLIP: Parameter-Efficient Multilingual Extension of Vision-Language Models with Unpaired Data
by: Chung, Dahyun, et al.
Published: (2025)
by: Chung, Dahyun, et al.
Published: (2025)
IsoCLIP: Decomposing CLIP Projectors for Efficient Intra-modal Alignment
by: Magistri, Simone, et al.
Published: (2026)
by: Magistri, Simone, et al.
Published: (2026)
ULF-Synth: Physics-Guided Ultra-Low-Field MRI Enhancement for Pediatric Neuroimaging
by: Musah, Toufiq, et al.
Published: (2026)
by: Musah, Toufiq, et al.
Published: (2026)
CLIP-RD: Relative Distillation for Efficient CLIP Knowledge Distillation
by: Chung, Jeannie, et al.
Published: (2026)
by: Chung, Jeannie, et al.
Published: (2026)
$λ$-ECLIPSE: Multi-Concept Personalized Text-to-Image Diffusion Models by Leveraging CLIP Latent Space
by: Patel, Maitreya, et al.
Published: (2024)
by: Patel, Maitreya, et al.
Published: (2024)
Maintaining Structural Integrity in Parameter Spaces for Parameter Efficient Fine-tuning
by: Si, Chongjie, et al.
Published: (2024)
by: Si, Chongjie, et al.
Published: (2024)
CLIPure: Purification in Latent Space via CLIP for Adversarially Robust Zero-Shot Classification
by: Zhang, Mingkun, et al.
Published: (2025)
by: Zhang, Mingkun, et al.
Published: (2025)
Dream2Learn: Structured Generative Dreaming for Continual Learning
by: Calcagno, Salvatore, et al.
Published: (2026)
by: Calcagno, Salvatore, et al.
Published: (2026)
Distilling Knowledge from Text-to-Image Generative Models Improves Visio-Linguistic Reasoning in CLIP
by: Basu, Samyadeep, et al.
Published: (2023)
by: Basu, Samyadeep, et al.
Published: (2023)
CLIP-CID: Efficient CLIP Distillation via Cluster-Instance Discrimination
by: Yang, Kaicheng, et al.
Published: (2024)
by: Yang, Kaicheng, et al.
Published: (2024)
MARBLE: Material Recomposition and Blending in CLIP-Space
by: Cheng, Ta-Ying, et al.
Published: (2025)
by: Cheng, Ta-Ying, et al.
Published: (2025)
Adaptive Weighted Parameter Fusion with CLIP for Class-Incremental Learning
by: Guo, Juncen, et al.
Published: (2025)
by: Guo, Juncen, et al.
Published: (2025)
PE-CLIP: A Parameter-Efficient Fine-Tuning of Vision Language Models for Dynamic Facial Expression Recognition
by: Saadi, Ibtissam, et al.
Published: (2025)
by: Saadi, Ibtissam, et al.
Published: (2025)
MoCLIP-Lite: Efficient Video Recognition by Fusing CLIP with Motion Vectors
by: Huang, Binhua, et al.
Published: (2025)
by: Huang, Binhua, et al.
Published: (2025)
Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space
by: Liu, Chengzhi, et al.
Published: (2025)
by: Liu, Chengzhi, et al.
Published: (2025)
MaskedCLIP: Bridging the Masked and CLIP Space for Semi-Supervised Medical Vision-Language Pre-training
by: Zhu, Lei, et al.
Published: (2025)
by: Zhu, Lei, et al.
Published: (2025)
Similar Items
-
SeeingSounds: Learning Audio-to-Visual Alignment via Text
by: Carnemolla, Simone, et al.
Published: (2025) -
DEXTER: Diffusion-Guided EXplanations with TExtual Reasoning for Vision Models
by: Carnemolla, Simone, et al.
Published: (2025) -
UNBOX: Unveiling Black-box visual models with Natural-language
by: Carnemolla, Simone, et al.
Published: (2026) -
Diffexplainer: Towards Cross-modal Global Explanations with Diffusion Models
by: Pennisi, Matteo, et al.
Published: (2024) -
Selective Attention-based Modulation for Continual Learning
by: Bellitto, Giovanni, et al.
Published: (2024)