RankCLIP: Ranking-Consistent Language-Image Pretraining
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Yiming, Zhao, Zhuokai, Chen, Zhaorun, Feng, Zhili, Ding, Zenghui, Sun, Yining |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DINORANKCLIP: DINOv3 Distillation and Injection for Vision-Language Pretraining with High-Order Ranking Consistency
di: Jiang, Shuyang, et al.
Pubblicazione: (2026)
di: Jiang, Shuyang, et al.
Pubblicazione: (2026)
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding
di: Zhang, Yiming, et al.
Pubblicazione: (2024)
di: Zhang, Yiming, et al.
Pubblicazione: (2024)
Enhancing Vision-Language Model Reliability with Uncertainty-Guided Dropout Decoding
di: Fang, Yixiong, et al.
Pubblicazione: (2024)
di: Fang, Yixiong, et al.
Pubblicazione: (2024)
HALC: Object Hallucination Reduction via Adaptive Focal-Contrast Decoding
di: Chen, Zhaorun, et al.
Pubblicazione: (2024)
di: Chen, Zhaorun, et al.
Pubblicazione: (2024)
GoldiCLIP: The Goldilocks Approach for Balancing Explicit Supervision for Language-Image Pretraining
di: Mohan, Deen Dayal, et al.
Pubblicazione: (2026)
di: Mohan, Deen Dayal, et al.
Pubblicazione: (2026)
CircuitProbe: Tracing Visual Temporal Evidence Flow in Video Language Models
di: Zhang, Yiming, et al.
Pubblicazione: (2025)
di: Zhang, Yiming, et al.
Pubblicazione: (2025)
Investigating and Mitigating Object Hallucinations in Pretrained Vision-Language (CLIP) Models
di: Liu, Yufang, et al.
Pubblicazione: (2024)
di: Liu, Yufang, et al.
Pubblicazione: (2024)
AdaRank: Adaptive Rank Pruning for Enhanced Model Merging
di: Lee, Chanhyuk, et al.
Pubblicazione: (2025)
di: Lee, Chanhyuk, et al.
Pubblicazione: (2025)
IDEA: Image Description Enhanced CLIP-Adapter
di: Ye, Zhipeng, et al.
Pubblicazione: (2025)
di: Ye, Zhipeng, et al.
Pubblicazione: (2025)
RAR: Retrieving And Ranking Augmented MLLMs for Visual Recognition
di: Liu, Ziyu, et al.
Pubblicazione: (2024)
di: Liu, Ziyu, et al.
Pubblicazione: (2024)
Not All Attention Heads Are What You Need: Refining CLIP's Image Representation with Attention Ablation
di: Lin, Feng, et al.
Pubblicazione: (2025)
di: Lin, Feng, et al.
Pubblicazione: (2025)
Attention Down-Sampling Transformer, Relative Ranking and Self-Consistency for Blind Image Quality Assessment
di: Alsaafin, Mohammed, et al.
Pubblicazione: (2024)
di: Alsaafin, Mohammed, et al.
Pubblicazione: (2024)
Text-to-Image GAN with Pretrained Representations
di: You, Xiaozhou, et al.
Pubblicazione: (2024)
di: You, Xiaozhou, et al.
Pubblicazione: (2024)
HyDRA: Hierarchical and Dynamic Rank Adaptation for Mobile Vision Language Model
di: Xi, Yuanhao, et al.
Pubblicazione: (2025)
di: Xi, Yuanhao, et al.
Pubblicazione: (2025)
Not All Layers Are Created Equal: Adaptive LoRA Ranks for Personalized Image Generation
di: Shenaj, Donald, et al.
Pubblicazione: (2026)
di: Shenaj, Donald, et al.
Pubblicazione: (2026)
Region-centric Image-Language Pretraining for Open-Vocabulary Detection
di: Kim, Dahun, et al.
Pubblicazione: (2023)
di: Kim, Dahun, et al.
Pubblicazione: (2023)
Enhancing Adversarial Robustness of Vision-Language Models through Low-Rank Adaptation
di: Ji, Yuheng, et al.
Pubblicazione: (2024)
di: Ji, Yuheng, et al.
Pubblicazione: (2024)
DiffLoRA: Generating Personalized Low-Rank Adaptation Weights with Diffusion
di: Wu, Yujia, et al.
Pubblicazione: (2024)
di: Wu, Yujia, et al.
Pubblicazione: (2024)
DL-QAT: Weight-Decomposed Low-Rank Quantization-Aware Training for Large Language Models
di: Ke, Wenjin, et al.
Pubblicazione: (2025)
di: Ke, Wenjin, et al.
Pubblicazione: (2025)
VeCLIP: Improving CLIP Training via Visual-enriched Captions
di: Lai, Zhengfeng, et al.
Pubblicazione: (2023)
di: Lai, Zhengfeng, et al.
Pubblicazione: (2023)
Data or Language Supervision: What Makes CLIP Better than DINO?
di: Liu, Yiming, et al.
Pubblicazione: (2025)
di: Liu, Yiming, et al.
Pubblicazione: (2025)
A Robust Incomplete Multimodal Low-Rank Adaptation Approach for Emotion Recognition
di: Zhao, Xinkui, et al.
Pubblicazione: (2025)
di: Zhao, Xinkui, et al.
Pubblicazione: (2025)
MoPE-CLIP: Structured Pruning for Efficient Vision-Language Models with Module-wise Pruning Error Metric
di: Lin, Haokun, et al.
Pubblicazione: (2024)
di: Lin, Haokun, et al.
Pubblicazione: (2024)
CLIP-RL: Surgical Scene Segmentation Using Contrastive Language-Vision Pretraining & Reinforcement Learning
di: Ahmed, Fatmaelzahraa Ali, et al.
Pubblicazione: (2025)
di: Ahmed, Fatmaelzahraa Ali, et al.
Pubblicazione: (2025)
HALOC: Hardware-Aware Automatic Low-Rank Compression for Compact Neural Networks
di: Xiao, Jinqi, et al.
Pubblicazione: (2023)
di: Xiao, Jinqi, et al.
Pubblicazione: (2023)
Structured Unrestricted-Rank Matrices for Parameter Efficient Fine-tuning
di: Sehanobish, Arijit, et al.
Pubblicazione: (2024)
di: Sehanobish, Arijit, et al.
Pubblicazione: (2024)
Integrating Saliency Ranking and Reinforcement Learning for Enhanced Object Detection
di: Bartolo, Matthias, et al.
Pubblicazione: (2024)
di: Bartolo, Matthias, et al.
Pubblicazione: (2024)
TULIP: Towards Unified Language-Image Pretraining
di: Tang, Zineng, et al.
Pubblicazione: (2025)
di: Tang, Zineng, et al.
Pubblicazione: (2025)
RL for Consistency Models: Faster Reward Guided Text-to-Image Generation
di: Oertell, Owen, et al.
Pubblicazione: (2024)
di: Oertell, Owen, et al.
Pubblicazione: (2024)
DiffCLIP: Differential Attention Meets CLIP
di: Hammoud, Hasan Abed Al Kader, et al.
Pubblicazione: (2025)
di: Hammoud, Hasan Abed Al Kader, et al.
Pubblicazione: (2025)
Interpreting CLIP: Insights on the Robustness to ImageNet Distribution Shifts
di: Crabbé, Jonathan, et al.
Pubblicazione: (2023)
di: Crabbé, Jonathan, et al.
Pubblicazione: (2023)
Patch Ranking: Efficient CLIP by Learning to Rank Local Patches
di: Wu, Cheng-En, et al.
Pubblicazione: (2024)
di: Wu, Cheng-En, et al.
Pubblicazione: (2024)
One Head Eight Arms: Block Matrix based Low Rank Adaptation for CLIP-based Few-Shot Learning
di: Zhou, Chunpeng, et al.
Pubblicazione: (2025)
di: Zhou, Chunpeng, et al.
Pubblicazione: (2025)
Ranking vs. Assignment: The Metric Mismatch in Multi-View Object Association
di: Shelukhan, Matvei, et al.
Pubblicazione: (2026)
di: Shelukhan, Matvei, et al.
Pubblicazione: (2026)
Training Neural Networks from Scratch with Parallel Low-Rank Adapters
di: Huh, Minyoung, et al.
Pubblicazione: (2024)
di: Huh, Minyoung, et al.
Pubblicazione: (2024)
CountCLIP -- [Re] Teaching CLIP to Count to Ten
di: Mestha, Harshvardhan, et al.
Pubblicazione: (2024)
di: Mestha, Harshvardhan, et al.
Pubblicazione: (2024)
On the Reproducibility of "FairCLIP: Harnessing Fairness in Vision-Language Learning''
di: Bakker, Hua Chang, et al.
Pubblicazione: (2025)
di: Bakker, Hua Chang, et al.
Pubblicazione: (2025)
MKOR: Momentum-Enabled Kronecker-Factor-Based Optimizer Using Rank-1 Updates
di: Mozaffari, Mohammad, et al.
Pubblicazione: (2023)
di: Mozaffari, Mohammad, et al.
Pubblicazione: (2023)
Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
di: Li, Qixiu, et al.
Pubblicazione: (2025)
di: Li, Qixiu, et al.
Pubblicazione: (2025)
Improving Image Captioning Descriptiveness by Ranking and LLM-based Fusion
di: Celona, Luigi, et al.
Pubblicazione: (2023)
di: Celona, Luigi, et al.
Pubblicazione: (2023)
Documenti analoghi
-
DINORANKCLIP: DINOv3 Distillation and Injection for Vision-Language Pretraining with High-Order Ranking Consistency
di: Jiang, Shuyang, et al.
Pubblicazione: (2026) -
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding
di: Zhang, Yiming, et al.
Pubblicazione: (2024) -
Enhancing Vision-Language Model Reliability with Uncertainty-Guided Dropout Decoding
di: Fang, Yixiong, et al.
Pubblicazione: (2024) -
HALC: Object Hallucination Reduction via Adaptive Focal-Contrast Decoding
di: Chen, Zhaorun, et al.
Pubblicazione: (2024) -
GoldiCLIP: The Goldilocks Approach for Balancing Explicit Supervision for Language-Image Pretraining
di: Mohan, Deen Dayal, et al.
Pubblicazione: (2026)