Centroid-centered Modeling for Efficient Vision Transformer Pre-training
Fuente:
arXiv
Saved in:
| Main Authors: | Yan, Xin, Li, Zuchao, Zhang, Lefei |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MGIMM: Multi-Granularity Instruction Multimodal Model for Attribute-Guided Remote Sensing Image Detailed Description
by: Yang, Cong, et al.
Published: (2024)
by: Yang, Cong, et al.
Published: (2024)
Enhancing Perception of Key Changes in Remote Sensing Image Change Captioning
by: Yang, Cong, et al.
Published: (2024)
by: Yang, Cong, et al.
Published: (2024)
Efficient Adaptation of Pre-trained Vision Transformer via Householder Transformation
by: Dong, Wei, et al.
Published: (2024)
by: Dong, Wei, et al.
Published: (2024)
Efficient and Long-Tailed Generalization for Pre-trained Vision-Language Model
by: Shi, Jiang-Xin, et al.
Published: (2024)
by: Shi, Jiang-Xin, et al.
Published: (2024)
VisionTS++: Cross-Modal Time Series Foundation Model with Continual Pre-trained Vision Backbones
by: Shen, Lefei, et al.
Published: (2025)
by: Shen, Lefei, et al.
Published: (2025)
Multi-modal Auto-regressive Modeling via Visual Words
by: Peng, Tianshuo, et al.
Published: (2024)
by: Peng, Tianshuo, et al.
Published: (2024)
CoViPAL: Layer-wise Contextualized Visual Token Pruning for Large Vision-Language Models
by: Tang, Zicong, et al.
Published: (2025)
by: Tang, Zicong, et al.
Published: (2025)
ECAMP: Entity-centered Context-aware Medical Vision Language Pre-training
by: Wang, Rongsheng, et al.
Published: (2023)
by: Wang, Rongsheng, et al.
Published: (2023)
BUS:Efficient and Effective Vision-language Pre-training with Bottom-Up Patch Summarization
by: Jiang, Chaoya, et al.
Published: (2023)
by: Jiang, Chaoya, et al.
Published: (2023)
Split Adaptation for Pre-trained Vision Transformers
by: Wang, Lixu, et al.
Published: (2025)
by: Wang, Lixu, et al.
Published: (2025)
Continual Forgetting for Pre-trained Vision Models
by: Zhao, Hongbo, et al.
Published: (2024)
by: Zhao, Hongbo, et al.
Published: (2024)
Efficient Vision-Language Pre-training by Cluster Masking
by: Wei, Zihao, et al.
Published: (2024)
by: Wei, Zihao, et al.
Published: (2024)
Efficient Adaptation of Pre-trained Vision Transformer underpinned by Approximately Orthogonal Fine-Tuning Strategy
by: Yang, Yiting, et al.
Published: (2025)
by: Yang, Yiting, et al.
Published: (2025)
Skip Tuning: Pre-trained Vision-Language Models are Effective and Efficient Adapters Themselves
by: Wu, Shihan, et al.
Published: (2024)
by: Wu, Shihan, et al.
Published: (2024)
An Empirical Study of Parameter Efficient Fine-tuning on Vision-Language Pre-train Model
by: Tian, Yuxin, et al.
Published: (2024)
by: Tian, Yuxin, et al.
Published: (2024)
NEVLP: Noise-Robust Framework for Efficient Vision-Language Pre-training
by: Tao, Yiyi, et al.
Published: (2024)
by: Tao, Yiyi, et al.
Published: (2024)
Revisiting Continual Semantic Segmentation with Pre-trained Vision Models
by: Zhang, Duzhen, et al.
Published: (2025)
by: Zhang, Duzhen, et al.
Published: (2025)
VITAL: Vision-Encoder-centered Pre-training for LMMs in Visual Quality Assessment
by: Jia, Ziheng, et al.
Published: (2025)
by: Jia, Ziheng, et al.
Published: (2025)
Semantics-enhanced Cross-modal Masked Image Modeling for Vision-Language Pre-training
by: Liu, Haowei, et al.
Published: (2024)
by: Liu, Haowei, et al.
Published: (2024)
Pre-trained Vision and Language Transformers Are Few-Shot Incremental Learners
by: Park, Keon-Hee, et al.
Published: (2024)
by: Park, Keon-Hee, et al.
Published: (2024)
OFL-SAM2: Prompt SAM2 with Online Few-shot Learner for Efficient Medical Image Segmentation
by: Lan, Meng, et al.
Published: (2025)
by: Lan, Meng, et al.
Published: (2025)
ELIP: Efficient Discriminative Language-Image Pre-training with Fewer Vision Tokens
by: Guo, Yangyang, et al.
Published: (2023)
by: Guo, Yangyang, et al.
Published: (2023)
Mind the Interference: Retaining Pre-trained Knowledge in Parameter Efficient Continual Learning of Vision-Language Models
by: Tang, Longxiang, et al.
Published: (2024)
by: Tang, Longxiang, et al.
Published: (2024)
Efficient and Effective Universal Adversarial Attack against Vision-Language Pre-training Models
by: Yang, Fan, et al.
Published: (2024)
by: Yang, Fan, et al.
Published: (2024)
Efficient Vision-and-Language Pre-training with Text-Relevant Image Patch Selection
by: Ye, Wei, et al.
Published: (2024)
by: Ye, Wei, et al.
Published: (2024)
FedEU: Evidential Uncertainty-Driven Federated Fine-Tuning of Vision Foundation Models for Remote Sensing Image Segmentation
by: Zhang, Xiaokang, et al.
Published: (2026)
by: Zhang, Xiaokang, et al.
Published: (2026)
Anatomical Structure-Guided Medical Vision-Language Pre-training
by: Li, Qingqiu, et al.
Published: (2024)
by: Li, Qingqiu, et al.
Published: (2024)
TRIPS: Efficient Vision-and-Language Pre-training with Text-Relevant Image Patch Selection
by: Jiang, Chaoya, et al.
Published: (2023)
by: Jiang, Chaoya, et al.
Published: (2023)
VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine
by: Zhang, Ziyang, et al.
Published: (2025)
by: Zhang, Ziyang, et al.
Published: (2025)
Sample-agnostic Adversarial Perturbation for Vision-Language Pre-training Models
by: Zheng, Haonan, et al.
Published: (2024)
by: Zheng, Haonan, et al.
Published: (2024)
GLID: Pre-training a Generalist Encoder-Decoder Vision Model
by: Liu, Jihao, et al.
Published: (2024)
by: Liu, Jihao, et al.
Published: (2024)
Cross-Modal Adapter: Parameter-Efficient Transfer Learning Approach for Vision-Language Models
by: Yang, Juncheng, et al.
Published: (2024)
by: Yang, Juncheng, et al.
Published: (2024)
OpenPath: Open-Set Active Learning for Pathology Image Classification via Pre-trained Vision-Language Models
by: Zhong, Lanfeng, et al.
Published: (2025)
by: Zhong, Lanfeng, et al.
Published: (2025)
Steering Vision-Language Pre-trained Models for Incremental Face Presentation Attack Detection
by: Li, Haoze, et al.
Published: (2025)
by: Li, Haoze, et al.
Published: (2025)
UniQA: Unified Vision-Language Pre-training for Image Quality and Aesthetic Assessment
by: Zhou, Hantao, et al.
Published: (2024)
by: Zhou, Hantao, et al.
Published: (2024)
Dual Selective Fusion Transformer Network for Hyperspectral Image Classification
by: Xu, Yichu, et al.
Published: (2024)
by: Xu, Yichu, et al.
Published: (2024)
MAA: Meticulous Adversarial Attack against Vision-Language Pre-trained Models
by: Zhang, Peng-Fei, et al.
Published: (2025)
by: Zhang, Peng-Fei, et al.
Published: (2025)
HyperVision: A Channel-Adaptive Ground-Based Hyperspectral Vision Pre-trained Backbone
by: Fu, Guanyiman, et al.
Published: (2026)
by: Fu, Guanyiman, et al.
Published: (2026)
MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models
by: Hua, Hang, et al.
Published: (2024)
by: Hua, Hang, et al.
Published: (2024)
FashionFAE: Fine-grained Attributes Enhanced Fashion Vision-Language Pre-training
by: Huang, Jiale, et al.
Published: (2024)
by: Huang, Jiale, et al.
Published: (2024)
Similar Items
-
MGIMM: Multi-Granularity Instruction Multimodal Model for Attribute-Guided Remote Sensing Image Detailed Description
by: Yang, Cong, et al.
Published: (2024) -
Enhancing Perception of Key Changes in Remote Sensing Image Change Captioning
by: Yang, Cong, et al.
Published: (2024) -
Efficient Adaptation of Pre-trained Vision Transformer via Householder Transformation
by: Dong, Wei, et al.
Published: (2024) -
Efficient and Long-Tailed Generalization for Pre-trained Vision-Language Model
by: Shi, Jiang-Xin, et al.
Published: (2024) -
VisionTS++: Cross-Modal Time Series Foundation Model with Continual Pre-trained Vision Backbones
by: Shen, Lefei, et al.
Published: (2025)