Refining CLIP's Spatial Awareness: A Visual-Centric Perspective
Fuente:
arXiv
Saved in:
| Main Authors: | Qiu, Congpei, Wu, Yanhao, Ke, Wei, Bai, Xiuxiu, Zhang, Tong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UniRefiner: Teaching Pre-trained ViTs to Self-Dispose Dross via Contrastive Register
by: Qiu, Congpei, et al.
Published: (2026)
by: Qiu, Congpei, et al.
Published: (2026)
Mitigating Object Dependencies: Improving Point Cloud Self-Supervised Learning through Object Exchange
by: Wu, Yanhao, et al.
Published: (2024)
by: Wu, Yanhao, et al.
Published: (2024)
Generating Multimodal Driving Scenes via Next-Scene Prediction
by: Wu, Yanhao, et al.
Published: (2025)
by: Wu, Yanhao, et al.
Published: (2025)
AlignDrive: Aligned Lateral-Longitudinal Planning for End-to-End Autonomous Driving
by: Wu, Yanhao, et al.
Published: (2026)
by: Wu, Yanhao, et al.
Published: (2026)
BadCLIP: Trigger-Aware Prompt Learning for Backdoor Attacks on CLIP
by: Bai, Jiawang, et al.
Published: (2023)
by: Bai, Jiawang, et al.
Published: (2023)
Rethinking Visual Content Refinement in Low-Shot CLIP Adaptation
by: Lu, Jinda, et al.
Published: (2024)
by: Lu, Jinda, et al.
Published: (2024)
CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization
by: Xia, Rui, et al.
Published: (2025)
by: Xia, Rui, et al.
Published: (2025)
CLIP-Guided Adaptable Self-Supervised Learning for Human-Centric Visual Tasks
by: Luo, Mingshuang, et al.
Published: (2026)
by: Luo, Mingshuang, et al.
Published: (2026)
Focus-Scan-Refine: From Human Visual Perception to Efficient Visual Token Pruning
by: Tong, Enwei, et al.
Published: (2026)
by: Tong, Enwei, et al.
Published: (2026)
Positional Prompt Tuning for Efficient 3D Representation Learning
by: Zhang, Shaochen, et al.
Published: (2024)
by: Zhang, Shaochen, et al.
Published: (2024)
Hand-Centric Motion Refinement for 3D Hand-Object Interaction via Hierarchical Spatial-Temporal Modeling
by: Hao, Yuze, et al.
Published: (2024)
by: Hao, Yuze, et al.
Published: (2024)
Spatial Orthogonal Refinement for Robust RGB-Event Visual Object Tracking
by: Huang, Dexing, et al.
Published: (2026)
by: Huang, Dexing, et al.
Published: (2026)
EntityCLIP: Entity-Centric Image-Text Matching via Multimodal Attentive Contrastive Learning
by: Wang, Yaxiong, et al.
Published: (2024)
by: Wang, Yaxiong, et al.
Published: (2024)
Making Better Mistakes in CLIP-Based Zero-Shot Classification with Hierarchy-Aware Language Prompts
by: Liang, Tong, et al.
Published: (2025)
by: Liang, Tong, et al.
Published: (2025)
Redundant Queries in DETR-Based 3D Detection Methods: Unnecessary and Prunable
by: Xu, Lizhen, et al.
Published: (2024)
by: Xu, Lizhen, et al.
Published: (2024)
Spatially-Weighted CLIP for Street-View Geo-localization
by: Han, Ting, et al.
Published: (2026)
by: Han, Ting, et al.
Published: (2026)
WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition
by: Ning, Shan, et al.
Published: (2026)
by: Ning, Shan, et al.
Published: (2026)
ProxyCLIP: Proxy Attention Improves CLIP for Open-Vocabulary Segmentation
by: Lan, Mengcheng, et al.
Published: (2024)
by: Lan, Mengcheng, et al.
Published: (2024)
ClearCLIP: Decomposing CLIP Representations for Dense Vision-Language Inference
by: Lan, Mengcheng, et al.
Published: (2024)
by: Lan, Mengcheng, et al.
Published: (2024)
Enhancing Zero-Shot Anomaly Detection: CLIP-SAM Collaboration with Cascaded Prompts
by: Hou, Yanning, et al.
Published: (2025)
by: Hou, Yanning, et al.
Published: (2025)
VRSO: Visual-Centric Reconstruction for Static Object Annotation
by: Yu, Chenyao, et al.
Published: (2024)
by: Yu, Chenyao, et al.
Published: (2024)
IPAD-CLIP: Teaching CLIP to Detect Image Local Perceptual Artifacts
by: Wang, Juan, et al.
Published: (2026)
by: Wang, Juan, et al.
Published: (2026)
MedP-CLIP: Medical CLIP with Region-Aware Prompt Integration
by: Peng, Jiahui, et al.
Published: (2026)
by: Peng, Jiahui, et al.
Published: (2026)
Visual Self-Refinement for Autoregressive Models
by: Wang, Jiamian, et al.
Published: (2025)
by: Wang, Jiamian, et al.
Published: (2025)
MediCLIP: Adapting CLIP for Few-shot Medical Image Anomaly Detection
by: Zhang, Ximiao, et al.
Published: (2024)
by: Zhang, Ximiao, et al.
Published: (2024)
WP-CLIP: Leveraging CLIP to Predict Wölfflin's Principles in Visual Art
by: Ghildyal, Abhijay, et al.
Published: (2025)
by: Ghildyal, Abhijay, et al.
Published: (2025)
CLIP-VG: Self-paced Curriculum Adapting of CLIP for Visual Grounding
by: Xiao, Linhui, et al.
Published: (2023)
by: Xiao, Linhui, et al.
Published: (2023)
un$^2$CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP
by: Li, Yinqi, et al.
Published: (2025)
by: Li, Yinqi, et al.
Published: (2025)
Appearance-Based Refinement for Object-Centric Motion Segmentation
by: Xie, Junyu, et al.
Published: (2023)
by: Xie, Junyu, et al.
Published: (2023)
VeCLIP: Improving CLIP Training via Visual-enriched Captions
by: Lai, Zhengfeng, et al.
Published: (2023)
by: Lai, Zhengfeng, et al.
Published: (2023)
SPACE-CLIP: Spatial Perception via Adaptive CLIP Embeddings for Monocular Depth Estimation
by: Cho, Taewan, et al.
Published: (2026)
by: Cho, Taewan, et al.
Published: (2026)
ContextRefine-CLIP for EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge 2025
by: He, Jing, et al.
Published: (2025)
by: He, Jing, et al.
Published: (2025)
Dual-Prompt CLIP with Hybrid Visual Encoders for Occluded Person Re-Identification
by: Ji, Zhangjian, et al.
Published: (2026)
by: Ji, Zhangjian, et al.
Published: (2026)
GatedCLIP: Gated Multimodal Fusion for Hateful Memes Detection
by: Guo, Yingying, et al.
Published: (2026)
by: Guo, Yingying, et al.
Published: (2026)
Detecting Deepfakes with Multivariate Soft Blending and CLIP-based Image-Text Alignment
by: Li, Jingwei, et al.
Published: (2026)
by: Li, Jingwei, et al.
Published: (2026)
CLIP-Powered TASS: Target-Aware Single-Stream Network for Audio-Visual Question Answering
by: Jiang, Yuanyuan, et al.
Published: (2024)
by: Jiang, Yuanyuan, et al.
Published: (2024)
Improved Visual-Spatial Reasoning via R1-Zero-Like Training
by: Liao, Zhenyi, et al.
Published: (2025)
by: Liao, Zhenyi, et al.
Published: (2025)
CLIP-UP: CLIP-Based Unanswerable Problem Detection for Visual Question Answering
by: Vardi, Ben, et al.
Published: (2025)
by: Vardi, Ben, et al.
Published: (2025)
CausalSpatial: A Benchmark for Object-Centric Causal Spatial Reasoning
by: Ma, Wenxin, et al.
Published: (2026)
by: Ma, Wenxin, et al.
Published: (2026)
OmniCLIP: Adapting CLIP for Video Recognition with Spatial-Temporal Omni-Scale Feature Learning
by: Liu, Mushui, et al.
Published: (2024)
by: Liu, Mushui, et al.
Published: (2024)
Similar Items
-
UniRefiner: Teaching Pre-trained ViTs to Self-Dispose Dross via Contrastive Register
by: Qiu, Congpei, et al.
Published: (2026) -
Mitigating Object Dependencies: Improving Point Cloud Self-Supervised Learning through Object Exchange
by: Wu, Yanhao, et al.
Published: (2024) -
Generating Multimodal Driving Scenes via Next-Scene Prediction
by: Wu, Yanhao, et al.
Published: (2025) -
AlignDrive: Aligned Lateral-Longitudinal Planning for End-to-End Autonomous Driving
by: Wu, Yanhao, et al.
Published: (2026) -
BadCLIP: Trigger-Aware Prompt Learning for Backdoor Attacks on CLIP
by: Bai, Jiawang, et al.
Published: (2023)