BiCLIP: Domain Canonicalization via Structured Geometric Transformation
Fuente:
arXiv
Saved in:
| Main Authors: | Mantini, Pranav, Shah, Shishir K. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GeoStack: A Framework for Quasi-Abelian Knowledge Composition in VLMs
by: Mantini, Pranav, et al.
Published: (2026)
by: Mantini, Pranav, et al.
Published: (2026)
MoDE: CLIP Data Experts via Clustering
by: Ma, Jiawei, et al.
Published: (2024)
by: Ma, Jiawei, et al.
Published: (2024)
Enhancing CLIP Conceptual Embedding through Knowledge Distillation
by: Kao, Kuei-Chun
Published: (2024)
by: Kao, Kuei-Chun
Published: (2024)
PerceptionCLIP: Visual Classification by Inferring and Conditioning on Contexts
by: An, Bang, et al.
Published: (2023)
by: An, Bang, et al.
Published: (2023)
MobileCLIP2: Improving Multi-Modal Reinforced Training
by: Faghri, Fartash, et al.
Published: (2025)
by: Faghri, Fartash, et al.
Published: (2025)
X-Transfer Attacks: Towards Super Transferable Adversarial Attacks on CLIP
by: Huang, Hanxun, et al.
Published: (2025)
by: Huang, Hanxun, et al.
Published: (2025)
GET: Unlocking the Multi-modal Potential of CLIP for Generalized Category Discovery
by: Wang, Enguang, et al.
Published: (2024)
by: Wang, Enguang, et al.
Published: (2024)
Mitigate the Gap: Investigating Approaches for Improving Cross-Modal Alignment in CLIP
by: Eslami, Sedigheh, et al.
Published: (2024)
by: Eslami, Sedigheh, et al.
Published: (2024)
Seeing is Believing: Mitigating Hallucination in Large Vision-Language Models via CLIP-Guided Decoding
by: Deng, Ailin, et al.
Published: (2024)
by: Deng, Ailin, et al.
Published: (2024)
Sim-CLIP: Unsupervised Siamese Adversarial Fine-Tuning for Robust and Semantically-Rich Vision-Language Models
by: Hossain, Md Zarif, et al.
Published: (2024)
by: Hossain, Md Zarif, et al.
Published: (2024)
Euclid's Gift: Enhancing Spatial Perception and Reasoning in Vision-Language Models via Geometric Surrogate Tasks
by: Lian, Shijie, et al.
Published: (2025)
by: Lian, Shijie, et al.
Published: (2025)
Set-CLIP: Exploring Aligned Semantic From Low-Alignment Multimodal Data Through A Distribution View
by: Song, Zijia, et al.
Published: (2024)
by: Song, Zijia, et al.
Published: (2024)
Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates
by: Ahn, Jaewoo, et al.
Published: (2025)
by: Ahn, Jaewoo, et al.
Published: (2025)
DiagrammerGPT: Generating Open-Domain, Open-Platform Diagrams via LLM Planning
by: Zala, Abhay, et al.
Published: (2023)
by: Zala, Abhay, et al.
Published: (2023)
A General and Efficient Training for Transformer via Token Expansion
by: Huang, Wenxuan, et al.
Published: (2024)
by: Huang, Wenxuan, et al.
Published: (2024)
ET tu, CLIP? Addressing Common Object Errors for Unseen Environments
by: Byun, Ye Won, et al.
Published: (2024)
by: Byun, Ye Won, et al.
Published: (2024)
Data or Language Supervision: What Makes CLIP Better than DINO?
by: Liu, Yiming, et al.
Published: (2025)
by: Liu, Yiming, et al.
Published: (2025)
CatLIP: CLIP-level Visual Recognition Accuracy with 2.7x Faster Pre-training on Web-scale Image-Text Data
by: Mehta, Sachin, et al.
Published: (2024)
by: Mehta, Sachin, et al.
Published: (2024)
Reproducibility Study of CDUL: CLIP-Driven Unsupervised Learning for Multi-Label Image Classification
by: Shah, Manan, et al.
Published: (2024)
by: Shah, Manan, et al.
Published: (2024)
VeCLIP: Improving CLIP Training via Visual-enriched Captions
by: Lai, Zhengfeng, et al.
Published: (2023)
by: Lai, Zhengfeng, et al.
Published: (2023)
Unveiling the Hidden Structure of Self-Attention via Kernel Principal Component Analysis
by: Teo, Rachel S. Y., et al.
Published: (2024)
by: Teo, Rachel S. Y., et al.
Published: (2024)
Transformers without Normalization
by: Zhu, Jiachen, et al.
Published: (2025)
by: Zhu, Jiachen, et al.
Published: (2025)
Robustness in Both Domains: CLIP Needs a Robust Text Encoder
by: Rocamora, Elias Abad, et al.
Published: (2025)
by: Rocamora, Elias Abad, et al.
Published: (2025)
Attention-based Shape and Gait Representations Learning for Video-based Cloth-Changing Person Re-Identification
by: Nguyen, Vuong D., et al.
Published: (2024)
by: Nguyen, Vuong D., et al.
Published: (2024)
Data Quality Aware Approaches for Addressing Model Drift of Semantic Segmentation Models
by: Mirza, Samiha, et al.
Published: (2024)
by: Mirza, Samiha, et al.
Published: (2024)
Stronger Normalization-Free Transformers
by: Chen, Mingzhi, et al.
Published: (2025)
by: Chen, Mingzhi, et al.
Published: (2025)
Bias Is a Subspace, Not a Coordinate: A Geometric Rethinking of Post-hoc Debiasing in Vision-Language Models
by: Zhao, Dachuan, et al.
Published: (2025)
by: Zhao, Dachuan, et al.
Published: (2025)
DoMIX: An Efficient Framework for Exploiting Domain Knowledge in Fine-Tuning
by: Kim, Dohoon, et al.
Published: (2025)
by: Kim, Dohoon, et al.
Published: (2025)
Transformers are Stateless Differentiable Neural Computers
by: Tang, Bo, et al.
Published: (2026)
by: Tang, Bo, et al.
Published: (2026)
Universal Approximation of Visual Autoregressive Transformers
by: Chen, Yifang, et al.
Published: (2025)
by: Chen, Yifang, et al.
Published: (2025)
Energy-Based Transformers are Scalable Learners and Thinkers
by: Gladstone, Alexi, et al.
Published: (2025)
by: Gladstone, Alexi, et al.
Published: (2025)
SeqPE: Transformer with Sequential Position Encoding
by: Li, Huayang, et al.
Published: (2025)
by: Li, Huayang, et al.
Published: (2025)
Impact of Layer Norm on Memorization and Generalization in Transformers
by: Singhal, Rishi, et al.
Published: (2025)
by: Singhal, Rishi, et al.
Published: (2025)
Bi-Level Optimization for Single Domain Generalization
by: Heidari, Marzi, et al.
Published: (2026)
by: Heidari, Marzi, et al.
Published: (2026)
A Primal-Dual Framework for Transformers and Neural Networks
by: Nguyen, Tan M., et al.
Published: (2024)
by: Nguyen, Tan M., et al.
Published: (2024)
Robust Pre-Training of Medical Vision-and-Language Models with Domain-Invariant Multi-Modal Masked Reconstruction
by: Filvantorkaman, Melika, et al.
Published: (2026)
by: Filvantorkaman, Melika, et al.
Published: (2026)
Guiding Vision-Language Model Selection for Visual Question-Answering Across Tasks, Domains, and Knowledge Types
by: Sinha, Neelabh, et al.
Published: (2024)
by: Sinha, Neelabh, et al.
Published: (2024)
DiffCLIP: Differential Attention Meets CLIP
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
Frozen Transformers in Language Models Are Effective Visual Encoder Layers
by: Pang, Ziqi, et al.
Published: (2023)
by: Pang, Ziqi, et al.
Published: (2023)
AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers
by: Achtibat, Reduan, et al.
Published: (2024)
by: Achtibat, Reduan, et al.
Published: (2024)
Similar Items
-
GeoStack: A Framework for Quasi-Abelian Knowledge Composition in VLMs
by: Mantini, Pranav, et al.
Published: (2026) -
MoDE: CLIP Data Experts via Clustering
by: Ma, Jiawei, et al.
Published: (2024) -
Enhancing CLIP Conceptual Embedding through Knowledge Distillation
by: Kao, Kuei-Chun
Published: (2024) -
PerceptionCLIP: Visual Classification by Inferring and Conditioning on Contexts
by: An, Bang, et al.
Published: (2023) -
MobileCLIP2: Improving Multi-Modal Reinforced Training
by: Faghri, Fartash, et al.
Published: (2025)