Geometry-Aware CLIP Retrieval via Local Cross-Modal Alignment and Steering
Fuente:
arXiv
Saved in:
| Main Authors: | Prakash, Nirmalendu, Oozeer, Narmeen Fatimah, Su, Xin, Howard, Phillip, Shah, Shaan, He, Zoe Wanying, Wu, Shuang, Raval, Shivam, Lee, Roy Ka-Wei, Khosla, Meenakshi, Abdullah, Amir |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DreamReader: An Interpretability Toolkit for Text-to-Image Models
by: Prakash, Nirmalendu, et al.
Published: (2026)
by: Prakash, Nirmalendu, et al.
Published: (2026)
Spectral Superposition: A Theory of Feature Geometry
by: Ivanov, Georgi, et al.
Published: (2026)
by: Ivanov, Georgi, et al.
Published: (2026)
Representational Alignment Across Model Layers and Brain Regions with Multi-Level Optimal Transport
by: Shah, Shaan, et al.
Published: (2025)
by: Shah, Shaan, et al.
Published: (2025)
Distribution-Aware Feature Selection for SAEs
by: Oozeer, Narmeen, et al.
Published: (2025)
by: Oozeer, Narmeen, et al.
Published: (2025)
Activation Space Interventions Can Be Transferred Between Large Language Models
by: Oozeer, Narmeen, et al.
Published: (2025)
by: Oozeer, Narmeen, et al.
Published: (2025)
Seeing Through Words, Speaking Through Pixels: Deep Representational Alignment Between Vision and Language Models
by: He, Zoe Wanying, et al.
Published: (2025)
by: He, Zoe Wanying, et al.
Published: (2025)
Barycentric alignment for instance-level comparison of neural representations
by: Saha, Shreya, et al.
Published: (2026)
by: Saha, Shreya, et al.
Published: (2026)
Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
by: Oozeer, Narmeen, et al.
Published: (2025)
by: Oozeer, Narmeen, et al.
Published: (2025)
Understanding and Mitigating Dataset Corruption in LLM Steering
by: Anderson, Cullen, et al.
Published: (2026)
by: Anderson, Cullen, et al.
Published: (2026)
Bilinear Convolution Decomposition for Causal RL Interpretability
by: Oozeer, Narmeen, et al.
Published: (2024)
by: Oozeer, Narmeen, et al.
Published: (2024)
Interpreting Bias in Large Language Models: A Feature-Based Approach
by: Prakash, Nirmalendu, et al.
Published: (2024)
by: Prakash, Nirmalendu, et al.
Published: (2024)
Narrow Fine-Tuning Erodes Safety Alignment in Vision-Language Agents
by: Gulati, Idhant, et al.
Published: (2026)
by: Gulati, Idhant, et al.
Published: (2026)
Position: Require Frontier AI Labs To Release Small "Analog" Models
by: Upadhyay, Shriyash, et al.
Published: (2025)
by: Upadhyay, Shriyash, et al.
Published: (2025)
Bridging Critical Gaps in Convergent Learning: How Representational Alignment Evolves Across Layers, Training, and Distribution Shifts
by: Kapoor, Chaitanya, et al.
Published: (2025)
by: Kapoor, Chaitanya, et al.
Published: (2025)
Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal
by: Prakash, Nirmalendu, et al.
Published: (2025)
by: Prakash, Nirmalendu, et al.
Published: (2025)
Curveball Steering: The Right Direction To Steer Isn't Always Linear
by: Raval, Shivam, et al.
Published: (2026)
by: Raval, Shivam, et al.
Published: (2026)
Approximating Human Preferences Using a Multi-Judge Learned System
by: Sprejer, Eitán, et al.
Published: (2025)
by: Sprejer, Eitán, et al.
Published: (2025)
Intersectional Sycophancy: How Perceived User Demographics Shape False Validation in Large Language Models
by: Maltbie, Benjamin, et al.
Published: (2026)
by: Maltbie, Benjamin, et al.
Published: (2026)
Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators
by: Roytburg, Dani, et al.
Published: (2025)
by: Roytburg, Dani, et al.
Published: (2025)
Modeling the Human Visual System: Comparative Insights from Response-Optimized and Task-Optimized Vision Models, Language Models, and different Readout Mechanisms
by: Saha, Shreya, et al.
Published: (2024)
by: Saha, Shreya, et al.
Published: (2024)
Superposition disentanglement of neural representations reveals hidden alignment
by: Longon, André, et al.
Published: (2025)
by: Longon, André, et al.
Published: (2025)
Measuring the Representational Alignment of Neural Systems in Superposition
by: Liu, Sunny, et al.
Published: (2026)
by: Liu, Sunny, et al.
Published: (2026)
Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations
by: Roytburg, Dani, et al.
Published: (2026)
by: Roytburg, Dani, et al.
Published: (2026)
Partial Soft-Matching Distance for Neural Representational Comparison with Partial Unit Correspondence
by: Kapoor, Chaitanya, et al.
Published: (2026)
by: Kapoor, Chaitanya, et al.
Published: (2026)
$β$-CLIP: Text-Conditioned Contrastive Learning for Multi-Granular Vision-Language Alignment
by: Zohra, Fatimah, et al.
Published: (2025)
by: Zohra, Fatimah, et al.
Published: (2025)
DialCLIP: Empowering CLIP as Multi-Modal Dialog Retriever
by: Yin, Zhichao, et al.
Published: (2024)
by: Yin, Zhichao, et al.
Published: (2024)
Quantifying and Enabling the Interpretability of CLIP-like Models
by: Madasu, Avinash, et al.
Published: (2024)
by: Madasu, Avinash, et al.
Published: (2024)
Linear probes rely on textual evidence: Results from leakage mitigation studies in language models
by: Boxo, Gerard, et al.
Published: (2025)
by: Boxo, Gerard, et al.
Published: (2025)
Dissociating Decodability and Causal Use in Bracket-Sequence Transformers
by: Sharma, Aryan, et al.
Published: (2026)
by: Sharma, Aryan, et al.
Published: (2026)
Beyond Monoliths: Expert Orchestration for More Capable, Democratic, and Safe Language Models
by: Quirke, Philip, et al.
Published: (2025)
by: Quirke, Philip, et al.
Published: (2025)
Comparing and Integrating Different Notions of Representational Correspondence in Neural Systems
by: Wu, Jialin, et al.
Published: (2025)
by: Wu, Jialin, et al.
Published: (2025)
Integrated representational signatures strengthen specificity in brains and models
by: Wu, Jialin, et al.
Published: (2025)
by: Wu, Jialin, et al.
Published: (2025)
Measuring the Measures: Discriminative Capacity of Representational Similarity Metrics Across Model Families
by: Wu, Jialin, et al.
Published: (2025)
by: Wu, Jialin, et al.
Published: (2025)
Evaluating Representational Similarity Measures from the Lens of Functional Correspondence
by: Bo, Yiqing, et al.
Published: (2024)
by: Bo, Yiqing, et al.
Published: (2024)
Sparse components distinguish visual pathways & their alignment to neural networks
by: Marvi, Ammar I, et al.
Published: (2025)
by: Marvi, Ammar I, et al.
Published: (2025)
Pruning the Paradox: How CLIP's Most Informative Heads Enhance Performance While Amplifying Bias
by: Madasu, Avinash, et al.
Published: (2025)
by: Madasu, Avinash, et al.
Published: (2025)
Enhancing CLIP Robustness via Cross-Modality Alignment
by: Zhu, Xingyu, et al.
Published: (2025)
by: Zhu, Xingyu, et al.
Published: (2025)
Understanding Refusal in Language Models with Sparse Autoencoders
by: Yeo, Wei Jie, et al.
Published: (2025)
by: Yeo, Wei Jie, et al.
Published: (2025)
Activation Steering for Bias Mitigation: An Interpretable Approach to Safer LLMs
by: Dubey, Shivam
Published: (2025)
by: Dubey, Shivam
Published: (2025)
SGHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Singapore
by: Ng, Ri Chi, et al.
Published: (2024)
by: Ng, Ri Chi, et al.
Published: (2024)
Similar Items
-
DreamReader: An Interpretability Toolkit for Text-to-Image Models
by: Prakash, Nirmalendu, et al.
Published: (2026) -
Spectral Superposition: A Theory of Feature Geometry
by: Ivanov, Georgi, et al.
Published: (2026) -
Representational Alignment Across Model Layers and Brain Regions with Multi-Level Optimal Transport
by: Shah, Shaan, et al.
Published: (2025) -
Distribution-Aware Feature Selection for SAEs
by: Oozeer, Narmeen, et al.
Published: (2025) -
Activation Space Interventions Can Be Transferred Between Large Language Models
by: Oozeer, Narmeen, et al.
Published: (2025)