SwordBench: Evaluating Orthogonality of Steering Image Representations
Fuente:
arXiv
Saved in:
| Main Authors: | Zaigrajew, Vladimir, Pludowski, Dawid, Baniecki, Hubert, Biecek, Przemyslaw |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Interpreting CLIP with Hierarchical Sparse Autoencoders
by: Zaigrajew, Vladimir, et al.
Published: (2025)
by: Zaigrajew, Vladimir, et al.
Published: (2025)
Your CLIP has 164 dimensions of noise: Exploring the embeddings covariance eigenspectrum of contrastively pretrained vision-language transformers
by: Grzywaczewski, Jakub, et al.
Published: (2026)
by: Grzywaczewski, Jakub, et al.
Published: (2026)
Adversarial attacks and defenses in explainable artificial intelligence: A survey
by: Baniecki, Hubert, et al.
Published: (2023)
by: Baniecki, Hubert, et al.
Published: (2023)
Birds look like cars: Adversarial analysis of intrinsically interpretable deep learning
by: Baniecki, Hubert, et al.
Published: (2025)
by: Baniecki, Hubert, et al.
Published: (2025)
Red Teaming Models for Hyperspectral Image Analysis Using Explainable AI
by: Zaigrajew, Vladimir, et al.
Published: (2024)
by: Zaigrajew, Vladimir, et al.
Published: (2024)
LINE: LLM-based Iterative Neuron Explanations for Vision Models
by: Zaigrajew, Vladimir, et al.
Published: (2026)
by: Zaigrajew, Vladimir, et al.
Published: (2026)
Explaining Similarity in Vision-Language Encoders with Weighted Banzhaf Interactions
by: Baniecki, Hubert, et al.
Published: (2025)
by: Baniecki, Hubert, et al.
Published: (2025)
Red-Teaming Segment Anything Model
by: Jankowski, Krzysztof, et al.
Published: (2024)
by: Jankowski, Krzysztof, et al.
Published: (2024)
Position: Do Not Explain Vision Models Without Context
by: Tomaszewska, Paulina, et al.
Published: (2024)
by: Tomaszewska, Paulina, et al.
Published: (2024)
Global Counterfactual Directions
by: Sobieski, Bartlomiej, et al.
Published: (2024)
by: Sobieski, Bartlomiej, et al.
Published: (2024)
Local Intrinsic Dimension Unveils Hallucinations in Diffusion Models
by: Sobieski, Bartlomiej, et al.
Published: (2026)
by: Sobieski, Bartlomiej, et al.
Published: (2026)
NormEnsembleXAI: Unveiling the Strengths and Weaknesses of XAI Ensemble Techniques
by: Hryniewska-Guzik, Weronika, et al.
Published: (2024)
by: Hryniewska-Guzik, Weronika, et al.
Published: (2024)
Attributions All the Way Down? The Metagame of Interpretability
by: Baniecki, Hubert, et al.
Published: (2026)
by: Baniecki, Hubert, et al.
Published: (2026)
X-ray transferable polyrepresentation learning
by: Hryniewska-Guzik, Weronika, et al.
Published: (2025)
by: Hryniewska-Guzik, Weronika, et al.
Published: (2025)
Auditing Sybil: Explaining Deep Lung Cancer Risk Prediction Through Generative Interventional Attributions
by: Sobieski, Bartlomiej, et al.
Published: (2026)
by: Sobieski, Bartlomiej, et al.
Published: (2026)
Interpretable machine learning for time-to-event prediction in medicine and healthcare
by: Baniecki, Hubert, et al.
Published: (2023)
by: Baniecki, Hubert, et al.
Published: (2023)
Aggregated Attributions for Explanatory Analysis of 3D Segmentation Models
by: Chrabaszcz, Maciej, et al.
Published: (2024)
by: Chrabaszcz, Maciej, et al.
Published: (2024)
Pre-training with Random Orthogonal Projection Image Modeling
by: Haghighat, Maryam, et al.
Published: (2023)
by: Haghighat, Maryam, et al.
Published: (2023)
STEP-Parts: Geometric Partitioning of Boundary Representations for Large-Scale CAD Processing
by: Fan, Shen, et al.
Published: (2026)
by: Fan, Shen, et al.
Published: (2026)
Refine and Purify: Orthogonal Basis Optimization with Null-Space Denoising for Conditional Representation Learning
by: Wang, Jiaquan, et al.
Published: (2026)
by: Wang, Jiaquan, et al.
Published: (2026)
CNN-based explanation ensembling for dataset, representation and explanations evaluation
by: Hryniewska-Guzik, Weronika, et al.
Published: (2024)
by: Hryniewska-Guzik, Weronika, et al.
Published: (2024)
Controlling Text-to-Image Diffusion by Orthogonal Finetuning
by: Qiu, Zeju, et al.
Published: (2023)
by: Qiu, Zeju, et al.
Published: (2023)
Revisiting FunnyBirds evaluation framework for prototypical parts networks
by: Opłatek, Szymon, et al.
Published: (2024)
by: Opłatek, Szymon, et al.
Published: (2024)
Steering to Say No: Configurable Refusal via Activation Steering in Vision Language Models
by: Yang, Jiaxi, et al.
Published: (2026)
by: Yang, Jiaxi, et al.
Published: (2026)
Steering Away from Memorization: Reachability-Constrained Reinforcement Learning for Text-to-Image Diffusion
by: Karnik, Sathwik, et al.
Published: (2026)
by: Karnik, Sathwik, et al.
Published: (2026)
MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective
by: Huang, Hailang, et al.
Published: (2024)
by: Huang, Hailang, et al.
Published: (2024)
Rethinking Inter-LoRA Orthogonality in Adapter Merging: Insights from Orthogonal Monte Carlo Dropout
by: Zhang, Andi, et al.
Published: (2025)
by: Zhang, Andi, et al.
Published: (2025)
Symbolic Disentangled Representations for Images
by: Korchemnyi, Alexandr, et al.
Published: (2024)
by: Korchemnyi, Alexandr, et al.
Published: (2024)
KAGE-Bench: Fast Known-Axis Visual Generalization Evaluation for Reinforcement Learning
by: Cherepanov, Egor, et al.
Published: (2026)
by: Cherepanov, Egor, et al.
Published: (2026)
Mars-Bench: A Benchmark for Evaluating Foundation Models for Mars Science Tasks
by: Purohit, Mirali, et al.
Published: (2025)
by: Purohit, Mirali, et al.
Published: (2025)
OMENN: One Matrix to Explain Neural Networks
by: Wróbel, Adam, et al.
Published: (2024)
by: Wróbel, Adam, et al.
Published: (2024)
SIDE: Sparse Information Disentanglement for Explainable Artificial Intelligence
by: Dubovik, Viktar, et al.
Published: (2025)
by: Dubovik, Viktar, et al.
Published: (2025)
T2I-ConBench: Text-to-Image Benchmark for Continual Post-training
by: Huang, Zhehao, et al.
Published: (2025)
by: Huang, Zhehao, et al.
Published: (2025)
The Grammar of Interactive Explanatory Model Analysis
by: Baniecki, Hubert, et al.
Published: (2020)
by: Baniecki, Hubert, et al.
Published: (2020)
Graphic-Design-Bench: A Comprehensive Benchmark for Evaluating AI on Graphic Design Tasks
by: Deganutti, Adrienne, et al.
Published: (2026)
by: Deganutti, Adrienne, et al.
Published: (2026)
Text-to-Image GAN with Pretrained Representations
by: You, Xiaozhou, et al.
Published: (2024)
by: You, Xiaozhou, et al.
Published: (2024)
Learning to Steer: Input-dependent Steering for Multimodal LLMs
by: Parekh, Jayneel, et al.
Published: (2025)
by: Parekh, Jayneel, et al.
Published: (2025)
TraSCE: Trajectory Steering for Concept Erasure
by: Jain, Anubhav, et al.
Published: (2024)
by: Jain, Anubhav, et al.
Published: (2024)
Steering CLIP's vision transformer with sparse autoencoders
by: Joseph, Sonia, et al.
Published: (2025)
by: Joseph, Sonia, et al.
Published: (2025)
StressDream: Steering Video World Models for Robust Policy Evaluation and Improvement
by: Seo, Junwon, et al.
Published: (2026)
by: Seo, Junwon, et al.
Published: (2026)
Similar Items
-
Interpreting CLIP with Hierarchical Sparse Autoencoders
by: Zaigrajew, Vladimir, et al.
Published: (2025) -
Your CLIP has 164 dimensions of noise: Exploring the embeddings covariance eigenspectrum of contrastively pretrained vision-language transformers
by: Grzywaczewski, Jakub, et al.
Published: (2026) -
Adversarial attacks and defenses in explainable artificial intelligence: A survey
by: Baniecki, Hubert, et al.
Published: (2023) -
Birds look like cars: Adversarial analysis of intrinsically interpretable deep learning
by: Baniecki, Hubert, et al.
Published: (2025) -
Red Teaming Models for Hyperspectral Image Analysis Using Explainable AI
by: Zaigrajew, Vladimir, et al.
Published: (2024)