Towards Unifying Interpretability and Control: Evaluation via Intervention
Fuente:
arXiv
Salvato in:
| Autori principali: | Bhalla, Usha, Srinivas, Suraj, Ghandeharioun, Asma, Lakkaraju, Himabindu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Discriminative Feature Attributions: Bridging Post Hoc Explainability and Inherent Interpretability
di: Bhalla, Usha, et al.
Pubblicazione: (2023)
di: Bhalla, Usha, et al.
Pubblicazione: (2023)
Evaluating Adversarial Robustness of Concept Representations in Sparse Autoencoders
di: Li, Aaron J., et al.
Pubblicazione: (2025)
di: Li, Aaron J., et al.
Pubblicazione: (2025)
Towards Unified Attribution in Explainable AI, Data-Centric AI, and Mechanistic Interpretability
di: Zhang, Shichang, et al.
Pubblicazione: (2025)
di: Zhang, Shichang, et al.
Pubblicazione: (2025)
Interpreting CLIP with Sparse Linear Concept Embeddings (SpLiCE)
di: Bhalla, Usha, et al.
Pubblicazione: (2024)
di: Bhalla, Usha, et al.
Pubblicazione: (2024)
All Roads Lead to Rome? Exploring Representational Similarities Between Latent Spaces of Generative Image Models
di: Badrinath, Charumathi, et al.
Pubblicazione: (2024)
di: Badrinath, Charumathi, et al.
Pubblicazione: (2024)
Characterizing Data Point Vulnerability via Average-Case Robustness
di: Han, Tessa, et al.
Pubblicazione: (2023)
di: Han, Tessa, et al.
Pubblicazione: (2023)
Operationalizing the Blueprint for an AI Bill of Rights: Recommendations for Practitioners, Researchers, and Policy Makers
di: Oesterling, Alex, et al.
Pubblicazione: (2024)
di: Oesterling, Alex, et al.
Pubblicazione: (2024)
Towards Interpretable Soft Prompts
di: Patel, Oam, et al.
Pubblicazione: (2025)
di: Patel, Oam, et al.
Pubblicazione: (2025)
Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
di: Bhalla, Usha, et al.
Pubblicazione: (2025)
di: Bhalla, Usha, et al.
Pubblicazione: (2025)
Generalized Group Data Attribution
di: Ley, Dan, et al.
Pubblicazione: (2024)
di: Ley, Dan, et al.
Pubblicazione: (2024)
Which Models have Perceptually-Aligned Gradients? An Explanation via Off-Manifold Robustness
di: Srinivas, Suraj, et al.
Pubblicazione: (2023)
di: Srinivas, Suraj, et al.
Pubblicazione: (2023)
Learning Recourse Costs from Pairwise Feature Comparisons
di: Rawal, Kaivalya, et al.
Pubblicazione: (2024)
di: Rawal, Kaivalya, et al.
Pubblicazione: (2024)
Certifying LLM Safety against Adversarial Prompting
di: Kumar, Aounon, et al.
Pubblicazione: (2023)
di: Kumar, Aounon, et al.
Pubblicazione: (2023)
Explaining the Model, Protecting Your Data: Revealing and Mitigating the Data Privacy Risks of Post-Hoc Model Explanations via Membership Inference
di: Huang, Catherine, et al.
Pubblicazione: (2024)
di: Huang, Catherine, et al.
Pubblicazione: (2024)
Interpretability Needs a New Paradigm
di: Madsen, Andreas, et al.
Pubblicazione: (2024)
di: Madsen, Andreas, et al.
Pubblicazione: (2024)
On the Trade-offs between Adversarial Robustness and Actionable Explanations
di: Krishna, Satyapriya, et al.
Pubblicazione: (2023)
di: Krishna, Satyapriya, et al.
Pubblicazione: (2023)
Interpretability Illusions in the Generalization of Simplified Models
di: Friedman, Dan, et al.
Pubblicazione: (2023)
di: Friedman, Dan, et al.
Pubblicazione: (2023)
Monitorability as a Free Gift: How RLVR Spontaneously Aligns Reasoning
di: Xiong, Zidi, et al.
Pubblicazione: (2026)
di: Xiong, Zidi, et al.
Pubblicazione: (2026)
In-Context Unlearning: Language Models as Few Shot Unlearners
di: Pawelczyk, Martin, et al.
Pubblicazione: (2023)
di: Pawelczyk, Martin, et al.
Pubblicazione: (2023)
Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models
di: Ghandeharioun, Asma, et al.
Pubblicazione: (2024)
di: Ghandeharioun, Asma, et al.
Pubblicazione: (2024)
Confronting LLMs with Traditional ML: Rethinking the Fairness of Large Language Models in Tabular Classifications
di: Liu, Yanchen, et al.
Pubblicazione: (2023)
di: Liu, Yanchen, et al.
Pubblicazione: (2023)
Data Poisoning Attacks on Off-Policy Policy Evaluation Methods
di: Lobo, Elita, et al.
Pubblicazione: (2024)
di: Lobo, Elita, et al.
Pubblicazione: (2024)
Who Gets Credit or Blame? Attributing Accountability in Modern AI Systems
di: Zhang, Shichang, et al.
Pubblicazione: (2025)
di: Zhang, Shichang, et al.
Pubblicazione: (2025)
OpenXAI: Towards a Transparent Evaluation of Model Explanations
di: Agarwal, Chirag, et al.
Pubblicazione: (2022)
di: Agarwal, Chirag, et al.
Pubblicazione: (2022)
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models
di: Pawelczyk, Martin, et al.
Pubblicazione: (2024)
di: Pawelczyk, Martin, et al.
Pubblicazione: (2024)
In-Context Explainers: Harnessing LLMs for Explaining Black Box Models
di: Kroeger, Nicholas, et al.
Pubblicazione: (2023)
di: Kroeger, Nicholas, et al.
Pubblicazione: (2023)
Inference-Time Reward Hacking in Large Language Models
di: Khalaf, Hadi, et al.
Pubblicazione: (2025)
di: Khalaf, Hadi, et al.
Pubblicazione: (2025)
The Disagreement Problem in Explainable Machine Learning: A Practitioner's Perspective
di: Krishna, Satyapriya, et al.
Pubblicazione: (2022)
di: Krishna, Satyapriya, et al.
Pubblicazione: (2022)
Follow My Instruction and Spill the Beans: Scalable Data Extraction from Retrieval-Augmented Generation Systems
di: Qi, Zhenting, et al.
Pubblicazione: (2024)
di: Qi, Zhenting, et al.
Pubblicazione: (2024)
Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control
di: Makelov, Aleksandar, et al.
Pubblicazione: (2024)
di: Makelov, Aleksandar, et al.
Pubblicazione: (2024)
When Can Transformers Count to n?
di: Yehudai, Gilad, et al.
Pubblicazione: (2024)
di: Yehudai, Gilad, et al.
Pubblicazione: (2024)
How Post-Training Reshapes LLMs: A Mechanistic View on Knowledge, Truthfulness, Refusal, and Confidence
di: Du, Hongzhe, et al.
Pubblicazione: (2025)
di: Du, Hongzhe, et al.
Pubblicazione: (2025)
A Study on the Calibration of In-context Learning
di: Zhang, Hanlin, et al.
Pubblicazione: (2023)
di: Zhang, Hanlin, et al.
Pubblicazione: (2023)
Manipulating Large Language Models to Increase Product Visibility
di: Kumar, Aounon, et al.
Pubblicazione: (2024)
di: Kumar, Aounon, et al.
Pubblicazione: (2024)
Interpretable Model Drift Detection
di: Panda, Pranoy, et al.
Pubblicazione: (2025)
di: Panda, Pranoy, et al.
Pubblicazione: (2025)
D-PACE: Dynamic Position-Aware Cross-Entropy for Parallel Speculative Drafting
di: Wu, Tianyu, et al.
Pubblicazione: (2026)
di: Wu, Tianyu, et al.
Pubblicazione: (2026)
EvoLM: In Search of Lost Language Model Training Dynamics
di: Qi, Zhenting, et al.
Pubblicazione: (2025)
di: Qi, Zhenting, et al.
Pubblicazione: (2025)
Think Before You Lie: How Reasoning Leads to Honesty
di: Yuan, Ann, et al.
Pubblicazione: (2026)
di: Yuan, Ann, et al.
Pubblicazione: (2026)
Toward Interpretable Evaluation Measures for Time Series Segmentation
di: Chavelli, Félix, et al.
Pubblicazione: (2025)
di: Chavelli, Félix, et al.
Pubblicazione: (2025)
How Much Can We Forget about Data Contamination?
di: Bordt, Sebastian, et al.
Pubblicazione: (2024)
di: Bordt, Sebastian, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Discriminative Feature Attributions: Bridging Post Hoc Explainability and Inherent Interpretability
di: Bhalla, Usha, et al.
Pubblicazione: (2023) -
Evaluating Adversarial Robustness of Concept Representations in Sparse Autoencoders
di: Li, Aaron J., et al.
Pubblicazione: (2025) -
Towards Unified Attribution in Explainable AI, Data-Centric AI, and Mechanistic Interpretability
di: Zhang, Shichang, et al.
Pubblicazione: (2025) -
Interpreting CLIP with Sparse Linear Concept Embeddings (SpLiCE)
di: Bhalla, Usha, et al.
Pubblicazione: (2024) -
All Roads Lead to Rome? Exploring Representational Similarities Between Latent Spaces of Generative Image Models
di: Badrinath, Charumathi, et al.
Pubblicazione: (2024)