The SaTML '24 CNN Interpretability Competition: New Innovations for Concept-Level Interpretability
Fuente:
arXiv
Saved in:
| Main Authors: | Casper, Stephen, Yun, Jieun, Baek, Joonhyuk, Jung, Yeseong, Kim, Minhwan, Kwon, Kiwan, Park, Saerom, Moore, Hayden, Shriver, David, Connor, Marissa, Grimes, Keltin, Nicolson, Angus, Tagade, Arush, Rumbelow, Jessica, Nguyen, Hieu Minh, Hadfield-Menell, Dylan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Concept-ROT: Poisoning Concepts in Large Language Models with Model Editing
by: Grimes, Keltin, et al.
Published: (2024)
by: Grimes, Keltin, et al.
Published: (2024)
JuliaTrustworthyAI/CounterfactualTraining.jl: SaTML camera-ready
by: Patrick Altmeyer, et al.
Published: (2026)
by: Patrick Altmeyer, et al.
Published: (2026)
Dataset and Lessons Learned from the 2024 SaTML LLM Capture-the-Flag Competition
by: Debenedetti, Edoardo, et al.
Published: (2024)
by: Debenedetti, Edoardo, et al.
Published: (2024)
MIDST Challenge at SaTML 2025: Membership Inference over Diffusion-models-based Synthetic Tabular data
by: Shafieinejad, Masoumeh, et al.
Published: (2026)
by: Shafieinejad, Masoumeh, et al.
Published: (2026)
LLM Assertiveness can be Mechanistically Decomposed into Emotional and Logical Components
by: Tsujimura, Hikaru, et al.
Published: (2025)
by: Tsujimura, Hikaru, et al.
Published: (2025)
Exemplar Partitioning for Mechanistic Interpretability
by: Rumbelow, Jessica
Published: (2026)
by: Rumbelow, Jessica
Published: (2026)
Benchmarking the Discovery Engine
by: Foxabbott, Jack, et al.
Published: (2025)
by: Foxabbott, Jack, et al.
Published: (2025)
Randomness, Not Representation: The Unreliability of Evaluating Cultural Alignment in LLMs
by: Khan, Ariba, et al.
Published: (2025)
by: Khan, Ariba, et al.
Published: (2025)
Pitfalls of Evidence-Based AI Policy
by: Casper, Stephen, et al.
Published: (2025)
by: Casper, Stephen, et al.
Published: (2025)
Explaining Surface Layer Theory Departures in Marine Flux Profiles with Data-Driven Discovery
by: Foxabbott, Jack, et al.
Published: (2025)
by: Foxabbott, Jack, et al.
Published: (2025)
Defending Against Unforeseen Failure Modes with Latent Adversarial Training
by: Casper, Stephen, et al.
Published: (2024)
by: Casper, Stephen, et al.
Published: (2024)
Eight Methods to Evaluate Robust Unlearning in LLMs
by: Lynch, Aengus, et al.
Published: (2024)
by: Lynch, Aengus, et al.
Published: (2024)
Gone but Not Forgotten: Improved Benchmarks for Machine Unlearning
by: Grimes, Keltin, et al.
Published: (2024)
by: Grimes, Keltin, et al.
Published: (2024)
Distributional Process Reward Models: Calibrated Prediction of Future Rewards via Conditional Optimal Transport
by: Ma, Rachel, et al.
Published: (2026)
by: Ma, Rachel, et al.
Published: (2026)
CALMA: A Process for Deriving Context-aligned Axes for Language Model Alignment
by: Soni, Prajna, et al.
Published: (2025)
by: Soni, Prajna, et al.
Published: (2025)
Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases
by: Hahm, Dongyoon, et al.
Published: (2026)
by: Hahm, Dongyoon, et al.
Published: (2026)
Disjoint Processing Mechanisms of Hierarchical and Linear Grammars in Large Language Models
by: Sankaranarayanan, Aruna, et al.
Published: (2025)
by: Sankaranarayanan, Aruna, et al.
Published: (2025)
Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF
by: Siththaranjan, Anand, et al.
Published: (2023)
by: Siththaranjan, Anand, et al.
Published: (2023)
Prompt Injection as Role Confusion
by: Ye, Charles, et al.
Published: (2026)
by: Ye, Charles, et al.
Published: (2026)
Generic Reduction-Based Interpreters (Extended Version)
by: Bach, Casper
Published: (2025)
by: Bach, Casper
Published: (2025)
Diverse Preference Learning for Capabilities and Alignment
by: Slocum, Stewart, et al.
Published: (2025)
by: Slocum, Stewart, et al.
Published: (2025)
Cooperative Inverse Reinforcement Learning
by: Hadfield-Menell, Dylan, et al.
Published: (2016)
by: Hadfield-Menell, Dylan, et al.
Published: (2016)
Flexible Agent Alignment with Goal Inference from Open-Ended Dialog
by: Ma, Rachel, et al.
Published: (2025)
by: Ma, Rachel, et al.
Published: (2025)
Layered Unlearning for Adversarial Relearning
by: Qian, Timothy, et al.
Published: (2025)
by: Qian, Timothy, et al.
Published: (2025)
Goal Inference from Open-Ended Dialog
by: Ma, Rachel, et al.
Published: (2024)
by: Ma, Rachel, et al.
Published: (2024)
Activation Steering via Generative Causal Mediation
by: Sankaranarayanan, Aruna, et al.
Published: (2026)
by: Sankaranarayanan, Aruna, et al.
Published: (2026)
Safety Drift After Fine-Tuning: Evidence from High-Stakes Domains
by: Khan, Emaan Bilal, et al.
Published: (2026)
by: Khan, Emaan Bilal, et al.
Published: (2026)
Back to Blackwell: Closing the Loop on Intransitivity in Multi-Objective Preference Fine-Tuning
by: Zhang, Jiahao, et al.
Published: (2026)
by: Zhang, Jiahao, et al.
Published: (2026)
Median Mishaps between Chirality and Spin-Orbit Torques via Asymmetric Hysteresis
by: Kim, Minhwan, et al.
Published: (2024)
by: Kim, Minhwan, et al.
Published: (2024)
AUTOCT: Automating Interpretable Clinical Trial Prediction with LLM Agents
by: Liu, Fengze, et al.
Published: (2025)
by: Liu, Fengze, et al.
Published: (2025)
TextCAVs: Debugging vision models using text
by: Nicolson, Angus, et al.
Published: (2024)
by: Nicolson, Angus, et al.
Published: (2024)
Teaching Counselors‐in‐Training to Broach Race: An Interpretative Phenomenological Analysis of Black Counselor Educators’ Experiences
by: Jacoby Loury, et al.
Published: (2026)
by: Jacoby Loury, et al.
Published: (2026)
Age determination of sediment core TML1
by: Lamy, Frank, et al.
Published: (2010)
by: Lamy, Frank, et al.
Published: (2010)
Formal Contracts Mitigate Social Dilemmas in Multi-Agent RL
by: Haupt, Andreas A., et al.
Published: (2022)
by: Haupt, Andreas A., et al.
Published: (2022)
$α$ Effect and Magnetic Diffusivity $β$ in Helical Plasma under Turbulence Growth
by: Park, Kiwan
Published: (2024)
by: Park, Kiwan
Published: (2024)
Effect of Turbulent Kinetic Helicity on Diffusive beta effect for Large Scale Dynamo
by: Park, Kiwan
Published: (2024)
by: Park, Kiwan
Published: (2024)
Magnetic Field Amplification and Reconstruction in Rotating Astrophysical Plasmas: Verifying the Roles of $α$ and $β$ in Dynamo Action
by: Park, Kiwan
Published: (2025)
by: Park, Kiwan
Published: (2025)
Analytical approach to the design of RF photoinjector
by: Park, Kiwan
Published: (2024)
by: Park, Kiwan
Published: (2024)
Open Problems in Mechanistic Interpretability
by: Sharkey, Lee, et al.
Published: (2025)
by: Sharkey, Lee, et al.
Published: (2025)
Tropical Geometric Tools for Machine Learning: the TML package
by: Barnhill, David, et al.
Published: (2023)
by: Barnhill, David, et al.
Published: (2023)
Similar Items
-
Concept-ROT: Poisoning Concepts in Large Language Models with Model Editing
by: Grimes, Keltin, et al.
Published: (2024) -
JuliaTrustworthyAI/CounterfactualTraining.jl: SaTML camera-ready
by: Patrick Altmeyer, et al.
Published: (2026) -
Dataset and Lessons Learned from the 2024 SaTML LLM Capture-the-Flag Competition
by: Debenedetti, Edoardo, et al.
Published: (2024) -
MIDST Challenge at SaTML 2025: Membership Inference over Diffusion-models-based Synthetic Tabular data
by: Shafieinejad, Masoumeh, et al.
Published: (2026) -
LLM Assertiveness can be Mechanistically Decomposed into Emotional and Logical Components
by: Tsujimura, Hikaru, et al.
Published: (2025)