From Attribution to Action: A Human-Centered Application of Activation Steering
Fuente:
arXiv
Saved in:
| Main Authors: | Labarta, Tobias, Dreyer, Maximilian, Weitz, Katharina, Samek, Wojciech, Lapuschkin, Sebastian |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ProfileXAI: User-Adaptive Explainable AI
by: Corrales, Gilber A., et al.
Published: (2025)
by: Corrales, Gilber A., et al.
Published: (2025)
X-SYS: A Reference Architecture for Interactive Explanation Systems
by: Labarta, Tobias, et al.
Published: (2026)
by: Labarta, Tobias, et al.
Published: (2026)
EchoLSTM: A Self-Reflective Recurrent Network for Stabilizing Long-Range Memory
by: K, Prasanth K, et al.
Published: (2025)
by: K, Prasanth K, et al.
Published: (2025)
Vis-CoT: A Human-in-the-Loop Framework for Interactive Visualization and Intervention in LLM Chain-of-Thought Reasoning
by: Pather, Kaviraj, et al.
Published: (2025)
by: Pather, Kaviraj, et al.
Published: (2025)
See What I Mean? CUE: A Cognitive Model of Understanding Explanations
by: Labarta, Tobias, et al.
Published: (2025)
by: Labarta, Tobias, et al.
Published: (2025)
Spectral Integrated Gradients for Coarse-to-Fine Feature Attribution
by: Kim, Soyeon, et al.
Published: (2026)
by: Kim, Soyeon, et al.
Published: (2026)
Manifold-Aligned Guided Integrated Gradients for Reliable Feature Attribution
by: Kim, Soyeon, et al.
Published: (2026)
by: Kim, Soyeon, et al.
Published: (2026)
Think Thrice Before You Speak: Dual knowledge-enhanced Theory-of-Mind Reasoning for Persuasive Agents
by: Ma, Minghui, et al.
Published: (2026)
by: Ma, Minghui, et al.
Published: (2026)
Normalisation and Initialisation Strategies for Graph Neural Networks in Blockchain Anomaly Detection
by: Duy, Dang Sy, et al.
Published: (2026)
by: Duy, Dang Sy, et al.
Published: (2026)
Structured Basis Function Networks: Loss-Centric Multi-Hypothesis Ensembles with Controllable Diversity
by: Dominguez, Alejandro Rodriguez, et al.
Published: (2025)
by: Dominguez, Alejandro Rodriguez, et al.
Published: (2025)
Think, Act, Learn: A Framework for Autonomous Robotic Agents using Closed-Loop Large Language Models
by: Menon, Anjali R., et al.
Published: (2025)
by: Menon, Anjali R., et al.
Published: (2025)
Multi-Agent Synergy-Driven Iterative Visual Narrative Synthesis
by: Xi, Wang, et al.
Published: (2025)
by: Xi, Wang, et al.
Published: (2025)
Trust and Trustworthiness from Human-Centered Perspective in HRI -- A Systematic Literature Review
by: de Souza, Debora Firmino, et al.
Published: (2025)
by: de Souza, Debora Firmino, et al.
Published: (2025)
Inducing Causal World Models in LLMs for Zero-Shot Physical Reasoning
by: Sharma, Aditya, et al.
Published: (2025)
by: Sharma, Aditya, et al.
Published: (2025)
MRMS-Net and LMRMS-Net: Scalable Multi-Representation Multi-Scale Networks for Time Series Classification
by: Alagöz, Celal, et al.
Published: (2026)
by: Alagöz, Celal, et al.
Published: (2026)
Grounded Gesture Generation: Language, Motion, and Space
by: Deichler, Anna, et al.
Published: (2025)
by: Deichler, Anna, et al.
Published: (2025)
Randomized Spline Trees for Functional Data Classification: Theory and Application to Environmental Time Series
by: Riccio, Donato, et al.
Published: (2024)
by: Riccio, Donato, et al.
Published: (2024)
What Should Explanations Contain? A Human-Centered Explanation Content Model for Local, Post-Hoc Explanations
by: Degen, Helmut
Published: (2026)
by: Degen, Helmut
Published: (2026)
Model Input-Output Configuration Search with Embedded Feature Selection for Sensor Time-series and Image Classification
by: Hoang, Anh T., et al.
Published: (2023)
by: Hoang, Anh T., et al.
Published: (2023)
Approaches to Semantic Textual Similarity in Slovak Language: From Algorithms to Transformers
by: Radosky, Lukas, et al.
Published: (2026)
by: Radosky, Lukas, et al.
Published: (2026)
Polarization-Based Eye Tracking with Personalized Siamese Architectures
by: Kalkanli, Beyza, et al.
Published: (2026)
by: Kalkanli, Beyza, et al.
Published: (2026)
CentaurTA Studio: A Self-Improving Human-Agent Collaboration System for Thematic Analysis
by: Wang, Lei, et al.
Published: (2026)
by: Wang, Lei, et al.
Published: (2026)
AI Model for Predicting Binding Affinity of Antidiabetic Compounds Targeting PPAR
by: Aman, La Ode, et al.
Published: (2024)
by: Aman, La Ode, et al.
Published: (2024)
Evaluating Prompting Strategies for Chart Question Answering with Large Language Models
by: Naikar, Ruthuparna, et al.
Published: (2026)
by: Naikar, Ruthuparna, et al.
Published: (2026)
KGroups: A Versatile Univariate Max-Relevance Min-Redundancy Feature Selection Algorithm for High-dimensional Biological Data
by: Ebiele, Malick, et al.
Published: (2026)
by: Ebiele, Malick, et al.
Published: (2026)
Towards Context-Aware Human-like Pointing Gestures with RL Motion Imitation
by: Deichler, Anna, et al.
Published: (2025)
by: Deichler, Anna, et al.
Published: (2025)
Cross-Domain Malware Detection via Probability-Level Fusion of Lightweight Gradient Boosting Models
by: Mohamed, Omar Khalid Ali
Published: (2025)
by: Mohamed, Omar Khalid Ali
Published: (2025)
On the Equivalence of Regression and Classification
by: Jayadeva, et al.
Published: (2025)
by: Jayadeva, et al.
Published: (2025)
Sparse Projection Oblique Randomer Forests
by: Tomita, Tyler M., et al.
Published: (2015)
by: Tomita, Tyler M., et al.
Published: (2015)
Convexity-Driven Projection for Point Cloud Dimensionality Reduction
by: Sanyal, Suman
Published: (2025)
by: Sanyal, Suman
Published: (2025)
Developing and Evaluating a Design Method for Positive Artificial Intelligence
by: van der Maden, Willem, et al.
Published: (2024)
by: van der Maden, Willem, et al.
Published: (2024)
Can Agentic AI Match the Performance of Human Data Scientists?
by: Luo, An, et al.
Published: (2025)
by: Luo, An, et al.
Published: (2025)
AgentDS Technical Report: Benchmarking the Future of Human-AI Collaboration in Domain-Specific Data Science
by: Luo, An, et al.
Published: (2026)
by: Luo, An, et al.
Published: (2026)
Named entity recognition for Serbian legal documents: Design, methodology and dataset development
by: Kalušev, Vladimir, et al.
Published: (2025)
by: Kalušev, Vladimir, et al.
Published: (2025)
Federated Learning and Class Imbalances
by: Zhu, Siqi, et al.
Published: (2026)
by: Zhu, Siqi, et al.
Published: (2026)
Context-dependent manifold learning: A neuromodulated constrained autoencoder approach
by: Adriaens, Jérôme, et al.
Published: (2026)
by: Adriaens, Jérôme, et al.
Published: (2026)
torchsom: The Reference PyTorch Library for Self-Organizing Maps
by: Berthier, Louis, et al.
Published: (2025)
by: Berthier, Louis, et al.
Published: (2025)
Hybrid activation functions for deep neural networks: S3 and S4 -- a novel approach to gradient flow optimization
by: Kavun, Sergii
Published: (2025)
by: Kavun, Sergii
Published: (2025)
Classifier Calibration at Scale: An Empirical Study of Model-Agnostic Post-Hoc Methods
by: Manokhin, Valery, et al.
Published: (2026)
by: Manokhin, Valery, et al.
Published: (2026)
SLEAN: Simple Lightweight Ensemble Analysis Network for Multi-Provider LLM Coordination: Design, Implementation, and Vibe Coding Bug Investigation Case Study
by: Vargas, Matheus J. T.
Published: (2025)
by: Vargas, Matheus J. T.
Published: (2025)
Similar Items
-
ProfileXAI: User-Adaptive Explainable AI
by: Corrales, Gilber A., et al.
Published: (2025) -
X-SYS: A Reference Architecture for Interactive Explanation Systems
by: Labarta, Tobias, et al.
Published: (2026) -
EchoLSTM: A Self-Reflective Recurrent Network for Stabilizing Long-Range Memory
by: K, Prasanth K, et al.
Published: (2025) -
Vis-CoT: A Human-in-the-Loop Framework for Interactive Visualization and Intervention in LLM Chain-of-Thought Reasoning
by: Pather, Kaviraj, et al.
Published: (2025) -
See What I Mean? CUE: A Cognitive Model of Understanding Explanations
by: Labarta, Tobias, et al.
Published: (2025)