Concept-Level Explainability for Auditing & Steering LLM Responses
Fuente:
arXiv
Saved in:
| Main Authors: | Amara, Kenza, Sevastjanova, Rita, El-Assady, Mennatallah |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Challenges and Opportunities in Text Generation Explainability
by: Amara, Kenza, et al.
Published: (2024)
by: Amara, Kenza, et al.
Published: (2024)
SyntaxShap: Syntax-aware Explainability Method for Text Generation
by: Amara, Kenza, et al.
Published: (2024)
by: Amara, Kenza, et al.
Published: (2024)
PleaSQLarify: Visual Pragmatic Repair for Natural Language Database Querying
by: Chan, Robin Shing Moon, et al.
Published: (2026)
by: Chan, Robin Shing Moon, et al.
Published: (2026)
TopoAlign: Topology-Aware Visual Representation Alignment
by: Yan, Xinyuan, et al.
Published: (2026)
by: Yan, Xinyuan, et al.
Published: (2026)
LayerFlow: Layer-wise Exploration of LLM Embeddings using Uncertainty-aware Interlinked Projections
by: Sevastjanova, Rita, et al.
Published: (2025)
by: Sevastjanova, Rita, et al.
Published: (2025)
Why context matters in VQA and Reasoning: Semantic interventions for VLM input modalities
by: Amara, Kenza, et al.
Published: (2024)
by: Amara, Kenza, et al.
Published: (2024)
Cross-Cultural Simulation of Citizen Emotional Responses to Bureaucratic Red Tape Using LLM Agents
by: Ni, Wanchun, et al.
Published: (2026)
by: Ni, Wanchun, et al.
Published: (2026)
Deconstructing Human-AI Collaboration: Agency, Interaction, and Adaptation
by: Holter, Steffen, et al.
Published: (2024)
by: Holter, Steffen, et al.
Published: (2024)
LayerFlow: Layer‐wise Exploration of LLM Embeddings using Uncertainty‐aware Interlinked Projections
by: Rita Sevastjanova, et al.
Published: (2025)
by: Rita Sevastjanova, et al.
Published: (2025)
Explainable Mapper: Charting LLM Embedding Spaces Using Perturbation-Based Explanation and Verification Agents
by: Yan, Xinyuan, et al.
Published: (2025)
by: Yan, Xinyuan, et al.
Published: (2025)
DxHF: Providing High-Quality Human Feedback for LLM Alignment via Interactive Decomposition
by: Shi, Danqing, et al.
Published: (2025)
by: Shi, Danqing, et al.
Published: (2025)
Dia-Lingle: A Gamified Interface for Dialectal Data Collection
by: Sun, Jiugeng, et al.
Published: (2025)
by: Sun, Jiugeng, et al.
Published: (2025)
RepIt: Steering Language Models with Concept-Specific Refusal Vectors
by: Siu, Vincent, et al.
Published: (2025)
by: Siu, Vincent, et al.
Published: (2025)
Understanding Large Language Model Behaviors through Interactive Counterfactual Generation and Analysis
by: Cheng, Furui, et al.
Published: (2024)
by: Cheng, Furui, et al.
Published: (2024)
EasySteer: A Unified Framework for High-Performance and Extensible LLM Steering
by: Xu, Haolei, et al.
Published: (2025)
by: Xu, Haolei, et al.
Published: (2025)
Steering LLM Thinking with Budget Guidance
by: Li, Junyan, et al.
Published: (2025)
by: Li, Junyan, et al.
Published: (2025)
Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering
by: Zhao, Haiyan, et al.
Published: (2025)
by: Zhao, Haiyan, et al.
Published: (2025)
Steering Language Models Before They Speak: Logit-Level Interventions
by: An, Hyeseon, et al.
Published: (2026)
by: An, Hyeseon, et al.
Published: (2026)
Steer Like the LLM: Activation Steering that Mimics Prompting
by: Heyman, Geert, et al.
Published: (2026)
by: Heyman, Geert, et al.
Published: (2026)
generAItor: Tree-in-the-Loop Text Generation for Language Model Explainability and Adaptation
by: Spinner, Thilo, et al.
Published: (2024)
by: Spinner, Thilo, et al.
Published: (2024)
Output Scouting: Auditing Large Language Models for Catastrophic Responses
by: Bell, Andrew, et al.
Published: (2024)
by: Bell, Andrew, et al.
Published: (2024)
Reward Learning from Multiple Feedback Types
by: Metz, Yannick, et al.
Published: (2025)
by: Metz, Yannick, et al.
Published: (2025)
Improving LLM Reasoning through Interpretable Role-Playing Steering
by: Wang, Anyi, et al.
Published: (2025)
by: Wang, Anyi, et al.
Published: (2025)
A Systematic Analysis of the Impact of Persona Steering on LLM Capabilities
by: Chen, Jiaqi, et al.
Published: (2026)
by: Chen, Jiaqi, et al.
Published: (2026)
Critical-Questions-of-Thought: Steering LLM reasoning with Argumentative Querying
by: Castagna, Federico, et al.
Published: (2024)
by: Castagna, Federico, et al.
Published: (2024)
Controllable LLM Reasoning via Sparse Autoencoder-Based Steering
by: Fang, Yi, et al.
Published: (2026)
by: Fang, Yi, et al.
Published: (2026)
When Wording Steers the Evaluation: Framing Bias in LLM judges
by: Hwang, Yerin, et al.
Published: (2026)
by: Hwang, Yerin, et al.
Published: (2026)
Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning
by: Casademunt, Helena, et al.
Published: (2025)
by: Casademunt, Helena, et al.
Published: (2025)
Prompt-Activation Duality: Improving Activation Steering via Attention-Level Interventions
by: Kang, Diancheng, et al.
Published: (2026)
by: Kang, Diancheng, et al.
Published: (2026)
Effectively Steer LLM To Follow Preference via Building Confident Directions
by: Song, Bingqing, et al.
Published: (2025)
by: Song, Bingqing, et al.
Published: (2025)
Equal Access, Unequal Interaction: A Counterfactual Audit of LLM Fairness
by: Amiri-Margavi, Alireza, et al.
Published: (2026)
by: Amiri-Margavi, Alireza, et al.
Published: (2026)
RELIC: Investigating Large Language Model Responses using Self-Consistency
by: Cheng, Furui, et al.
Published: (2023)
by: Cheng, Furui, et al.
Published: (2023)
Steer LLM Latents for Hallucination Detection
by: Park, Seongheon, et al.
Published: (2025)
by: Park, Seongheon, et al.
Published: (2025)
Toward Subtrait-Level Model Explainability in Automated Writing Evaluation
by: Andrade-Lotero, Alejandro, et al.
Published: (2025)
by: Andrade-Lotero, Alejandro, et al.
Published: (2025)
MIRA: A Bilingual Benchmark for Medical Information Response Audit
by: Xu, Mengyu, et al.
Published: (2026)
by: Xu, Mengyu, et al.
Published: (2026)
Explainability-Based Token Replacement on LLM-Generated Text
by: Mohammadi, Hadi, et al.
Published: (2025)
by: Mohammadi, Hadi, et al.
Published: (2025)
Steering Awareness: Detecting Activation Steering from Within
by: Rivera, Joshua Fonseca, et al.
Published: (2025)
by: Rivera, Joshua Fonseca, et al.
Published: (2025)
SEval-Ex: A Statement-Level Framework for Explainable Summarization Evaluation
by: Herserant, Tanguy, et al.
Published: (2025)
by: Herserant, Tanguy, et al.
Published: (2025)
VACP: Visual Analytics Context Protocol
by: Stähle, Tobias, et al.
Published: (2026)
by: Stähle, Tobias, et al.
Published: (2026)
Understanding and Mitigating Dataset Corruption in LLM Steering
by: Anderson, Cullen, et al.
Published: (2026)
by: Anderson, Cullen, et al.
Published: (2026)
Similar Items
-
Challenges and Opportunities in Text Generation Explainability
by: Amara, Kenza, et al.
Published: (2024) -
SyntaxShap: Syntax-aware Explainability Method for Text Generation
by: Amara, Kenza, et al.
Published: (2024) -
PleaSQLarify: Visual Pragmatic Repair for Natural Language Database Querying
by: Chan, Robin Shing Moon, et al.
Published: (2026) -
TopoAlign: Topology-Aware Visual Representation Alignment
by: Yan, Xinyuan, et al.
Published: (2026) -
LayerFlow: Layer-wise Exploration of LLM Embeddings using Uncertainty-aware Interlinked Projections
by: Sevastjanova, Rita, et al.
Published: (2025)