Activation Steering via Generative Causal Mediation
Fuente:
arXiv
Saved in:
| Main Authors: | Sankaranarayanan, Aruna, Zur, Amir, Geiger, Atticus, Hadfield-Menell, Dylan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Causal Reasoning and Large Language Models: Opening a New Frontier for Causality
by: Kıcıman, Emre, et al.
Published: (2023)
by: Kıcıman, Emre, et al.
Published: (2023)
Watching the Watchers: A Comparative Fairness Audit of Cloud-based Content Moderation Services
by: Hartmann, David, et al.
Published: (2024)
by: Hartmann, David, et al.
Published: (2024)
Does Writing with Language Models Reduce Content Diversity?
by: Padmakumar, Vishakh, et al.
Published: (2023)
by: Padmakumar, Vishakh, et al.
Published: (2023)
Humanizing LLMs: A Survey of Psychological Measurements with Tools, Datasets, and Human-Agent Applications
by: Dong, Wenhan, et al.
Published: (2025)
by: Dong, Wenhan, et al.
Published: (2025)
The Moral Gap of Large Language Models
by: Skorski, Maciej, et al.
Published: (2025)
by: Skorski, Maciej, et al.
Published: (2025)
Attention to Non-Adopters
by: Zhou, Kaitlyn, et al.
Published: (2025)
by: Zhou, Kaitlyn, et al.
Published: (2025)
The "Colonial Impulse" of Natural Language Processing: An Audit of Bengali Sentiment Analysis Tools and Their Identity-based Biases
by: Das, Dipto, et al.
Published: (2024)
by: Das, Dipto, et al.
Published: (2024)
LLM-based Cognitive Models of Students with Misconceptions
by: Sonkar, Shashank, et al.
Published: (2024)
by: Sonkar, Shashank, et al.
Published: (2024)
Toward Cultural Interpretability: A Linguistic Anthropological Framework for Describing and Evaluating Large Language Models (LLMs)
by: Jones, Graham M., et al.
Published: (2024)
by: Jones, Graham M., et al.
Published: (2024)
Disjoint Processing Mechanisms of Hierarchical and Linear Grammars in Large Language Models
by: Sankaranarayanan, Aruna, et al.
Published: (2025)
by: Sankaranarayanan, Aruna, et al.
Published: (2025)
Properties and Challenges of LLM-Generated Explanations
by: Kunz, Jenny, et al.
Published: (2024)
by: Kunz, Jenny, et al.
Published: (2024)
EduAgent: Generative Student Agents in Learning
by: Xu, Songlin, et al.
Published: (2024)
by: Xu, Songlin, et al.
Published: (2024)
Beyond Accuracy: Rethinking Hallucination and Regulatory Response in Generative AI
by: Li, Zihao, et al.
Published: (2025)
by: Li, Zihao, et al.
Published: (2025)
ABLEIST: Intersectional Disability Bias in LLM-Generated Hiring Scenarios
by: Phutane, Mahika, et al.
Published: (2025)
by: Phutane, Mahika, et al.
Published: (2025)
"They are uncultured": Unveiling Covert Harms and Social Threats in LLM Generated Conversations
by: Dammu, Preetam Prabhu Srikar, et al.
Published: (2024)
by: Dammu, Preetam Prabhu Srikar, et al.
Published: (2024)
The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas
by: Si, Chenglei, et al.
Published: (2025)
by: Si, Chenglei, et al.
Published: (2025)
Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
by: Si, Chenglei, et al.
Published: (2024)
by: Si, Chenglei, et al.
Published: (2024)
The Odyssey of the Fittest: Can Agents Survive and Still Be Good?
by: Waldner, Dylan, et al.
Published: (2025)
by: Waldner, Dylan, et al.
Published: (2025)
Designing Computational Tools for Exploring Causal Relationships in Qualitative Data
by: Meng, Han, et al.
Published: (2026)
by: Meng, Han, et al.
Published: (2026)
Red-Teaming for Generative AI: Silver Bullet or Security Theater?
by: Feffer, Michael, et al.
Published: (2024)
by: Feffer, Michael, et al.
Published: (2024)
Mapping how LLMs debate societal issues when shadowing human personality traits, sociodemographics and social media behavior
by: Ardebili, Ali Aghazadeh, et al.
Published: (2026)
by: Ardebili, Ali Aghazadeh, et al.
Published: (2026)
GRASP: Deterministic argument ranking in interaction graphs
by: Misra, Diganta, et al.
Published: (2026)
by: Misra, Diganta, et al.
Published: (2026)
MentalChat16K: A Benchmark Dataset for Conversational Mental Health Assistance
by: Xu, Jia, et al.
Published: (2025)
by: Xu, Jia, et al.
Published: (2025)
Think Like a Person Before Responding: A Multi-Faceted Evaluation of Persona-Guided LLMs for Countering Hate
by: Ngueajio, Mikel K., et al.
Published: (2025)
by: Ngueajio, Mikel K., et al.
Published: (2025)
EmoAgent: Assessing and Safeguarding Human-AI Interaction for Mental Health Safety
by: Qiu, Jiahao, et al.
Published: (2025)
by: Qiu, Jiahao, et al.
Published: (2025)
AmarDoctor: An AI-Driven, Multilingual, Voice-Interactive Digital Health Application for Primary Care Triage and Patient Management to Bridge the Digital Health Divide for Bengali Speakers
by: Nahar, Nazmun, et al.
Published: (2025)
by: Nahar, Nazmun, et al.
Published: (2025)
From Measurement to Expertise: Empathetic Expert Adapters for Context-Based Empathy in Conversational AI Agents
by: Shayegani, Erfan, et al.
Published: (2025)
by: Shayegani, Erfan, et al.
Published: (2025)
ProgressGym: Alignment with a Millennium of Moral Progress
by: Qiu, Tianyi, et al.
Published: (2024)
by: Qiu, Tianyi, et al.
Published: (2024)
Sociodemographic Prompting is Not Yet an Effective Approach for Simulating Subjective Judgments with LLMs
by: Sun, Huaman, et al.
Published: (2023)
by: Sun, Huaman, et al.
Published: (2023)
LLM4PM: A case study on using Large Language Models for Process Modeling in Enterprise Organizations
by: Ziche, Clara, et al.
Published: (2024)
by: Ziche, Clara, et al.
Published: (2024)
Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmas
by: Chiu, Yu Ying, et al.
Published: (2025)
by: Chiu, Yu Ying, et al.
Published: (2025)
Asking For It: Question-Answering for Predicting Rule Infractions in Online Content Moderation
by: Samory, Mattia, et al.
Published: (2025)
by: Samory, Mattia, et al.
Published: (2025)
What are human values, and how do we align AI to them?
by: Klingefjord, Oliver, et al.
Published: (2024)
by: Klingefjord, Oliver, et al.
Published: (2024)
The Good, the Bad, and the Ugly: The Role of AI Quality Disclosure in Lie Detection
by: Bhattacharya, Haimanti, et al.
Published: (2024)
by: Bhattacharya, Haimanti, et al.
Published: (2024)
Who is a Better Matchmaker? Human vs. Algorithmic Judge Assignment in a High-Stakes Startup Competition
by: Xi, Sarina, et al.
Published: (2025)
by: Xi, Sarina, et al.
Published: (2025)
Large Language Models Can Infer Personality from Free-Form User Interactions
by: Peters, Heinrich, et al.
Published: (2024)
by: Peters, Heinrich, et al.
Published: (2024)
When "A Helpful Assistant" Is Not Really Helpful: Personas in System Prompts Do Not Improve Performances of Large Language Models
by: Zheng, Mingqian, et al.
Published: (2023)
by: Zheng, Mingqian, et al.
Published: (2023)
Wisdom from Diversity: Bias Mitigation Through Hybrid Human-LLM Crowds
by: Abels, Axel, et al.
Published: (2025)
by: Abels, Axel, et al.
Published: (2025)
Llms, Virtual Users, and Bias: Predicting Any Survey Question Without Human Data
by: Sinacola, Enzo, et al.
Published: (2025)
by: Sinacola, Enzo, et al.
Published: (2025)
PsychoGAT: A Novel Psychological Measurement Paradigm through Interactive Fiction Games with LLM Agents
by: Yang, Qisen, et al.
Published: (2024)
by: Yang, Qisen, et al.
Published: (2024)
Similar Items
-
Causal Reasoning and Large Language Models: Opening a New Frontier for Causality
by: Kıcıman, Emre, et al.
Published: (2023) -
Watching the Watchers: A Comparative Fairness Audit of Cloud-based Content Moderation Services
by: Hartmann, David, et al.
Published: (2024) -
Does Writing with Language Models Reduce Content Diversity?
by: Padmakumar, Vishakh, et al.
Published: (2023) -
Humanizing LLMs: A Survey of Psychological Measurements with Tools, Datasets, and Human-Agent Applications
by: Dong, Wenhan, et al.
Published: (2025) -
The Moral Gap of Large Language Models
by: Skorski, Maciej, et al.
Published: (2025)