Neural steering vectors reveal dose and exposure-dependent impacts of human-AI relationships
Fuente:
arXiv
Saved in:
| Main Authors: | Kirk, Hannah Rose, Davidson, Henry, Saunders, Ed, Luettgau, Lennart, Vidgen, Bertie, Hale, Scott A., Summerfield, Christopher |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Why human-AI relationships need socioaffective alignment
by: Kirk, Hannah Rose, et al.
Published: (2025)
by: Kirk, Hannah Rose, et al.
Published: (2025)
PRISM-X: Experiments on Personalised Fine-Tuning with Human and Simulated Users
by: Kirk, Hannah Rose, et al.
Published: (2026)
by: Kirk, Hannah Rose, et al.
Published: (2026)
Conversational AI increases political knowledge as effectively as self-directed internet search
by: Luettgau, Lennart, et al.
Published: (2025)
by: Luettgau, Lennart, et al.
Published: (2025)
SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models
by: Vidgen, Bertie, et al.
Published: (2023)
by: Vidgen, Bertie, et al.
Published: (2023)
When Do LLM Preferences Predict Downstream Behavior?
by: Slama, Katarina, et al.
Published: (2026)
by: Slama, Katarina, et al.
Published: (2026)
Ask don't tell: Reducing sycophancy in large language models
by: Dubois, Magda, et al.
Published: (2026)
by: Dubois, Magda, et al.
Published: (2026)
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
by: Röttger, Paul, et al.
Published: (2023)
by: Röttger, Paul, et al.
Published: (2023)
People readily follow personal advice from AI but it does not improve their well-being
by: Luettgau, Lennart, et al.
Published: (2025)
by: Luettgau, Lennart, et al.
Published: (2025)
HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics
by: Luettgau, Lennart, et al.
Published: (2025)
by: Luettgau, Lennart, et al.
Published: (2025)
Multilingual != Multicultural: Evaluating Gaps Between Multilingual Capabilities and Cultural Alignment in LLMs
by: Rystrøm, Jonathan, et al.
Published: (2025)
by: Rystrøm, Jonathan, et al.
Published: (2025)
Measuring and Mitigating Persona Distortions from AI Writing Assistance
by: Röttger, Paul, et al.
Published: (2026)
by: Röttger, Paul, et al.
Published: (2026)
The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models
by: Kirk, Hannah Rose, et al.
Published: (2024)
by: Kirk, Hannah Rose, et al.
Published: (2024)
Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language
by: Summerfield, Christopher, et al.
Published: (2025)
by: Summerfield, Christopher, et al.
Published: (2025)
Vulnerability-Amplifying Interaction Loops: a systematic failure mode in AI chatbot mental-health interactions
by: Weilnhammer, Veith, et al.
Published: (2026)
by: Weilnhammer, Veith, et al.
Published: (2026)
SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety
by: Röttger, Paul, et al.
Published: (2024)
by: Röttger, Paul, et al.
Published: (2024)
Classification is a RAG problem: A case study on hate speech detection
by: Willats, Richard, et al.
Published: (2025)
by: Willats, Richard, et al.
Published: (2025)
Disclosure By Design: Identity Transparency as a Behavioural Property of Conversational AI Models
by: Gausen, Anna, et al.
Published: (2026)
by: Gausen, Anna, et al.
Published: (2026)
One-shot emergency psychiatric triage across 15 frontier AI chatbots
by: Weilnhammer, Veith, et al.
Published: (2026)
by: Weilnhammer, Veith, et al.
Published: (2026)
Reward Model Interpretability via Optimal and Pessimal Tokens
by: Christian, Brian, et al.
Published: (2025)
by: Christian, Brian, et al.
Published: (2025)
Oregon County Library Service, 1965.
by: Davidson, Rose, Ed.
Published: (1966)
by: Davidson, Rose, Ed.
Published: (1966)
Indian-BhED: A Dataset for Measuring India-Centric Biases in Large Language Models
by: Khandelwal, Khyati, et al.
Published: (2023)
by: Khandelwal, Khyati, et al.
Published: (2023)
Technological folie à deux: Feedback Loops Between AI Chatbots and Mental Illness
by: Dohnány, Sebastian, et al.
Published: (2025)
by: Dohnány, Sebastian, et al.
Published: (2025)
WorkBench: a Benchmark Dataset for Agents in a Realistic Workplace Setting
by: Styles, Olly, et al.
Published: (2024)
by: Styles, Olly, et al.
Published: (2024)
Reward Models Inherit Value Biases from Pretraining
by: Christian, Brian, et al.
Published: (2026)
by: Christian, Brian, et al.
Published: (2026)
Robust estimation of causal dose-response relationship using exposure data with dose as an instrumental variable
by: Wang, Jixian, et al.
Published: (2025)
by: Wang, Jixian, et al.
Published: (2025)
The AI Consumer Index (ACE)
by: Benchek, Julien, et al.
Published: (2025)
by: Benchek, Julien, et al.
Published: (2025)
Early learning of the optimal constant solution in neural networks and humans
by: Rubruck, Jirko, et al.
Published: (2024)
by: Rubruck, Jirko, et al.
Published: (2024)
The benefit of dose-exposure-response modeling in the estimation of dose-response relationship and dose optimization: some theoretical and simulation evidence
by: Wang, Jixian, et al.
Published: (2025)
by: Wang, Jixian, et al.
Published: (2025)
Skewed Score: A statistical framework to assess autograders
by: Dubois, Magda, et al.
Published: (2025)
by: Dubois, Magda, et al.
Published: (2025)
The AI Community Building the Future? A Quantitative Analysis of Development Activity on Hugging Face Hub
by: Osborne, Cailean, et al.
Published: (2024)
by: Osborne, Cailean, et al.
Published: (2024)
LINGOLY: A Benchmark of Olympiad-Level Linguistic Reasoning Puzzles in Low-Resource and Extinct Languages
by: Bean, Andrew M., et al.
Published: (2024)
by: Bean, Andrew M., et al.
Published: (2024)
The paradoxical cycle of neutral genesis
by: Summerfield, Joel
Published: (2025)
by: Summerfield, Joel
Published: (2025)
Effects of the Changing Employment Situation on Urban Chinese Woman
by: Summerfield, Gale
Published: (1994)
by: Summerfield, Gale
Published: (1994)
Lie Algebra Decomposition Classes for Reductive Algebraic Groups in Arbitrary Characteristic
by: Summerfield, Joel
Published: (2025)
by: Summerfield, Joel
Published: (2025)
Topics in English for the Secondary School.
by: Summerfield, Geoffrey
Published: (1965)
by: Summerfield, Geoffrey
Published: (1965)
Can sparse autoencoders be used to decompose and interpret steering vectors?
by: Mayne, Harry, et al.
Published: (2024)
by: Mayne, Harry, et al.
Published: (2024)
LMUnit: Fine-grained Evaluation with Natural Language Unit Tests
by: Saad-Falcon, Jon, et al.
Published: (2024)
by: Saad-Falcon, Jon, et al.
Published: (2024)
‘Cheering on from the side‐lines’: The perceived impact of romantic partner's commentary and behaviour on maintaining women's appearance anxiety
by: Gemma Stephanie Lumsdale, et al.
Published: (2024)
by: Gemma Stephanie Lumsdale, et al.
Published: (2024)
Analysis of human steering behavior differences in human-in-control and autonomy-in-control driving
by: Mai, Rene, et al.
Published: (2024)
by: Mai, Rene, et al.
Published: (2024)
Real‐world evidence for computerized insulin dose‐adjustment algorithms in the effective use of continuous glucose monitoring by primary care clinicians
by: Mayer B. Davidson, et al.
Published: (2024)
by: Mayer B. Davidson, et al.
Published: (2024)
Similar Items
-
Why human-AI relationships need socioaffective alignment
by: Kirk, Hannah Rose, et al.
Published: (2025) -
PRISM-X: Experiments on Personalised Fine-Tuning with Human and Simulated Users
by: Kirk, Hannah Rose, et al.
Published: (2026) -
Conversational AI increases political knowledge as effectively as self-directed internet search
by: Luettgau, Lennart, et al.
Published: (2025) -
SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models
by: Vidgen, Bertie, et al.
Published: (2023) -
When Do LLM Preferences Predict Downstream Behavior?
by: Slama, Katarina, et al.
Published: (2026)