Direction-Flipped Influence Audits Reveal Hidden Structure in Moral Choices of LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Blandfort, Phil, Karayil, Tushar, McKenzie, Alex, Pawar, Urja, Graham, Robert, Krasheninnikov, Dmitrii |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Detecting High-Stakes Interactions with Activation Probes
von: McKenzie, Alex, et al.
Veröffentlicht: (2025)
von: McKenzie, Alex, et al.
Veröffentlicht: (2025)
Red-teaming Activation Probes using Prompted LLMs
von: Blandfort, Phil, et al.
Veröffentlicht: (2025)
von: Blandfort, Phil, et al.
Veröffentlicht: (2025)
IT Students Career Confidence and Career Identity During COVID-19
von: McKenzie, Sophie
Veröffentlicht: (2025)
von: McKenzie, Sophie
Veröffentlicht: (2025)
Hidden Persuaders: LLMs' Political Leaning and Their Influence on Voters
von: Potter, Yujin, et al.
Veröffentlicht: (2024)
von: Potter, Yujin, et al.
Veröffentlicht: (2024)
Stress-Testing Capability Elicitation With Password-Locked Models
von: Greenblatt, Ryan, et al.
Veröffentlicht: (2024)
von: Greenblatt, Ryan, et al.
Veröffentlicht: (2024)
Fresh in memory: Training-order recency is linearly encoded in language model activations
von: Krasheninnikov, Dmitrii, et al.
Veröffentlicht: (2025)
von: Krasheninnikov, Dmitrii, et al.
Veröffentlicht: (2025)
AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviors
von: Sheshadri, Abhay, et al.
Veröffentlicht: (2026)
von: Sheshadri, Abhay, et al.
Veröffentlicht: (2026)
DefVerify: Do Hate Speech Models Reflect Their Dataset's Definition?
von: Khurana, Urja, et al.
Veröffentlicht: (2024)
von: Khurana, Urja, et al.
Veröffentlicht: (2024)
Map the Flow: Revealing Hidden Pathways of Information in VideoLLMs
von: Kim, Minji, et al.
Veröffentlicht: (2025)
von: Kim, Minji, et al.
Veröffentlicht: (2025)
Not What, But How: A Communicative Audit of LLM Response Framing
von: Pawar, Siddhesh Milind, et al.
Veröffentlicht: (2026)
von: Pawar, Siddhesh Milind, et al.
Veröffentlicht: (2026)
Direct Ascent Synthesis: Revealing Hidden Generative Capabilities in Discriminative Models
von: Fort, Stanislav, et al.
Veröffentlicht: (2025)
von: Fort, Stanislav, et al.
Veröffentlicht: (2025)
Trust in foundation models and GenAI: A geographic perspective
von: McKenzie, Grant, et al.
Veröffentlicht: (2025)
von: McKenzie, Grant, et al.
Veröffentlicht: (2025)
Normalizing Basis Functions: Approximate Stationary Models for Large Spatial Data
von: Sikorski, Antony, et al.
Veröffentlicht: (2024)
von: Sikorski, Antony, et al.
Veröffentlicht: (2024)
Crowd-Calibrator: Can Annotator Disagreement Inform Calibration in Subjective Tasks?
von: Khurana, Urja, et al.
Veröffentlicht: (2024)
von: Khurana, Urja, et al.
Veröffentlicht: (2024)
Game mechanics for cyber-harm awareness in the metaverse
von: McKenzie, Sophie, et al.
Veröffentlicht: (2025)
von: McKenzie, Sophie, et al.
Veröffentlicht: (2025)
Moral Mazes in the Era of LLMs
von: Nguyen, Dang, et al.
Veröffentlicht: (2026)
von: Nguyen, Dang, et al.
Veröffentlicht: (2026)
Audit Me If You Can: Query-Efficient Active Fairness Auditing of Black-Box LLMs
von: Hartmann, David, et al.
Veröffentlicht: (2026)
von: Hartmann, David, et al.
Veröffentlicht: (2026)
FinAuditing: A Financial Taxonomy-Structured Multi-Document Benchmark for Evaluating LLMs
von: Wang, Yan, et al.
Veröffentlicht: (2025)
von: Wang, Yan, et al.
Veröffentlicht: (2025)
Offscript: Automated Auditing of Instruction Adherence in LLMs
von: Clark, Nicholas, et al.
Veröffentlicht: (2025)
von: Clark, Nicholas, et al.
Veröffentlicht: (2025)
Counterfactual Evaluation Reveals Hidden Capability Profiles in Clinical LLMs and Agents
von: Turk, Matt
Veröffentlicht: (2026)
von: Turk, Matt
Veröffentlicht: (2026)
Morally Programmed LLMs Reshape Human Morality
von: Lyu, Pengzhao, et al.
Veröffentlicht: (2026)
von: Lyu, Pengzhao, et al.
Veröffentlicht: (2026)
From Morality Installation in LLMs to LLMs in Morality-as-a-System
von: Bombaerts, Gunter
Veröffentlicht: (2026)
von: Bombaerts, Gunter
Veröffentlicht: (2026)
Exploring the psychology of LLMs' Moral and Legal Reasoning
von: Almeida, Guilherme F. C. F., et al.
Veröffentlicht: (2023)
von: Almeida, Guilherme F. C. F., et al.
Veröffentlicht: (2023)
BengaliMoralBench: A Benchmark for Auditing Moral Reasoning in Large Language Models within Bengali Language and Culture
von: Ridoy, Shahriyar Zaman, et al.
Veröffentlicht: (2025)
von: Ridoy, Shahriyar Zaman, et al.
Veröffentlicht: (2025)
Sentence-Anchored Gist Compression for Long-Context LLMs
von: Tarasov, Dmitrii, et al.
Veröffentlicht: (2025)
von: Tarasov, Dmitrii, et al.
Veröffentlicht: (2025)
Formal Methods Meet LLMs: Auditing, Monitoring, and Intervention for Compliance of Advanced AI Systems
von: Alamdari, Parand A., et al.
Veröffentlicht: (2026)
von: Alamdari, Parand A., et al.
Veröffentlicht: (2026)
Recognition Without Authorization: LLMs and the Moral Order of Online Advice
von: van Nuenen, Tom
Veröffentlicht: (2026)
von: van Nuenen, Tom
Veröffentlicht: (2026)
MoralBench: Moral Evaluation of LLMs
von: Ji, Jianchao, et al.
Veröffentlicht: (2024)
von: Ji, Jianchao, et al.
Veröffentlicht: (2024)
Lexical Anthropomorphization Influences on Moral Judgments of AI Bad Behavior
von: Banks, Jaime, et al.
Veröffentlicht: (2026)
von: Banks, Jaime, et al.
Veröffentlicht: (2026)
Why are all LLMs Obsessed with Japanese Culture? On the Hidden Cultural and Regional Biases of LLMs
von: de Landa, Joseba Fernandez, et al.
Veröffentlicht: (2026)
von: de Landa, Joseba Fernandez, et al.
Veröffentlicht: (2026)
Libraries of the Future.
von: McKenzie, Jamie
Veröffentlicht: (1996)
von: McKenzie, Jamie
Veröffentlicht: (1996)
Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment
von: Laban, Philippe, et al.
Veröffentlicht: (2023)
von: Laban, Philippe, et al.
Veröffentlicht: (2023)
States Hidden in Hidden States: LLMs Emerge Discrete State Representations Implicitly
von: Chen, Junhao, et al.
Veröffentlicht: (2024)
von: Chen, Junhao, et al.
Veröffentlicht: (2024)
The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs
von: Berezin, Sergey, et al.
Veröffentlicht: (2025)
von: Berezin, Sergey, et al.
Veröffentlicht: (2025)
Disaggregation Reveals Hidden Training Dynamics: The Case of Agreement Attraction
von: Michaelov, James A., et al.
Veröffentlicht: (2025)
von: Michaelov, James A., et al.
Veröffentlicht: (2025)
InverseVis: Revealing the Hidden with Curved Sphere Tracing
von: Kai Lawonn, et al.
Veröffentlicht: (2024)
von: Kai Lawonn, et al.
Veröffentlicht: (2024)
Inside-Out: Hidden Factual Knowledge in LLMs
von: Gekhman, Zorik, et al.
Veröffentlicht: (2025)
von: Gekhman, Zorik, et al.
Veröffentlicht: (2025)
Flipping Knowledge Distillation: Leveraging Small Models' Expertise to Enhance LLMs in Text Matching
von: Li, Mingzhe, et al.
Veröffentlicht: (2025)
von: Li, Mingzhe, et al.
Veröffentlicht: (2025)
Hidden in Plain Text: Emergence & Mitigation of Steganographic Collusion in LLMs
von: Mathew, Yohan, et al.
Veröffentlicht: (2024)
von: Mathew, Yohan, et al.
Veröffentlicht: (2024)
Widespread Gender and Pronoun Bias in Moral Judgments Across LLMs
von: Fernandes, Gustavo Lúcius, et al.
Veröffentlicht: (2026)
von: Fernandes, Gustavo Lúcius, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Detecting High-Stakes Interactions with Activation Probes
von: McKenzie, Alex, et al.
Veröffentlicht: (2025) -
Red-teaming Activation Probes using Prompted LLMs
von: Blandfort, Phil, et al.
Veröffentlicht: (2025) -
IT Students Career Confidence and Career Identity During COVID-19
von: McKenzie, Sophie
Veröffentlicht: (2025) -
Hidden Persuaders: LLMs' Political Leaning and Their Influence on Voters
von: Potter, Yujin, et al.
Veröffentlicht: (2024) -
Stress-Testing Capability Elicitation With Password-Locked Models
von: Greenblatt, Ryan, et al.
Veröffentlicht: (2024)