"Dark Triad" Model Organisms of Misalignment: Narrow Fine-Tuning Mirrors Human Antisocial Behavior
Fuente:
arXiv
Saved in:
| Main Authors: | Lulla, Roshni, Collins, Fiona, Parekh, Sanaya, Hagendorff, Thilo, Kaplan, Jonas |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Subject of Emergent Misalignment in Superintelligence: An Anthropological, Cognitive Neuropsychological, Machine-Learning, and Ontological Perspective
by: Imran, Muhammad Osama, et al.
Published: (2025)
by: Imran, Muhammad Osama, et al.
Published: (2025)
A Spatio-Temporal Point Process for Fine-Grained Modeling of Reading Behavior
by: Re, Francesco Ignazio, et al.
Published: (2025)
by: Re, Francesco Ignazio, et al.
Published: (2025)
Mapping of Subjective Accounts into Interpreted Clusters (MOSAIC): Topic Modelling and LLM applied to Stroboscopic Phenomenology
by: Beauté, Romy, et al.
Published: (2025)
by: Beauté, Romy, et al.
Published: (2025)
Exploitation Without Deception: Dark Triad Feature Steering Reveals Separable Antisocial Circuits in Language Models
by: Berg, Cameron, et al.
Published: (2026)
by: Berg, Cameron, et al.
Published: (2026)
Beyond Human-Like Processing: Large Language Models Perform Equivalently on Forward and Backward Scientific Text
by: Luo, Xiaoliang, et al.
Published: (2024)
by: Luo, Xiaoliang, et al.
Published: (2024)
Linguistics and Human Brain: A Perspective of Computational Neuroscience
by: Zhang, Fudong, et al.
Published: (2026)
by: Zhang, Fudong, et al.
Published: (2026)
VLMs May Not Globally Enhance Human Alignment over LLMs During Natural Reading
by: Wu, Jinzhou, et al.
Published: (2026)
by: Wu, Jinzhou, et al.
Published: (2026)
Fine-Tuning Language Models to Know What They Know
by: Park, Sangjun, et al.
Published: (2026)
by: Park, Sangjun, et al.
Published: (2026)
On Agentic Behavioral Modeling
by: Ostwald, Dirk, et al.
Published: (2026)
by: Ostwald, Dirk, et al.
Published: (2026)
Investigating the Timescales of Language Processing with EEG and Language Models
by: Turco, Davide, et al.
Published: (2024)
by: Turco, Davide, et al.
Published: (2024)
The Prompting Brain: Neurocognitive Markers of Expertise in Guiding Large Language Models
by: Al-Khalifa, Hend, et al.
Published: (2025)
by: Al-Khalifa, Hend, et al.
Published: (2025)
LITcoder: A General-Purpose Library for Building and Comparing Encoding Models
by: Binhuraib, Taha, et al.
Published: (2025)
by: Binhuraib, Taha, et al.
Published: (2025)
Independent-Component-Based Encoding Models of Brain Activity During Story Comprehension
by: Hari, Kamya, et al.
Published: (2026)
by: Hari, Kamya, et al.
Published: (2026)
LinBridge: A Learnable Framework for Interpreting Nonlinear Neural Encoding Models
by: Gao, Xiaohui, et al.
Published: (2024)
by: Gao, Xiaohui, et al.
Published: (2024)
Meow-Omni 1: A Multimodal Large Language Model for Feline Ethology
by: Hu, Jucheng, et al.
Published: (2026)
by: Hu, Jucheng, et al.
Published: (2026)
Training-Driven Representational Geometry Modularization Predicts Brain Alignment in Language Models
by: Liu, Yixuan, et al.
Published: (2026)
by: Liu, Yixuan, et al.
Published: (2026)
Mind the Gap: Aligning the Brain with Language Models Requires a Nonlinear and Multimodal Approach
by: Han, Danny Dongyeop, et al.
Published: (2025)
by: Han, Danny Dongyeop, et al.
Published: (2025)
Large Language Model-based FMRI Encoding of Language Functions for Subjects with Neurocognitive Disorder
by: Wang, Yuejiao, et al.
Published: (2024)
by: Wang, Yuejiao, et al.
Published: (2024)
Graph Representations for Reading Comprehension Analysis using Large Language Model and Eye-Tracking Biomarker
by: Zhang, Yuhong, et al.
Published: (2025)
by: Zhang, Yuhong, et al.
Published: (2025)
The Application of Large Language Models on Major Depressive Disorder Support Based on African Natural Products
by: Zou, Linyan
Published: (2025)
by: Zou, Linyan
Published: (2025)
Decoding Probing: Revealing Internal Linguistic Structures in Neural Language Models using Minimal Pairs
by: He, Linyang, et al.
Published: (2024)
by: He, Linyang, et al.
Published: (2024)
Do Large Language Models Think Like the Brain? Sentence-Level Evidences from Layer-Wise Embeddings and fMRI
by: Lei, Yu, et al.
Published: (2025)
by: Lei, Yu, et al.
Published: (2025)
Large Language Models Align with the Human Brain during Creative Thinking
by: Ismayilzada, Mete, et al.
Published: (2026)
by: Ismayilzada, Mete, et al.
Published: (2026)
Comparing Abstraction in Humans and Large Language Models Using Multimodal Serial Reproduction
by: Kumar, Sreejan, et al.
Published: (2024)
by: Kumar, Sreejan, et al.
Published: (2024)
Grounded Computation & Consciousness: A Framework for Exploring Consciousness in Machines & Other Organisms
by: Williams, Ryan
Published: (2024)
by: Williams, Ryan
Published: (2024)
BrainStratify: Coarse-to-Fine Disentanglement of Intracranial Neural Dynamics
by: Zheng, Hui, et al.
Published: (2025)
by: Zheng, Hui, et al.
Published: (2025)
Large Language Models Show Signs of Alignment with Human Neurocognition During Abstract Reasoning
by: Pinier, Christopher, et al.
Published: (2025)
by: Pinier, Christopher, et al.
Published: (2025)
A Comparison of Large Language Model and Human Performance on Random Number Generation Tasks
by: Harrison, Rachel M.
Published: (2024)
by: Harrison, Rachel M.
Published: (2024)
Misalignment Between Backpropagation and the Hierarchy of Brain Responses to Images
by: Raugel, Joséphine, et al.
Published: (2026)
by: Raugel, Joséphine, et al.
Published: (2026)
Speaker effects in language comprehension: An integrative model of language and speaker processing
by: Wu, Hanlin, et al.
Published: (2024)
by: Wu, Hanlin, et al.
Published: (2024)
Decoding Predictive Inference in Visual Language Processing via Spatiotemporal Neural Coherence
by: Borneman, Sean C., et al.
Published: (2025)
by: Borneman, Sean C., et al.
Published: (2025)
Language Writ Large: LLMs, ChatGPT, Grounding, Meaning and Understanding
by: Harnad, Stevan
Published: (2024)
by: Harnad, Stevan
Published: (2024)
The Way We Prompt: Conceptual Blending, Neural Dynamics, and Prompt-Induced Transitions in LLMs
by: Sato, Makoto
Published: (2025)
by: Sato, Makoto
Published: (2025)
Far from the Shallow: Brain-Predictive Reasoning Embedding through Residual Disentanglement
by: He, Linyang, et al.
Published: (2025)
by: He, Linyang, et al.
Published: (2025)
Word Class Representations Spontaneously Emerge from Successor Representations Trained on Natural Language
by: Immertreu, Mathis, et al.
Published: (2026)
by: Immertreu, Mathis, et al.
Published: (2026)
BrainCodec: Neural fMRI codec for the decoding of cognitive brain states
by: Nishimura, Yuto, et al.
Published: (2024)
by: Nishimura, Yuto, et al.
Published: (2024)
Motion Capture Analysis of Verb and Adjective Types in Austrian Sign Language
by: Krebs, Julia, et al.
Published: (2024)
by: Krebs, Julia, et al.
Published: (2024)
Language models and brains align due to more than next-word prediction and word-level information
by: Merlin, Gabriele, et al.
Published: (2022)
by: Merlin, Gabriele, et al.
Published: (2022)
Generative causal testing to bridge data-driven models and scientific theories in language neuroscience
by: Antonello, Richard, et al.
Published: (2024)
by: Antonello, Richard, et al.
Published: (2024)
Large language models are not about natural language
by: Bolhuis, Johan J., et al.
Published: (2025)
by: Bolhuis, Johan J., et al.
Published: (2025)
Similar Items
-
The Subject of Emergent Misalignment in Superintelligence: An Anthropological, Cognitive Neuropsychological, Machine-Learning, and Ontological Perspective
by: Imran, Muhammad Osama, et al.
Published: (2025) -
A Spatio-Temporal Point Process for Fine-Grained Modeling of Reading Behavior
by: Re, Francesco Ignazio, et al.
Published: (2025) -
Mapping of Subjective Accounts into Interpreted Clusters (MOSAIC): Topic Modelling and LLM applied to Stroboscopic Phenomenology
by: Beauté, Romy, et al.
Published: (2025) -
Exploitation Without Deception: Dark Triad Feature Steering Reveals Separable Antisocial Circuits in Language Models
by: Berg, Cameron, et al.
Published: (2026) -
Beyond Human-Like Processing: Large Language Models Perform Equivalently on Forward and Backward Scientific Text
by: Luo, Xiaoliang, et al.
Published: (2024)