Machine Psychology
Fuente:
arXiv
Saved in:
| Main Authors: | Hagendorff, Thilo, Dasgupta, Ishita, Binz, Marcel, Chan, Stephanie C. Y., Lampinen, Andrew, Wang, Jane X., Akata, Zeynep, Schulz, Eric |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Deception Abilities Emerged in Large Language Models
by: Hagendorff, Thilo
Published: (2023)
by: Hagendorff, Thilo
Published: (2023)
Inducing anxiety in large language models can induce bias
by: Coda-Forno, Julian, et al.
Published: (2023)
by: Coda-Forno, Julian, et al.
Published: (2023)
CogBench: a large language model walks into a psychology lab
by: Coda-Forno, Julian, et al.
Published: (2024)
by: Coda-Forno, Julian, et al.
Published: (2024)
A Looming Replication Crisis in Evaluating Behavior in Language Models? Evidence and Solutions
by: Vaugrante, Laurène, et al.
Published: (2024)
by: Vaugrante, Laurène, et al.
Published: (2024)
Just-in-time and distributed task representations in language models
by: Li, Yuxuan, et al.
Published: (2025)
by: Li, Yuxuan, et al.
Published: (2025)
Language models show human-like content effects on reasoning tasks
by: Dasgupta, Ishita, et al.
Published: (2022)
by: Dasgupta, Ishita, et al.
Published: (2022)
Large Reasoning Models Are Autonomous Jailbreak Agents
by: Hagendorff, Thilo, et al.
Published: (2025)
by: Hagendorff, Thilo, et al.
Published: (2025)
Mapping the Ethics of Generative AI: A Comprehensive Scoping Review
by: Hagendorff, Thilo
Published: (2024)
by: Hagendorff, Thilo
Published: (2024)
Building, Reusing, and Generalizing Abstract Representations from Concrete Sequences
by: Wu, Shuchen, et al.
Published: (2024)
by: Wu, Shuchen, et al.
Published: (2024)
Compromising Honesty and Harmlessness in Language Models via Deception Attacks
by: Vaugrante, Laurène, et al.
Published: (2025)
by: Vaugrante, Laurène, et al.
Published: (2025)
Reference-Free Rating of LLM Responses via Latent Information
by: Girrbach, Leander, et al.
Published: (2025)
by: Girrbach, Leander, et al.
Published: (2025)
"Dark Triad" Model Organisms of Misalignment: Narrow Fine-Tuning Mirrors Human Antisocial Behavior
by: Lulla, Roshni, et al.
Published: (2026)
by: Lulla, Roshni, et al.
Published: (2026)
Fairness Hacking: The Malicious Practice of Shrouding Unfairness in Algorithms
by: Meding, Kristof, et al.
Published: (2023)
by: Meding, Kristof, et al.
Published: (2023)
On the Inevitability of Left-Leaning Political Bias in Aligned Language Models
by: Hagendorff, Thilo
Published: (2025)
by: Hagendorff, Thilo
Published: (2025)
Discovering Chunks in Neural Embeddings for Interpretability
by: Wu, Shuchen, et al.
Published: (2025)
by: Wu, Shuchen, et al.
Published: (2025)
SemioLLM: Evaluating Large Language Models for Diagnostic Reasoning from Unstructured Clinical Narratives in Epilepsy
by: Dani, Meghal, et al.
Published: (2024)
by: Dani, Meghal, et al.
Published: (2024)
How do Large Language Models Navigate Conflicts between Honesty and Helpfulness?
by: Liu, Ryan, et al.
Published: (2024)
by: Liu, Ryan, et al.
Published: (2024)
Evaluating Spatial Understanding of Large Language Models
by: Yamada, Yutaro, et al.
Published: (2023)
by: Yamada, Yutaro, et al.
Published: (2023)
The in-context inductive biases of vision-language models differ across modalities
by: Allen, Kelsey, et al.
Published: (2025)
by: Allen, Kelsey, et al.
Published: (2025)
Is This Just Fantasy? Language Model Representations Reflect Human Judgments of Event Plausibility
by: Lepori, Michael A., et al.
Published: (2025)
by: Lepori, Michael A., et al.
Published: (2025)
On the generalization of language models from in-context learning and finetuning: a controlled study
by: Lampinen, Andrew K., et al.
Published: (2025)
by: Lampinen, Andrew K., et al.
Published: (2025)
Beyond Chains of Thought: Benchmarking Latent-Space Reasoning Abilities in Large Language Models
by: Hagendorff, Thilo, et al.
Published: (2025)
by: Hagendorff, Thilo, et al.
Published: (2025)
Context Structure Reshapes the Representational Geometry of Language Models
by: Hosseini, Eghbal A., et al.
Published: (2026)
by: Hosseini, Eghbal A., et al.
Published: (2026)
PRIDE -- Parameter-Efficient Reduction of Identity Discrimination for Equality in LLMs
by: Menke, Maluna, et al.
Published: (2025)
by: Menke, Maluna, et al.
Published: (2025)
Distinct Computations Emerge From Compositional Curricula in In-Context Learning
by: Lee, Jin Hwa, et al.
Published: (2025)
by: Lee, Jin Hwa, et al.
Published: (2025)
A Systematic Comparison of Syllogistic Reasoning in Humans and Language Models
by: Eisape, Tiwalayo, et al.
Published: (2023)
by: Eisape, Tiwalayo, et al.
Published: (2023)
An evolutionary perspective on modes of learning in Transformers
by: Ku, Alexander Y., et al.
Published: (2025)
by: Ku, Alexander Y., et al.
Published: (2025)
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training
by: Kim, Sanghwan, et al.
Published: (2024)
by: Kim, Sanghwan, et al.
Published: (2024)
Are Reasoning LLMs Robust to Interventions on Their Chain-of-Thought?
by: von Recum, Alexander, et al.
Published: (2026)
by: von Recum, Alexander, et al.
Published: (2026)
Meta-learning ecological priors from large language models explains human learning and decision making
by: Jagadish, Akshay K., et al.
Published: (2025)
by: Jagadish, Akshay K., et al.
Published: (2025)
Human-like Category Learning by Injecting Ecological Priors from Large Language Models into Neural Networks
by: Jagadish, Akshay K., et al.
Published: (2024)
by: Jagadish, Akshay K., et al.
Published: (2024)
The Illusion of Latent Generalization: Bi-directionality and the Reversal Curse
by: Coda-Forno, Julian, et al.
Published: (2026)
by: Coda-Forno, Julian, et al.
Published: (2026)
Concept-Guided Interpretability via Neural Chunking
by: Wu, Shuchen, et al.
Published: (2025)
by: Wu, Shuchen, et al.
Published: (2025)
Multimodality and Attention Increase Alignment in Natural Language Prediction Between Humans and Computational Models
by: Kewenig, Viktor, et al.
Published: (2023)
by: Kewenig, Viktor, et al.
Published: (2023)
Naturalistic Computational Cognitive Science: Towards generalizable models and theories that capture the full range of natural behavior
by: Carvalho, Wilka, et al.
Published: (2025)
by: Carvalho, Wilka, et al.
Published: (2025)
Can we automatize scientific discovery in the cognitive sciences?
by: Jagadish, Akshay K., et al.
Published: (2026)
by: Jagadish, Akshay K., et al.
Published: (2026)
Towards a Psychology of Machines: Large Language Models Predict Human Memory
by: Huff, Markus, et al.
Published: (2024)
by: Huff, Markus, et al.
Published: (2024)
Nürnberg NLP at PsyDefDetect: Multi-Axis Voter Ensembles for Psychological Defence Mechanism Classification
by: Steigerwald, Philipp, et al.
Published: (2026)
by: Steigerwald, Philipp, et al.
Published: (2026)
Machine Psychometrics: A Mathematical Psychology of Artificial Intelligence
by: Bogdan, Alex, et al.
Published: (2026)
by: Bogdan, Alex, et al.
Published: (2026)
Evaluation Awareness in Language Models Has Limited Effect on Behaviour
by: Knecht, Amelie, et al.
Published: (2026)
by: Knecht, Amelie, et al.
Published: (2026)
Similar Items
-
Deception Abilities Emerged in Large Language Models
by: Hagendorff, Thilo
Published: (2023) -
Inducing anxiety in large language models can induce bias
by: Coda-Forno, Julian, et al.
Published: (2023) -
CogBench: a large language model walks into a psychology lab
by: Coda-Forno, Julian, et al.
Published: (2024) -
A Looming Replication Crisis in Evaluating Behavior in Language Models? Evidence and Solutions
by: Vaugrante, Laurène, et al.
Published: (2024) -
Just-in-time and distributed task representations in language models
by: Li, Yuxuan, et al.
Published: (2025)