Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Summerfield, Christopher, Luettgau, Lennart, Dubois, Magda, Kirk, Hannah Rose, Hackenburg, Kobi, Fist, Catherine, Slama, Katarina, Ding, Nicola, Anselmetti, Rebecca, Strait, Andrew, Giulianelli, Mario, Ududec, Cozmin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Ask don't tell: Reducing sycophancy in large language models
von: Dubois, Magda, et al.
Veröffentlicht: (2026)
von: Dubois, Magda, et al.
Veröffentlicht: (2026)
HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics
von: Luettgau, Lennart, et al.
Veröffentlicht: (2025)
von: Luettgau, Lennart, et al.
Veröffentlicht: (2025)
Skewed Score: A statistical framework to assess autograders
von: Dubois, Magda, et al.
Veröffentlicht: (2025)
von: Dubois, Magda, et al.
Veröffentlicht: (2025)
Measuring and Mitigating Persona Distortions from AI Writing Assistance
von: Röttger, Paul, et al.
Veröffentlicht: (2026)
von: Röttger, Paul, et al.
Veröffentlicht: (2026)
Conversational AI increases political knowledge as effectively as self-directed internet search
von: Luettgau, Lennart, et al.
Veröffentlicht: (2025)
von: Luettgau, Lennart, et al.
Veröffentlicht: (2025)
When Do LLM Preferences Predict Downstream Behavior?
von: Slama, Katarina, et al.
Veröffentlicht: (2026)
von: Slama, Katarina, et al.
Veröffentlicht: (2026)
From Forest to Zoo: Great Ape Behavior Recognition with ChimpBehave
von: Fuchs, Michael, et al.
Veröffentlicht: (2024)
von: Fuchs, Michael, et al.
Veröffentlicht: (2024)
Neural steering vectors reveal dose and exposure-dependent impacts of human-AI relationships
von: Kirk, Hannah Rose, et al.
Veröffentlicht: (2025)
von: Kirk, Hannah Rose, et al.
Veröffentlicht: (2025)
People readily follow personal advice from AI but it does not improve their well-being
von: Luettgau, Lennart, et al.
Veröffentlicht: (2025)
von: Luettgau, Lennart, et al.
Veröffentlicht: (2025)
Artificial intelligence can persuade people to take political actions
von: Hackenburg, Kobi, et al.
Veröffentlicht: (2026)
von: Hackenburg, Kobi, et al.
Veröffentlicht: (2026)
The Levers of Political Persuasion with Conversational AI
von: Hackenburg, Kobi, et al.
Veröffentlicht: (2025)
von: Hackenburg, Kobi, et al.
Veröffentlicht: (2025)
Quest of Identity in Eugene O’ Neil’s The Hairy Ape
von: Peer Salim Jahangeer
Veröffentlicht: (2017)
von: Peer Salim Jahangeer
Veröffentlicht: (2017)
Log analysis is necessary for credible evaluation of AI agents
von: Kirgis, Peter, et al.
Veröffentlicht: (2026)
von: Kirgis, Peter, et al.
Veröffentlicht: (2026)
Vulnerability-Amplifying Interaction Loops: a systematic failure mode in AI chatbot mental-health interactions
von: Weilnhammer, Veith, et al.
Veröffentlicht: (2026)
von: Weilnhammer, Veith, et al.
Veröffentlicht: (2026)
Seven simple steps for log analysis in AI systems
von: Dubois, Magda, et al.
Veröffentlicht: (2026)
von: Dubois, Magda, et al.
Veröffentlicht: (2026)
A Multi-Turn Framework for Evaluating AI Misuse in Fraud and Cybercrime Scenarios
von: Mai, Kimberly T., et al.
Veröffentlicht: (2026)
von: Mai, Kimberly T., et al.
Veröffentlicht: (2026)
Disclosure By Design: Identity Transparency as a Behavioural Property of Conversational AI Models
von: Gausen, Anna, et al.
Veröffentlicht: (2026)
von: Gausen, Anna, et al.
Veröffentlicht: (2026)
One-shot emergency psychiatric triage across 15 frontier AI chatbots
von: Weilnhammer, Veith, et al.
Veröffentlicht: (2026)
von: Weilnhammer, Veith, et al.
Veröffentlicht: (2026)
Why human-AI relationships need socioaffective alignment
von: Kirk, Hannah Rose, et al.
Veröffentlicht: (2025)
von: Kirk, Hannah Rose, et al.
Veröffentlicht: (2025)
Parallel and Distributed Chimp-Optimized LSTM
von: Zisong Wang
Veröffentlicht: (2025)
von: Zisong Wang
Veröffentlicht: (2025)
Reward Model Interpretability via Optimal and Pessimal Tokens
von: Christian, Brian, et al.
Veröffentlicht: (2025)
von: Christian, Brian, et al.
Veröffentlicht: (2025)
The Spooning Ape
von: Robin, Nicolas
Veröffentlicht: (2025)
von: Robin, Nicolas
Veröffentlicht: (2025)
Xouth, The Ape
von: Pitsipios, Iakovos
Veröffentlicht: (2025)
von: Pitsipios, Iakovos
Veröffentlicht: (2025)
AlphaChimp: Tracking and Behavior Recognition of Chimpanzees
von: Ma, Xiaoxuan, et al.
Veröffentlicht: (2024)
von: Ma, Xiaoxuan, et al.
Veröffentlicht: (2024)
Evidence of a log scaling law for political persuasion with large language models
von: Hackenburg, Kobi, et al.
Veröffentlicht: (2024)
von: Hackenburg, Kobi, et al.
Veröffentlicht: (2024)
ChimpVLM: Ethogram-Enhanced Chimpanzee Behaviour Recognition
von: Brookes, Otto, et al.
Veröffentlicht: (2024)
von: Brookes, Otto, et al.
Veröffentlicht: (2024)
Technological folie à deux: Feedback Loops Between AI Chatbots and Mental Illness
von: Dohnány, Sebastian, et al.
Veröffentlicht: (2025)
von: Dohnány, Sebastian, et al.
Veröffentlicht: (2025)
IssueBench: Millions of Realistic Prompts for Measuring Issue Bias in LLM Writing Assistance
von: Röttger, Paul, et al.
Veröffentlicht: (2025)
von: Röttger, Paul, et al.
Veröffentlicht: (2025)
PRISM-X: Experiments on Personalised Fine-Tuning with Human and Simulated Users
von: Kirk, Hannah Rose, et al.
Veröffentlicht: (2026)
von: Kirk, Hannah Rose, et al.
Veröffentlicht: (2026)
La Red ¿democrática? en la sociedad del aislamiento
von: Santiago Giulianelli
Veröffentlicht: (2005)
von: Santiago Giulianelli
Veröffentlicht: (2005)
Reward Models Inherit Value Biases from Pretraining
von: Christian, Brian, et al.
Veröffentlicht: (2026)
von: Christian, Brian, et al.
Veröffentlicht: (2026)
Dopamine D1Aa and D2a Receptor Expression in the Auditory System of a Vocal Fish
von: Kobi Kobi, et al.
Veröffentlicht: (2026)
von: Kobi Kobi, et al.
Veröffentlicht: (2026)
The AI Community Building the Future? A Quantitative Analysis of Development Activity on Hugging Face Hub
von: Osborne, Cailean, et al.
Veröffentlicht: (2024)
von: Osborne, Cailean, et al.
Veröffentlicht: (2024)
Multilingual != Multicultural: Evaluating Gaps Between Multilingual Capabilities and Cultural Alignment in LLMs
von: Rystrøm, Jonathan, et al.
Veröffentlicht: (2025)
von: Rystrøm, Jonathan, et al.
Veröffentlicht: (2025)
SEB-ChOA: An Improved Chimp Optimization Algorithm Using Spiral Exploitation Behavior
von: Qian, Leren, et al.
Veröffentlicht: (2025)
von: Qian, Leren, et al.
Veröffentlicht: (2025)
The paradoxical cycle of neutral genesis
von: Summerfield, Joel
Veröffentlicht: (2025)
von: Summerfield, Joel
Veröffentlicht: (2025)
Effects of the Changing Employment Situation on Urban Chinese Woman
von: Summerfield, Gale
Veröffentlicht: (1994)
von: Summerfield, Gale
Veröffentlicht: (1994)
Lie Algebra Decomposition Classes for Reductive Algebraic Groups in Arbitrary Characteristic
von: Summerfield, Joel
Veröffentlicht: (2025)
von: Summerfield, Joel
Veröffentlicht: (2025)
Topics in English for the Secondary School.
von: Summerfield, Geoffrey
Veröffentlicht: (1965)
von: Summerfield, Geoffrey
Veröffentlicht: (1965)
Open-World Evaluations for Measuring Frontier AI Capabilities
von: Kapoor, Sayash, et al.
Veröffentlicht: (2026)
von: Kapoor, Sayash, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Ask don't tell: Reducing sycophancy in large language models
von: Dubois, Magda, et al.
Veröffentlicht: (2026) -
HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics
von: Luettgau, Lennart, et al.
Veröffentlicht: (2025) -
Skewed Score: A statistical framework to assess autograders
von: Dubois, Magda, et al.
Veröffentlicht: (2025) -
Measuring and Mitigating Persona Distortions from AI Writing Assistance
von: Röttger, Paul, et al.
Veröffentlicht: (2026) -
Conversational AI increases political knowledge as effectively as self-directed internet search
von: Luettgau, Lennart, et al.
Veröffentlicht: (2025)