Deception in LLMs: Self-Preservation and Autonomous Goals in Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Barkur, Sudarshan Kamath, Schacht, Sigurd, Scholl, Johannes |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PHOENIX: Open-Source Language Adaption for Direct Preference Optimization
von: Uhlig, Matthias, et al.
Veröffentlicht: (2024)
von: Uhlig, Matthias, et al.
Veröffentlicht: (2024)
Inference Optimizations for Large Language Models: Effects, Challenges, and Practical Considerations
von: Donisch, Leo, et al.
Veröffentlicht: (2024)
von: Donisch, Leo, et al.
Veröffentlicht: (2024)
A Review of Common Online Speaker Diarization Methods
von: Aperdannier, Roman, et al.
Veröffentlicht: (2024)
von: Aperdannier, Roman, et al.
Veröffentlicht: (2024)
Systematic Evaluation of Online Speaker Diarization Systems Regarding their Latency
von: Aperdannier, Roman, et al.
Veröffentlicht: (2024)
von: Aperdannier, Roman, et al.
Veröffentlicht: (2024)
An approach to optimize inference of the DIART speaker diarization pipeline
von: Aperdannier, Roman, et al.
Veröffentlicht: (2024)
von: Aperdannier, Roman, et al.
Veröffentlicht: (2024)
Do Large Language Models Exhibit Spontaneous Rational Deception?
von: Taylor, Samuel M., et al.
Veröffentlicht: (2025)
von: Taylor, Samuel M., et al.
Veröffentlicht: (2025)
Deception Abilities Emerged in Large Language Models
von: Hagendorff, Thilo
Veröffentlicht: (2023)
von: Hagendorff, Thilo
Veröffentlicht: (2023)
Scope Ambiguities in Large Language Models
von: Kamath, Gaurav, et al.
Veröffentlicht: (2024)
von: Kamath, Gaurav, et al.
Veröffentlicht: (2024)
D-REX: A Benchmark for Detecting Deceptive Reasoning in Large Language Models
von: Krishna, Satyapriya, et al.
Veröffentlicht: (2025)
von: Krishna, Satyapriya, et al.
Veröffentlicht: (2025)
Self-Steering Optimization: Autonomous Preference Optimization for Large Language Models
von: Xiang, Hao, et al.
Veröffentlicht: (2024)
von: Xiang, Hao, et al.
Veröffentlicht: (2024)
eDIF: A European Deep Inference Fabric for Remote Interpretability of LLM
von: Guggenberger, Irma Heithoff. Marc, et al.
Veröffentlicht: (2025)
von: Guggenberger, Irma Heithoff. Marc, et al.
Veröffentlicht: (2025)
Unmasking the Shadows of AI: Investigating Deceptive Capabilities in Large Language Models
von: Guo, Linge
Veröffentlicht: (2024)
von: Guo, Linge
Veröffentlicht: (2024)
Language Models Largely Exhibit Human-like Constituent Ordering Preferences
von: Tur, Ada Defne, et al.
Veröffentlicht: (2025)
von: Tur, Ada Defne, et al.
Veröffentlicht: (2025)
To Tell The Truth: Language of Deception and Language Models
von: Hazra, Sanchaita, et al.
Veröffentlicht: (2023)
von: Hazra, Sanchaita, et al.
Veröffentlicht: (2023)
Preserving Knowledge in Large Language Model with Model-Agnostic Self-Decompression
von: Zhang, Zilun, et al.
Veröffentlicht: (2024)
von: Zhang, Zilun, et al.
Veröffentlicht: (2024)
From Deception to Detection: The Dual Roles of Large Language Models in Fake News
von: Sallami, Dorsaf, et al.
Veröffentlicht: (2024)
von: Sallami, Dorsaf, et al.
Veröffentlicht: (2024)
Creating a Fine Grained Entity Type Taxonomy Using LLMs
von: Gunn, Michael, et al.
Veröffentlicht: (2024)
von: Gunn, Michael, et al.
Veröffentlicht: (2024)
Evaluating the Goal-Directedness of Large Language Models
von: Everitt, Tom, et al.
Veröffentlicht: (2025)
von: Everitt, Tom, et al.
Veröffentlicht: (2025)
Advancing NLP Security by Leveraging LLMs as Adversarial Engines
von: Srinivasan, Sudarshan, et al.
Veröffentlicht: (2024)
von: Srinivasan, Sudarshan, et al.
Veröffentlicht: (2024)
Defending Large Language Models Against Jailbreaking Attacks Through Goal Prioritization
von: Zhang, Zhexin, et al.
Veröffentlicht: (2023)
von: Zhang, Zhexin, et al.
Veröffentlicht: (2023)
Goal Hijacking Attack on Large Language Models via Pseudo-Conversation Injection
von: Chen, Zheng, et al.
Veröffentlicht: (2024)
von: Chen, Zheng, et al.
Veröffentlicht: (2024)
Does Synthetic Data Help Named Entity Recognition for Low-Resource Languages?
von: Kamath, Gaurav, et al.
Veröffentlicht: (2025)
von: Kamath, Gaurav, et al.
Veröffentlicht: (2025)
Unmasking Deceptive Visuals: Benchmarking Multimodal Large Language Models on Misleading Chart Question Answering
von: Chen, Zixin, et al.
Veröffentlicht: (2025)
von: Chen, Zixin, et al.
Veröffentlicht: (2025)
Application of Multimodal Large Language Models in Autonomous Driving
von: Islam, Md Robiul
Veröffentlicht: (2024)
von: Islam, Md Robiul
Veröffentlicht: (2024)
Context-Preserving Tensorial Reconfiguration in Large Language Model Training
von: Tonix, Larin, et al.
Veröffentlicht: (2025)
von: Tonix, Larin, et al.
Veröffentlicht: (2025)
Self-Prompt Tuning: Enable Autonomous Role-Playing in LLMs
von: Kong, Aobo, et al.
Veröffentlicht: (2024)
von: Kong, Aobo, et al.
Veröffentlicht: (2024)
Compromising Honesty and Harmlessness in Language Models via Deception Attacks
von: Vaugrante, Laurène, et al.
Veröffentlicht: (2025)
von: Vaugrante, Laurène, et al.
Veröffentlicht: (2025)
The Perils of Chart Deception: How Misleading Visualizations Affect Vision-Language Models
von: Mahbub, Ridwan, et al.
Veröffentlicht: (2025)
von: Mahbub, Ridwan, et al.
Veröffentlicht: (2025)
An Assessment of Model-On-Model Deception
von: Heitkoetter, Julius, et al.
Veröffentlicht: (2024)
von: Heitkoetter, Julius, et al.
Veröffentlicht: (2024)
Hidden in Plain Sight: Evaluation of the Deception Detection Capabilities of LLMs in Multimodal Settings
von: Miah, Md Messal Monem, et al.
Veröffentlicht: (2025)
von: Miah, Md Messal Monem, et al.
Veröffentlicht: (2025)
Plan-Grounded Large Language Models for Dual Goal Conversational Settings
von: Glória-Silva, Diogo, et al.
Veröffentlicht: (2024)
von: Glória-Silva, Diogo, et al.
Veröffentlicht: (2024)
Seamless Deception: Larger Language Models Are Better Knowledge Concealers
von: Ashok, Dhananjay, et al.
Veröffentlicht: (2026)
von: Ashok, Dhananjay, et al.
Veröffentlicht: (2026)
SelfGoal: Your Language Agents Already Know How to Achieve High-level Goals
von: Yang, Ruihan, et al.
Veröffentlicht: (2024)
von: Yang, Ruihan, et al.
Veröffentlicht: (2024)
Why is "Chicago" Predictive of Deceptive Reviews? Using LLMs to Discover Language Phenomena from Lexical Cues
von: Qu, Jiaming, et al.
Veröffentlicht: (2025)
von: Qu, Jiaming, et al.
Veröffentlicht: (2025)
Robust Utility-Preserving Text Anonymization Based on Large Language Models
von: Yang, Tianyu, et al.
Veröffentlicht: (2024)
von: Yang, Tianyu, et al.
Veröffentlicht: (2024)
Privacy-Preserving Instructions for Aligning Large Language Models
von: Yu, Da, et al.
Veröffentlicht: (2024)
von: Yu, Da, et al.
Veröffentlicht: (2024)
Too Big to Fool: Resisting Deception in Language Models
von: Samsami, Mohammad Reza, et al.
Veröffentlicht: (2024)
von: Samsami, Mohammad Reza, et al.
Veröffentlicht: (2024)
Can Large Language Model Summarizers Adapt to Diverse Scientific Communication Goals?
von: Fonseca, Marcio, et al.
Veröffentlicht: (2024)
von: Fonseca, Marcio, et al.
Veröffentlicht: (2024)
Towards Goal-oriented Prompt Engineering for Large Language Models: A Survey
von: Li, Haochen, et al.
Veröffentlicht: (2024)
von: Li, Haochen, et al.
Veröffentlicht: (2024)
Large Language Models Reveal Information Operation Goals, Tactics, and Narrative Frames
von: Burghardt, Keith, et al.
Veröffentlicht: (2024)
von: Burghardt, Keith, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
PHOENIX: Open-Source Language Adaption for Direct Preference Optimization
von: Uhlig, Matthias, et al.
Veröffentlicht: (2024) -
Inference Optimizations for Large Language Models: Effects, Challenges, and Practical Considerations
von: Donisch, Leo, et al.
Veröffentlicht: (2024) -
A Review of Common Online Speaker Diarization Methods
von: Aperdannier, Roman, et al.
Veröffentlicht: (2024) -
Systematic Evaluation of Online Speaker Diarization Systems Regarding their Latency
von: Aperdannier, Roman, et al.
Veröffentlicht: (2024) -
An approach to optimize inference of the DIART speaker diarization pipeline
von: Aperdannier, Roman, et al.
Veröffentlicht: (2024)