Empirical Evidence for Alignment Faking in a Small LLM and Prompt-Based Mitigation Techniques
Fuente:
arXiv
Saved in:
| Main Author: | Koorndijk, Jeanice |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Societal Alignment Frameworks Can Improve LLM Alignment
by: Stańczak, Karolina, et al.
Published: (2025)
by: Stańczak, Karolina, et al.
Published: (2025)
A Cross-Cultural Assessment of Human Ability to Detect LLM-Generated Fake News about South Africa
by: Schlippe, Tim, et al.
Published: (2025)
by: Schlippe, Tim, et al.
Published: (2025)
SocialNLP Fake-EmoReact 2021 Challenge Overview: Predicting Fake Tweets from Their Replies and GIFs
by: Huang, Chien-Kun, et al.
Published: (2024)
by: Huang, Chien-Kun, et al.
Published: (2024)
Beyond Single-Sentence Prompts: Upgrading Value Alignment Benchmarks with Dialogues and Stories
by: Zhang, Yazhou, et al.
Published: (2025)
by: Zhang, Yazhou, et al.
Published: (2025)
Prompt Engineering Techniques for Mitigating Cultural Bias Against Arabs and Muslims in Large Language Models: A Systematic Review
by: Asseri, Bushra, et al.
Published: (2025)
by: Asseri, Bushra, et al.
Published: (2025)
Dropouts in Confidence: Moral Uncertainty in Human-LLM Alignment
by: Kwon, Jea, et al.
Published: (2025)
by: Kwon, Jea, et al.
Published: (2025)
Exploring Persona-dependent LLM Alignment for the Moral Machine Experiment
by: Kim, Jiseon, et al.
Published: (2025)
by: Kim, Jiseon, et al.
Published: (2025)
Gender and Positional Biases in LLM-Based Hiring Decisions: Evidence from Comparative CV/Résumé Evaluations
by: Rozado, David
Published: (2025)
by: Rozado, David
Published: (2025)
Lightweight Prompt Engineering for Cognitive Alignment in Educational AI: A OneClickQuiz Case Study
by: Yaacoub, Antoun, et al.
Published: (2025)
by: Yaacoub, Antoun, et al.
Published: (2025)
From Descriptive to Prescriptive: Uncover the Social Value Alignment of LLM-based Agents
by: Qu, Jinxian, et al.
Published: (2026)
by: Qu, Jinxian, et al.
Published: (2026)
That's So FETCH: Fashioning Ensemble Techniques for LLM Classification in Civil Legal Intake and Referral
by: Steenhuis, Quinten
Published: (2025)
by: Steenhuis, Quinten
Published: (2025)
Evaluating the Impact of Advanced LLM Techniques on AI-Lecture Tutors for a Robotics Course
by: Kahl, Sebastian, et al.
Published: (2024)
by: Kahl, Sebastian, et al.
Published: (2024)
Fake Artificial Intelligence Generated Contents (FAIGC): A Survey of Theories, Detection Methods, and Opportunities
by: Yu, Xiaomin, et al.
Published: (2024)
by: Yu, Xiaomin, et al.
Published: (2024)
Bad Actor, Good Advisor: Exploring the Role of Large Language Models in Fake News Detection
by: Hu, Beizhe, et al.
Published: (2023)
by: Hu, Beizhe, et al.
Published: (2023)
No Free Lunch in Language Model Bias Mitigation? Targeted Bias Reduction Can Exacerbate Unmitigated LLM Biases
by: Chand, Shireen, et al.
Published: (2025)
by: Chand, Shireen, et al.
Published: (2025)
Evaluating Prompt Engineering Techniques for Accuracy and Confidence Elicitation in Medical LLMs
by: Naderi, Nariman, et al.
Published: (2025)
by: Naderi, Nariman, et al.
Published: (2025)
Mechanical Enforcement for LLM Governance:Evidence of Governance-Task Decoupling in Financial Decision Systems
by: Rodríguez, José Manuel de la Chica, et al.
Published: (2026)
by: Rodríguez, José Manuel de la Chica, et al.
Published: (2026)
Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook
by: Lee, Jaehyeok, et al.
Published: (2026)
by: Lee, Jaehyeok, et al.
Published: (2026)
German General Social Survey Personas: A Survey-Derived Persona Prompt Collection for Population-Aligned LLM Studies
by: Rupprecht, Jens, et al.
Published: (2025)
by: Rupprecht, Jens, et al.
Published: (2025)
Prompt-Based Cost-Effective Evaluation and Operation of ChatGPT as a Computer Programming Teaching Assistant
by: Ballestero-Ribó, Marc, et al.
Published: (2025)
by: Ballestero-Ribó, Marc, et al.
Published: (2025)
Scopes of Alignment
by: Varshney, Kush R., et al.
Published: (2025)
by: Varshney, Kush R., et al.
Published: (2025)
Prompt and Prejudice
by: Berlincioni, Lorenzo, et al.
Published: (2024)
by: Berlincioni, Lorenzo, et al.
Published: (2024)
Social Catalysts, Not Moral Agents: The Illusion of Alignment in LLM Societies
by: Hu, Yueqing, et al.
Published: (2026)
by: Hu, Yueqing, et al.
Published: (2026)
Chat Bankman-Fried: an Exploration of LLM Alignment in Finance
by: Biancotti, Claudia, et al.
Published: (2024)
by: Biancotti, Claudia, et al.
Published: (2024)
CR4T: Rewrite-Based Guardrails for Adolescent LLM Safety
by: An, Heajun, et al.
Published: (2026)
by: An, Heajun, et al.
Published: (2026)
Toxic HallucinAItions: Perturbing Prompts and Tracing LLM Circuits
by: Shimgekar, Soorya Ram, et al.
Published: (2026)
by: Shimgekar, Soorya Ram, et al.
Published: (2026)
Explore the Potential of LLMs in Misinformation Detection: An Empirical Study
by: Chen, Mengyang, et al.
Published: (2023)
by: Chen, Mengyang, et al.
Published: (2023)
From Feature-Based Models to Generative AI: Validity Evidence for Constructed Response Scoring
by: Casabianca, Jodi M., et al.
Published: (2026)
by: Casabianca, Jodi M., et al.
Published: (2026)
Beyond Translation: LLM-Based Data Generation for Multilingual Fact-Checking
by: Chung, Yi-Ling, et al.
Published: (2025)
by: Chung, Yi-Ling, et al.
Published: (2025)
Place Matters: Comparing LLM Hallucination Rates for Place-Based Legal Queries
by: Curran, Damian, et al.
Published: (2025)
by: Curran, Damian, et al.
Published: (2025)
From Prompts to Constructs: A Dual-Validity Framework for LLM Research in Psychology
by: Lin, Zhicheng
Published: (2025)
by: Lin, Zhicheng
Published: (2025)
Mitigating Gender Bias via Fostering Exploratory Thinking in LLMs
by: Wei, Kangda, et al.
Published: (2025)
by: Wei, Kangda, et al.
Published: (2025)
Position is Power: System Prompts as a Mechanism of Bias in Large Language Models (LLMs)
by: Neumann, Anna, et al.
Published: (2025)
by: Neumann, Anna, et al.
Published: (2025)
Fluent but Foreign: Even Regional LLMs Lack Cultural Alignment
by: Agarwal, Dhruv, et al.
Published: (2025)
by: Agarwal, Dhruv, et al.
Published: (2025)
Detection and Mitigation of Hallucination in Large Reasoning Models: A Mechanistic Perspective
by: Sun, Zhongxiang, et al.
Published: (2025)
by: Sun, Zhongxiang, et al.
Published: (2025)
A Cross-Domain Study of the Use of Persuasion Techniques in Online Disinformation
by: Leite, João A., et al.
Published: (2024)
by: Leite, João A., et al.
Published: (2024)
How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States
by: Zhou, Zhenhong, et al.
Published: (2024)
by: Zhou, Zhenhong, et al.
Published: (2024)
CLST: Cold-Start Mitigation in Knowledge Tracing by Aligning a Generative Language Model as a Students' Knowledge Tracer
by: Jung, Heeseok, et al.
Published: (2024)
by: Jung, Heeseok, et al.
Published: (2024)
Alignment Drift in Long-Term Human-LLM Interaction: A Mechanism-Oriented Framework
by: Yao, Xintong
Published: (2026)
by: Yao, Xintong
Published: (2026)
Measuring What Matters -- or What's Convenient?: Robustness of LLM-Based Scoring Systems to Construct-Irrelevant Factors
by: Walsh, Cole, et al.
Published: (2026)
by: Walsh, Cole, et al.
Published: (2026)
Similar Items
-
Societal Alignment Frameworks Can Improve LLM Alignment
by: Stańczak, Karolina, et al.
Published: (2025) -
A Cross-Cultural Assessment of Human Ability to Detect LLM-Generated Fake News about South Africa
by: Schlippe, Tim, et al.
Published: (2025) -
SocialNLP Fake-EmoReact 2021 Challenge Overview: Predicting Fake Tweets from Their Replies and GIFs
by: Huang, Chien-Kun, et al.
Published: (2024) -
Beyond Single-Sentence Prompts: Upgrading Value Alignment Benchmarks with Dialogues and Stories
by: Zhang, Yazhou, et al.
Published: (2025) -
Prompt Engineering Techniques for Mitigating Cultural Bias Against Arabs and Muslims in Large Language Models: A Systematic Review
by: Asseri, Bushra, et al.
Published: (2025)