Empirical Evidence for Alignment Faking in a Small LLM and Prompt-Based Mitigation Techniques
Fuente:
arXiv
Salvato in:
| Autore principale: | Koorndijk, Jeanice |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Societal Alignment Frameworks Can Improve LLM Alignment
di: Stańczak, Karolina, et al.
Pubblicazione: (2025)
di: Stańczak, Karolina, et al.
Pubblicazione: (2025)
A Cross-Cultural Assessment of Human Ability to Detect LLM-Generated Fake News about South Africa
di: Schlippe, Tim, et al.
Pubblicazione: (2025)
di: Schlippe, Tim, et al.
Pubblicazione: (2025)
SocialNLP Fake-EmoReact 2021 Challenge Overview: Predicting Fake Tweets from Their Replies and GIFs
di: Huang, Chien-Kun, et al.
Pubblicazione: (2024)
di: Huang, Chien-Kun, et al.
Pubblicazione: (2024)
Beyond Single-Sentence Prompts: Upgrading Value Alignment Benchmarks with Dialogues and Stories
di: Zhang, Yazhou, et al.
Pubblicazione: (2025)
di: Zhang, Yazhou, et al.
Pubblicazione: (2025)
Prompt Engineering Techniques for Mitigating Cultural Bias Against Arabs and Muslims in Large Language Models: A Systematic Review
di: Asseri, Bushra, et al.
Pubblicazione: (2025)
di: Asseri, Bushra, et al.
Pubblicazione: (2025)
Dropouts in Confidence: Moral Uncertainty in Human-LLM Alignment
di: Kwon, Jea, et al.
Pubblicazione: (2025)
di: Kwon, Jea, et al.
Pubblicazione: (2025)
Exploring Persona-dependent LLM Alignment for the Moral Machine Experiment
di: Kim, Jiseon, et al.
Pubblicazione: (2025)
di: Kim, Jiseon, et al.
Pubblicazione: (2025)
Gender and Positional Biases in LLM-Based Hiring Decisions: Evidence from Comparative CV/Résumé Evaluations
di: Rozado, David
Pubblicazione: (2025)
di: Rozado, David
Pubblicazione: (2025)
Lightweight Prompt Engineering for Cognitive Alignment in Educational AI: A OneClickQuiz Case Study
di: Yaacoub, Antoun, et al.
Pubblicazione: (2025)
di: Yaacoub, Antoun, et al.
Pubblicazione: (2025)
From Descriptive to Prescriptive: Uncover the Social Value Alignment of LLM-based Agents
di: Qu, Jinxian, et al.
Pubblicazione: (2026)
di: Qu, Jinxian, et al.
Pubblicazione: (2026)
That's So FETCH: Fashioning Ensemble Techniques for LLM Classification in Civil Legal Intake and Referral
di: Steenhuis, Quinten
Pubblicazione: (2025)
di: Steenhuis, Quinten
Pubblicazione: (2025)
Evaluating the Impact of Advanced LLM Techniques on AI-Lecture Tutors for a Robotics Course
di: Kahl, Sebastian, et al.
Pubblicazione: (2024)
di: Kahl, Sebastian, et al.
Pubblicazione: (2024)
Fake Artificial Intelligence Generated Contents (FAIGC): A Survey of Theories, Detection Methods, and Opportunities
di: Yu, Xiaomin, et al.
Pubblicazione: (2024)
di: Yu, Xiaomin, et al.
Pubblicazione: (2024)
Bad Actor, Good Advisor: Exploring the Role of Large Language Models in Fake News Detection
di: Hu, Beizhe, et al.
Pubblicazione: (2023)
di: Hu, Beizhe, et al.
Pubblicazione: (2023)
No Free Lunch in Language Model Bias Mitigation? Targeted Bias Reduction Can Exacerbate Unmitigated LLM Biases
di: Chand, Shireen, et al.
Pubblicazione: (2025)
di: Chand, Shireen, et al.
Pubblicazione: (2025)
Evaluating Prompt Engineering Techniques for Accuracy and Confidence Elicitation in Medical LLMs
di: Naderi, Nariman, et al.
Pubblicazione: (2025)
di: Naderi, Nariman, et al.
Pubblicazione: (2025)
Mechanical Enforcement for LLM Governance:Evidence of Governance-Task Decoupling in Financial Decision Systems
di: Rodríguez, José Manuel de la Chica, et al.
Pubblicazione: (2026)
di: Rodríguez, José Manuel de la Chica, et al.
Pubblicazione: (2026)
Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook
di: Lee, Jaehyeok, et al.
Pubblicazione: (2026)
di: Lee, Jaehyeok, et al.
Pubblicazione: (2026)
German General Social Survey Personas: A Survey-Derived Persona Prompt Collection for Population-Aligned LLM Studies
di: Rupprecht, Jens, et al.
Pubblicazione: (2025)
di: Rupprecht, Jens, et al.
Pubblicazione: (2025)
Prompt-Based Cost-Effective Evaluation and Operation of ChatGPT as a Computer Programming Teaching Assistant
di: Ballestero-Ribó, Marc, et al.
Pubblicazione: (2025)
di: Ballestero-Ribó, Marc, et al.
Pubblicazione: (2025)
Scopes of Alignment
di: Varshney, Kush R., et al.
Pubblicazione: (2025)
di: Varshney, Kush R., et al.
Pubblicazione: (2025)
Prompt and Prejudice
di: Berlincioni, Lorenzo, et al.
Pubblicazione: (2024)
di: Berlincioni, Lorenzo, et al.
Pubblicazione: (2024)
Social Catalysts, Not Moral Agents: The Illusion of Alignment in LLM Societies
di: Hu, Yueqing, et al.
Pubblicazione: (2026)
di: Hu, Yueqing, et al.
Pubblicazione: (2026)
Chat Bankman-Fried: an Exploration of LLM Alignment in Finance
di: Biancotti, Claudia, et al.
Pubblicazione: (2024)
di: Biancotti, Claudia, et al.
Pubblicazione: (2024)
CR4T: Rewrite-Based Guardrails for Adolescent LLM Safety
di: An, Heajun, et al.
Pubblicazione: (2026)
di: An, Heajun, et al.
Pubblicazione: (2026)
Toxic HallucinAItions: Perturbing Prompts and Tracing LLM Circuits
di: Shimgekar, Soorya Ram, et al.
Pubblicazione: (2026)
di: Shimgekar, Soorya Ram, et al.
Pubblicazione: (2026)
Explore the Potential of LLMs in Misinformation Detection: An Empirical Study
di: Chen, Mengyang, et al.
Pubblicazione: (2023)
di: Chen, Mengyang, et al.
Pubblicazione: (2023)
From Feature-Based Models to Generative AI: Validity Evidence for Constructed Response Scoring
di: Casabianca, Jodi M., et al.
Pubblicazione: (2026)
di: Casabianca, Jodi M., et al.
Pubblicazione: (2026)
Beyond Translation: LLM-Based Data Generation for Multilingual Fact-Checking
di: Chung, Yi-Ling, et al.
Pubblicazione: (2025)
di: Chung, Yi-Ling, et al.
Pubblicazione: (2025)
Place Matters: Comparing LLM Hallucination Rates for Place-Based Legal Queries
di: Curran, Damian, et al.
Pubblicazione: (2025)
di: Curran, Damian, et al.
Pubblicazione: (2025)
From Prompts to Constructs: A Dual-Validity Framework for LLM Research in Psychology
di: Lin, Zhicheng
Pubblicazione: (2025)
di: Lin, Zhicheng
Pubblicazione: (2025)
Mitigating Gender Bias via Fostering Exploratory Thinking in LLMs
di: Wei, Kangda, et al.
Pubblicazione: (2025)
di: Wei, Kangda, et al.
Pubblicazione: (2025)
Position is Power: System Prompts as a Mechanism of Bias in Large Language Models (LLMs)
di: Neumann, Anna, et al.
Pubblicazione: (2025)
di: Neumann, Anna, et al.
Pubblicazione: (2025)
Fluent but Foreign: Even Regional LLMs Lack Cultural Alignment
di: Agarwal, Dhruv, et al.
Pubblicazione: (2025)
di: Agarwal, Dhruv, et al.
Pubblicazione: (2025)
Detection and Mitigation of Hallucination in Large Reasoning Models: A Mechanistic Perspective
di: Sun, Zhongxiang, et al.
Pubblicazione: (2025)
di: Sun, Zhongxiang, et al.
Pubblicazione: (2025)
A Cross-Domain Study of the Use of Persuasion Techniques in Online Disinformation
di: Leite, João A., et al.
Pubblicazione: (2024)
di: Leite, João A., et al.
Pubblicazione: (2024)
How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States
di: Zhou, Zhenhong, et al.
Pubblicazione: (2024)
di: Zhou, Zhenhong, et al.
Pubblicazione: (2024)
CLST: Cold-Start Mitigation in Knowledge Tracing by Aligning a Generative Language Model as a Students' Knowledge Tracer
di: Jung, Heeseok, et al.
Pubblicazione: (2024)
di: Jung, Heeseok, et al.
Pubblicazione: (2024)
Alignment Drift in Long-Term Human-LLM Interaction: A Mechanism-Oriented Framework
di: Yao, Xintong
Pubblicazione: (2026)
di: Yao, Xintong
Pubblicazione: (2026)
Measuring What Matters -- or What's Convenient?: Robustness of LLM-Based Scoring Systems to Construct-Irrelevant Factors
di: Walsh, Cole, et al.
Pubblicazione: (2026)
di: Walsh, Cole, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Societal Alignment Frameworks Can Improve LLM Alignment
di: Stańczak, Karolina, et al.
Pubblicazione: (2025) -
A Cross-Cultural Assessment of Human Ability to Detect LLM-Generated Fake News about South Africa
di: Schlippe, Tim, et al.
Pubblicazione: (2025) -
SocialNLP Fake-EmoReact 2021 Challenge Overview: Predicting Fake Tweets from Their Replies and GIFs
di: Huang, Chien-Kun, et al.
Pubblicazione: (2024) -
Beyond Single-Sentence Prompts: Upgrading Value Alignment Benchmarks with Dialogues and Stories
di: Zhang, Yazhou, et al.
Pubblicazione: (2025) -
Prompt Engineering Techniques for Mitigating Cultural Bias Against Arabs and Muslims in Large Language Models: A Systematic Review
di: Asseri, Bushra, et al.
Pubblicazione: (2025)