AI Sandbagging: Language Models can Strategically Underperform on Evaluations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | van der Weij, Teun, Hofstätter, Felix, Jaffe, Ollie, Brown, Samuel F., Ward, Francis Rhys |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Extending Activation Steering to Broad Skills and Multiple Behaviours
von: van der Weij, Teun, et al.
Veröffentlicht: (2024)
von: van der Weij, Teun, et al.
Veröffentlicht: (2024)
The Elicitation Game: Evaluating Capability Elicitation Techniques
von: Hofstätter, Felix, et al.
Veröffentlicht: (2025)
von: Hofstätter, Felix, et al.
Veröffentlicht: (2025)
The Aftermath of DrawEduMath: Vision Language Models Underperform with Struggling Students and Misdiagnose Errors
von: Lucy, Li, et al.
Veröffentlicht: (2026)
von: Lucy, Li, et al.
Veröffentlicht: (2026)
Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models
von: Tice, Cameron, et al.
Veröffentlicht: (2024)
von: Tice, Cameron, et al.
Veröffentlicht: (2024)
Modeling Motivated Reasoning in Law: Evaluating Strategic Role Conditioning in LLM Summarization
von: Cho, Eunjung, et al.
Veröffentlicht: (2025)
von: Cho, Eunjung, et al.
Veröffentlicht: (2025)
DaKultur: Evaluating the Cultural Awareness of Language Models for Danish with Native Speakers
von: Müller-Eberstein, Max, et al.
Veröffentlicht: (2025)
von: Müller-Eberstein, Max, et al.
Veröffentlicht: (2025)
Speciesism in AI: Evaluating Discrimination Against Animals in Large Language Models
von: Jotautaitė, Monika, et al.
Veröffentlicht: (2025)
von: Jotautaitė, Monika, et al.
Veröffentlicht: (2025)
Strategic Insights in Human and Large Language Model Tactics at Word Guessing Games
von: Rikters, Matīss, et al.
Veröffentlicht: (2024)
von: Rikters, Matīss, et al.
Veröffentlicht: (2024)
Generative Language Models Exhibit Social Identity Biases
von: Hu, Tiancheng, et al.
Veröffentlicht: (2023)
von: Hu, Tiancheng, et al.
Veröffentlicht: (2023)
Evaluating Language Model Character Traits
von: Ward, Francis Rhys, et al.
Veröffentlicht: (2024)
von: Ward, Francis Rhys, et al.
Veröffentlicht: (2024)
Probing and Steering Evaluation Awareness of Language Models
von: Nguyen, Jord, et al.
Veröffentlicht: (2025)
von: Nguyen, Jord, et al.
Veröffentlicht: (2025)
MalAlgoQA: Pedagogical Evaluation of Counterfactual Reasoning in Large Language Models and Implications for AI in Education
von: Liu, Naiming, et al.
Veröffentlicht: (2024)
von: Liu, Naiming, et al.
Veröffentlicht: (2024)
AI Safety in Generative AI Large Language Models: A Survey
von: Chua, Jaymari, et al.
Veröffentlicht: (2024)
von: Chua, Jaymari, et al.
Veröffentlicht: (2024)
Evaluating Digital Inclusiveness of Digital Agri-Food Tools Using Large Language Models: A Comparative Analysis Between Human and AI-Based Evaluations
von: Pewinya, Githma, et al.
Veröffentlicht: (2026)
von: Pewinya, Githma, et al.
Veröffentlicht: (2026)
Generalization in Healthcare AI: Evaluation of a Clinical Large Language Model
von: Rahman, Salman, et al.
Veröffentlicht: (2024)
von: Rahman, Salman, et al.
Veröffentlicht: (2024)
Voice Under Revision: Large Language Models and the Normalization of Personal Narrative
von: van Nuenen, Tom
Veröffentlicht: (2026)
von: van Nuenen, Tom
Veröffentlicht: (2026)
The World of Generative AI: Deepfakes and Large Language Models
von: Mitra, Alakananda, et al.
Veröffentlicht: (2024)
von: Mitra, Alakananda, et al.
Veröffentlicht: (2024)
"Mirror" Language AI Models of Depression are Criterion-Contaminated
von: Li, Tong, et al.
Veröffentlicht: (2025)
von: Li, Tong, et al.
Veröffentlicht: (2025)
Evaluating Proactive Risk Awareness of Large Language Models
von: Luo, Xuan, et al.
Veröffentlicht: (2026)
von: Luo, Xuan, et al.
Veröffentlicht: (2026)
Extrinsic Evaluation of Cultural Competence in Large Language Models
von: Bhatt, Shaily, et al.
Veröffentlicht: (2024)
von: Bhatt, Shaily, et al.
Veröffentlicht: (2024)
Dual Use Concerns of Generative AI and Large Language Models
von: Grinbaum, Alexei, et al.
Veröffentlicht: (2023)
von: Grinbaum, Alexei, et al.
Veröffentlicht: (2023)
How malicious AI swarms can threaten democracy: The fusion of agentic AI and LLMs marks a new frontier in information warfare
von: Schroeder, Daniel Thilo, et al.
Veröffentlicht: (2025)
von: Schroeder, Daniel Thilo, et al.
Veröffentlicht: (2025)
AI Knows When It's Being Watched: Functional Strategic Action and Contextual Register Modulation in Large Language Models
von: Covas, Vinicius, et al.
Veröffentlicht: (2026)
von: Covas, Vinicius, et al.
Veröffentlicht: (2026)
Evaluation of AI Ethics Tools in Language Models: A Developers' Perspective Case Study
von: Silva, Jhessica, et al.
Veröffentlicht: (2025)
von: Silva, Jhessica, et al.
Veröffentlicht: (2025)
Which Type of Students can LLMs Act? Investigating Authentic Simulation with Graph-based Human-AI Collaborative System
von: Li, Haoxuan, et al.
Veröffentlicht: (2025)
von: Li, Haoxuan, et al.
Veröffentlicht: (2025)
People Make Better Edits: Measuring the Efficacy of LLM-Generated Counterfactually Augmented Data for Harmful Language Detection
von: Sen, Indira, et al.
Veröffentlicht: (2023)
von: Sen, Indira, et al.
Veröffentlicht: (2023)
Evaluation Awareness in Language Models Has Limited Effect on Behaviour
von: Knecht, Amelie, et al.
Veröffentlicht: (2026)
von: Knecht, Amelie, et al.
Veröffentlicht: (2026)
Towards Equitable AI: Detecting Bias in Using Large Language Models for Marketing
von: Yilmaz, Berk, et al.
Veröffentlicht: (2025)
von: Yilmaz, Berk, et al.
Veröffentlicht: (2025)
AGGA: A Dataset of Academic Guidelines for Generative AI and Large Language Models
von: Jiao, Junfeng, et al.
Veröffentlicht: (2025)
von: Jiao, Junfeng, et al.
Veröffentlicht: (2025)
The Language of Trauma: Modeling Traumatic Event Descriptions Across Domains with Explainable AI
von: Schirmer, Miriam, et al.
Veröffentlicht: (2024)
von: Schirmer, Miriam, et al.
Veröffentlicht: (2024)
Evaluating open-source Large Language Models for automated fact-checking
von: Fontana, Nicolo', et al.
Veröffentlicht: (2025)
von: Fontana, Nicolo', et al.
Veröffentlicht: (2025)
Human Preferences for Constructive Interactions in Language Model Alignment
von: Kyrychenko, Yara, et al.
Veröffentlicht: (2025)
von: Kyrychenko, Yara, et al.
Veröffentlicht: (2025)
Large Language Models and Forensic Linguistics: Navigating Opportunities and Threats in the Age of Generative AI
von: Mikros, George
Veröffentlicht: (2025)
von: Mikros, George
Veröffentlicht: (2025)
Cancer-Myth: Evaluating Large Language Models on Patient Questions with False Presuppositions
von: Zhu, Wang Bill, et al.
Veröffentlicht: (2025)
von: Zhu, Wang Bill, et al.
Veröffentlicht: (2025)
Not All Jokes Land: Evaluating Large Language Models Understanding of Workplace Humor
von: Shafiei, Mohammadamin, et al.
Veröffentlicht: (2025)
von: Shafiei, Mohammadamin, et al.
Veröffentlicht: (2025)
Leveraging Large Language Models for Actionable Course Evaluation Student Feedback to Lecturers
von: Zhang, Mike, et al.
Veröffentlicht: (2024)
von: Zhang, Mike, et al.
Veröffentlicht: (2024)
Evaluating Large Language Models on Spatial Tasks: A Multi-Task Benchmarking Study
von: Xu, Liuchang, et al.
Veröffentlicht: (2024)
von: Xu, Liuchang, et al.
Veröffentlicht: (2024)
Adesua: Development and Feasibility Study of an AI WhatsApp Bot for Science Learning in West Africa
von: Boateng, George, et al.
Veröffentlicht: (2026)
von: Boateng, George, et al.
Veröffentlicht: (2026)
SALAD: Smart AI Language Assistant Daily
von: Nihal, Ragib Amin, et al.
Veröffentlicht: (2024)
von: Nihal, Ragib Amin, et al.
Veröffentlicht: (2024)
AI-VERDE: A Gateway for Egalitarian Access to Large Language Model-Based Resources For Educational Institutions
von: Mithun, Paul, et al.
Veröffentlicht: (2025)
von: Mithun, Paul, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Extending Activation Steering to Broad Skills and Multiple Behaviours
von: van der Weij, Teun, et al.
Veröffentlicht: (2024) -
The Elicitation Game: Evaluating Capability Elicitation Techniques
von: Hofstätter, Felix, et al.
Veröffentlicht: (2025) -
The Aftermath of DrawEduMath: Vision Language Models Underperform with Struggling Students and Misdiagnose Errors
von: Lucy, Li, et al.
Veröffentlicht: (2026) -
Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models
von: Tice, Cameron, et al.
Veröffentlicht: (2024) -
Modeling Motivated Reasoning in Law: Evaluating Strategic Role Conditioning in LLM Summarization
von: Cho, Eunjung, et al.
Veröffentlicht: (2025)