Large Language Model Hacking: Quantifying the Hidden Risks of Using LLMs for Text Annotation
Fuente:
arXiv
Saved in:
| Main Authors: | Baumann, Joachim, Röttger, Paul, Urman, Aleksandra, Wendsjö, Albert, Plaza-del-Arco, Flor Miriam, Gruber, Johannes B., Hovy, Dirk |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Wisdom of Instruction-Tuned Language Model Crowds. Exploring Model Label Variation
by: Plaza-del-Arco, Flor Miriam, et al.
Published: (2023)
by: Plaza-del-Arco, Flor Miriam, et al.
Published: (2023)
No for Some, Yes for Others: Persona Prompts and Other Sources of False Refusal in Language Models
by: Plaza-del-Arco, Flor Miriam, et al.
Published: (2025)
by: Plaza-del-Arco, Flor Miriam, et al.
Published: (2025)
Exploring Subjective Tasks in Farsi: A Survey Analysis and Evaluation of Language Models
by: Rooein, Donya, et al.
Published: (2025)
by: Rooein, Donya, et al.
Published: (2025)
Emotion Analysis in NLP: Trends, Gaps and Roadmap for Future Directions
by: Plaza-del-Arco, Flor Miriam, et al.
Published: (2024)
by: Plaza-del-Arco, Flor Miriam, et al.
Published: (2024)
Divine LLaMAs: Bias, Stereotypes, Stigmatization, and Emotion Representation of Religion in Large Language Models
by: Plaza-del-Arco, Flor Miriam, et al.
Published: (2024)
by: Plaza-del-Arco, Flor Miriam, et al.
Published: (2024)
Angry Men, Sad Women: Large Language Models Reflect Gendered Stereotypes in Emotion Attribution
by: Plaza-del-Arco, Flor Miriam, et al.
Published: (2024)
by: Plaza-del-Arco, Flor Miriam, et al.
Published: (2024)
COPOS: Corpus Of Patient Opinions in Spanish. Application of Sentiment Analysis Techniques
by: Flor Miriam Plaza-del-Arco
Published: (2016)
by: Flor Miriam Plaza-del-Arco
Published: (2016)
SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors
by: Hu, Tiancheng, et al.
Published: (2025)
by: Hu, Tiancheng, et al.
Published: (2025)
Compromesso! Italian Many-Shot Jailbreaks Undermine the Safety of Large Language Models
by: Pernisi, Fabio, et al.
Published: (2024)
by: Pernisi, Fabio, et al.
Published: (2024)
GoogleTrendArchive: A Year-Long Archive of Real-Time Web Search Trends Worldwide
by: Urman, Aleksandra, et al.
Published: (2026)
by: Urman, Aleksandra, et al.
Published: (2026)
The Ecological Fallacy in Annotation: Modelling Human Label Variation goes beyond Sociodemographics
by: Orlikowski, Matthias, et al.
Published: (2023)
by: Orlikowski, Matthias, et al.
Published: (2023)
Beyond Flesch-Kincaid: Prompt-based Metrics Improve Difficulty Classification of Educational Texts
by: Rooein, Donya, et al.
Published: (2024)
by: Rooein, Donya, et al.
Published: (2024)
The Pluralistic Moral Gap: Understanding Judgment and Value Differences between Humans and Large Language Models
by: Russo, Giuseppe, et al.
Published: (2025)
by: Russo, Giuseppe, et al.
Published: (2025)
SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety
by: Röttger, Paul, et al.
Published: (2024)
by: Röttger, Paul, et al.
Published: (2024)
Beyond Demographics: Fine-tuning Large Language Models to Predict Individuals' Subjective Text Perceptions
by: Orlikowski, Matthias, et al.
Published: (2025)
by: Orlikowski, Matthias, et al.
Published: (2025)
SINAI at eRisk@CLEF 2023: Approaching Early Detection of Gambling with Natural Language Processing
by: Marmol-Romero, Alba Maria, et al.
Published: (2025)
by: Marmol-Romero, Alba Maria, et al.
Published: (2025)
Stop Automating Peer Review Without Rigorous Evaluation
by: Baumann, Joachim, et al.
Published: (2026)
by: Baumann, Joachim, et al.
Published: (2026)
FLANS at SemEval-2026 Task 7: RAG with Open-Sourced Smaller LLMs for Everyday Knowledge Across Diverse Languages and Cultures
by: Bogdanova, Liliia, et al.
Published: (2026)
by: Bogdanova, Liliia, et al.
Published: (2026)
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
by: Röttger, Paul, et al.
Published: (2023)
by: Röttger, Paul, et al.
Published: (2023)
Principled Personas: Defining and Measuring the Intended Effects of Persona Prompting on Task Performance
by: de Araujo, Pedro Henrique Luz, et al.
Published: (2025)
by: de Araujo, Pedro Henrique Luz, et al.
Published: (2025)
Language Model Council: Democratically Benchmarking Foundation Models on Highly Subjective Tasks
by: Zhao, Justin, et al.
Published: (2024)
by: Zhao, Justin, et al.
Published: (2024)
Diffusion Language Models Are Natively Length-Aware
by: Rossi, Vittorio, et al.
Published: (2026)
by: Rossi, Vittorio, et al.
Published: (2026)
From Chatbots to Confidants: A Cross-Cultural Study of LLM Adoption for Emotional Support
by: Amat-Lefort, Natalia, et al.
Published: (2026)
by: Amat-Lefort, Natalia, et al.
Published: (2026)
Think Like a Person Before Responding: A Multi-Faceted Evaluation of Persona-Guided LLMs for Countering Hate
by: Ngueajio, Mikel K., et al.
Published: (2025)
by: Ngueajio, Mikel K., et al.
Published: (2025)
Reduced AI Acceptance After the Generative AI Boom: Evidence From a Two-Wave Survey Study
by: Baumann, Joachim, et al.
Published: (2025)
by: Baumann, Joachim, et al.
Published: (2025)
Auditing Google's AI Overviews and Featured Snippets: A Case Study on Baby Care and Pregnancy
by: Hu, Desheng, et al.
Published: (2025)
by: Hu, Desheng, et al.
Published: (2025)
Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models
by: Röttger, Paul, et al.
Published: (2024)
by: Röttger, Paul, et al.
Published: (2024)
"My Answer is C": First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models
by: Wang, Xinpeng, et al.
Published: (2024)
by: Wang, Xinpeng, et al.
Published: (2024)
WEIRD Audits? Research Trends, Linguistic and Geographical Disparities in the Algorithm Audits of Online Platforms -- A Systematic Literature Review
by: Urman, Aleksandra, et al.
Published: (2024)
by: Urman, Aleksandra, et al.
Published: (2024)
The right to audit and power asymmetries in algorithm auditing
by: Urman, Aleksandra, et al.
Published: (2023)
by: Urman, Aleksandra, et al.
Published: (2023)
IssueBench: Millions of Realistic Prompts for Measuring Issue Bias in LLM Writing Assistance
by: Röttger, Paul, et al.
Published: (2025)
by: Röttger, Paul, et al.
Published: (2025)
Towards Human-Level Text Coding with LLMs: The Case of Fatherhood Roles in Public Policy Documents
by: Lupo, Lorenzo, et al.
Published: (2023)
by: Lupo, Lorenzo, et al.
Published: (2023)
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
by: Taylor, Mia, et al.
Published: (2025)
by: Taylor, Mia, et al.
Published: (2025)
Improving Adversarial Data Collection by Supporting Annotators: Lessons from GAHD, a German Hate Speech Dataset
by: Goldzycher, Janis, et al.
Published: (2024)
by: Goldzycher, Janis, et al.
Published: (2024)
MFTCXplain: A Multilingual Benchmark Dataset for Evaluating the Moral Reasoning of LLMs through Multi-hop Hate Speech Explanation
by: Trager, Jackson, et al.
Published: (2025)
by: Trager, Jackson, et al.
Published: (2025)
Conversations as a Source for Teaching Scientific Concepts at Different Education Levels
by: Rooein, Donya, et al.
Published: (2024)
by: Rooein, Donya, et al.
Published: (2024)
Narratives at Conflict: Computational Analysis of News Framing in Multilingual Disinformation Campaigns
by: Sinelnik, Antonina, et al.
Published: (2024)
by: Sinelnik, Antonina, et al.
Published: (2024)
The Hidden Cost of Straight Lines: Quantifying Misallocation Risk in Voronoi-based Service Area Models
by: Pinero, JA Torrecilla, et al.
Published: (2025)
by: Pinero, JA Torrecilla, et al.
Published: (2025)
SINAI at eRisk@CLEF 2022: Approaching Early Detection of Gambling and Eating Disorders with Natural Language Processing
by: Marmol-Romero, Alba Maria, et al.
Published: (2025)
by: Marmol-Romero, Alba Maria, et al.
Published: (2025)
Googling the Big Lie: Search Engines, News Media, and the US 2020 Election Conspiracy
by: de León, Ernesto, et al.
Published: (2024)
by: de León, Ernesto, et al.
Published: (2024)
Similar Items
-
Wisdom of Instruction-Tuned Language Model Crowds. Exploring Model Label Variation
by: Plaza-del-Arco, Flor Miriam, et al.
Published: (2023) -
No for Some, Yes for Others: Persona Prompts and Other Sources of False Refusal in Language Models
by: Plaza-del-Arco, Flor Miriam, et al.
Published: (2025) -
Exploring Subjective Tasks in Farsi: A Survey Analysis and Evaluation of Language Models
by: Rooein, Donya, et al.
Published: (2025) -
Emotion Analysis in NLP: Trends, Gaps and Roadmap for Future Directions
by: Plaza-del-Arco, Flor Miriam, et al.
Published: (2024) -
Divine LLaMAs: Bias, Stereotypes, Stigmatization, and Emotion Representation of Religion in Large Language Models
by: Plaza-del-Arco, Flor Miriam, et al.
Published: (2024)