IssueBench: Millions of Realistic Prompts for Measuring Issue Bias in LLM Writing Assistance
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Röttger, Paul, Hinck, Musashi, Hofmann, Valentin, Hackenburg, Kobi, Pyatkin, Valentina, Brahman, Faeze, Hovy, Dirk |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models
par: Röttger, Paul, et autres
Publié: (2024)
par: Röttger, Paul, et autres
Publié: (2024)
Measuring and Mitigating Persona Distortions from AI Writing Assistance
par: Röttger, Paul, et autres
Publié: (2026)
par: Röttger, Paul, et autres
Publié: (2026)
Beyond Flesch-Kincaid: Prompt-based Metrics Improve Difficulty Classification of Educational Texts
par: Rooein, Donya, et autres
Publié: (2024)
par: Rooein, Donya, et autres
Publié: (2024)
Principled Personas: Defining and Measuring the Intended Effects of Persona Prompting on Task Performance
par: de Araujo, Pedro Henrique Luz, et autres
Publié: (2025)
par: de Araujo, Pedro Henrique Luz, et autres
Publié: (2025)
SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety
par: Röttger, Paul, et autres
Publié: (2024)
par: Röttger, Paul, et autres
Publié: (2024)
Compromesso! Italian Many-Shot Jailbreaks Undermine the Safety of Large Language Models
par: Pernisi, Fabio, et autres
Publié: (2024)
par: Pernisi, Fabio, et autres
Publié: (2024)
Evidence of a log scaling law for political persuasion with large language models
par: Hackenburg, Kobi, et autres
Publié: (2024)
par: Hackenburg, Kobi, et autres
Publié: (2024)
SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors
par: Hu, Tiancheng, et autres
Publié: (2025)
par: Hu, Tiancheng, et autres
Publié: (2025)
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
par: Lin, Bill Yuchen, et autres
Publié: (2024)
par: Lin, Bill Yuchen, et autres
Publié: (2024)
The Ecological Fallacy in Annotation: Modelling Human Label Variation goes beyond Sociodemographics
par: Orlikowski, Matthias, et autres
Publié: (2023)
par: Orlikowski, Matthias, et autres
Publié: (2023)
The Pluralistic Moral Gap: Understanding Judgment and Value Differences between Humans and Large Language Models
par: Russo, Giuseppe, et autres
Publié: (2025)
par: Russo, Giuseppe, et autres
Publié: (2025)
Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement
par: Jung, Jaehun, et autres
Publié: (2024)
par: Jung, Jaehun, et autres
Publié: (2024)
What Makes it Ok to Set a Fire? Iterative Self-distillation of Contexts and Rationales for Disambiguating Defeasible Social and Moral Situations
par: Rao, Kavel, et autres
Publié: (2023)
par: Rao, Kavel, et autres
Publié: (2023)
No for Some, Yes for Others: Persona Prompts and Other Sources of False Refusal in Language Models
par: Plaza-del-Arco, Flor Miriam, et autres
Publié: (2025)
par: Plaza-del-Arco, Flor Miriam, et autres
Publié: (2025)
Using Imperfect Surrogates for Downstream Inference: Design-based Supervised Learning for Social Science Applications of Large Language Models
par: Egami, Naoki, et autres
Publié: (2023)
par: Egami, Naoki, et autres
Publié: (2023)
Do Prompts Reshape Representations? An Empirical Study of Prompting Effects on Embeddings
par: Gonzalez-Gutierrez, Cesar, et autres
Publié: (2025)
par: Gonzalez-Gutierrez, Cesar, et autres
Publié: (2025)
Promptly Predicting Structures: The Return of Inference
par: Mehta, Maitrey, et autres
Publié: (2024)
par: Mehta, Maitrey, et autres
Publié: (2024)
Diffusion Language Models Are Natively Length-Aware
par: Rossi, Vittorio, et autres
Publié: (2026)
par: Rossi, Vittorio, et autres
Publié: (2026)
Beyond Demographics: Fine-tuning Large Language Models to Predict Individuals' Subjective Text Perceptions
par: Orlikowski, Matthias, et autres
Publié: (2025)
par: Orlikowski, Matthias, et autres
Publié: (2025)
AutoPersuade: A Framework for Evaluating and Explaining Persuasive Arguments
par: Saenger, Till Raphael, et autres
Publié: (2024)
par: Saenger, Till Raphael, et autres
Publié: (2024)
Hybrid Preferences: Learning to Route Instances for Human vs. AI Feedback
par: Miranda, Lester James V., et autres
Publié: (2024)
par: Miranda, Lester James V., et autres
Publié: (2024)
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
par: Röttger, Paul, et autres
Publié: (2023)
par: Röttger, Paul, et autres
Publié: (2023)
Creativity Support in the Age of Large Language Models: An Empirical Study Involving Emerging Writers
par: Chakrabarty, Tuhin, et autres
Publié: (2023)
par: Chakrabarty, Tuhin, et autres
Publié: (2023)
Steering Large Language Models to Evaluate and Amplify Creativity
par: Olson, Matthew Lyle, et autres
Publié: (2024)
par: Olson, Matthew Lyle, et autres
Publié: (2024)
LLaVA-Gemma: Accelerating Multimodal Foundation Models with a Compact Language Model
par: Hinck, Musashi, et autres
Publié: (2024)
par: Hinck, Musashi, et autres
Publié: (2024)
PlaSma: Making Small Language Models Better Procedural Knowledge Models for (Counterfactual) Planning
par: Brahman, Faeze, et autres
Publié: (2023)
par: Brahman, Faeze, et autres
Publié: (2023)
Artificial intelligence can persuade people to take political actions
par: Hackenburg, Kobi, et autres
Publié: (2026)
par: Hackenburg, Kobi, et autres
Publié: (2026)
Biased Tales: Cultural and Topic Bias in Generating Children's Stories
par: Rooein, Donya, et autres
Publié: (2025)
par: Rooein, Donya, et autres
Publié: (2025)
Dopamine D1Aa and D2a Receptor Expression in the Auditory System of a Vocal Fish
par: Kobi Kobi, et autres
Publié: (2026)
par: Kobi Kobi, et autres
Publié: (2026)
Train for Truth, Keep the Skills: Binary Retrieval-Augmented Reward Mitigates Hallucinations
par: Chen, Tong, et autres
Publié: (2025)
par: Chen, Tong, et autres
Publié: (2025)
AI-LieDar: Examine the Trade-off Between Utility and Truthfulness in LLM Agents
par: Su, Zhe, et autres
Publié: (2024)
par: Su, Zhe, et autres
Publié: (2024)
Debiasing Large Vision-Language Models by Ablating Protected Attribute Representations
par: Ratzlaff, Neale, et autres
Publié: (2024)
par: Ratzlaff, Neale, et autres
Publié: (2024)
Conversations as a Source for Teaching Scientific Concepts at Different Education Levels
par: Rooein, Donya, et autres
Publié: (2024)
par: Rooein, Donya, et autres
Publié: (2024)
Narratives at Conflict: Computational Analysis of News Framing in Multilingual Disinformation Campaigns
par: Sinelnik, Antonina, et autres
Publié: (2024)
par: Sinelnik, Antonina, et autres
Publié: (2024)
"My Answer is C": First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models
par: Wang, Xinpeng, et autres
Publié: (2024)
par: Wang, Xinpeng, et autres
Publié: (2024)
Probing Semantic Routing in Large Mixture-of-Expert Models
par: Olson, Matthew Lyle, et autres
Publié: (2025)
par: Olson, Matthew Lyle, et autres
Publié: (2025)
Reasoning Up the Instruction Ladder for Controllable Language Models
par: Zheng, Zishuo, et autres
Publié: (2025)
par: Zheng, Zishuo, et autres
Publié: (2025)
Large-Scale Data Selection for Instruction Tuning
par: Ivison, Hamish, et autres
Publié: (2025)
par: Ivison, Hamish, et autres
Publié: (2025)
SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning
par: Limozin, Alexis, et autres
Publié: (2026)
par: Limozin, Alexis, et autres
Publié: (2026)
Large Language Model Hacking: Quantifying the Hidden Risks of Using LLMs for Text Annotation
par: Baumann, Joachim, et autres
Publié: (2025)
par: Baumann, Joachim, et autres
Publié: (2025)
Documents similaires
-
Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models
par: Röttger, Paul, et autres
Publié: (2024) -
Measuring and Mitigating Persona Distortions from AI Writing Assistance
par: Röttger, Paul, et autres
Publié: (2026) -
Beyond Flesch-Kincaid: Prompt-based Metrics Improve Difficulty Classification of Educational Texts
par: Rooein, Donya, et autres
Publié: (2024) -
Principled Personas: Defining and Measuring the Intended Effects of Persona Prompting on Task Performance
par: de Araujo, Pedro Henrique Luz, et autres
Publié: (2025) -
SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety
par: Röttger, Paul, et autres
Publié: (2024)