Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Röttger, Paul, Hofmann, Valentin, Pyatkin, Valentina, Hinck, Musashi, Kirk, Hannah Rose, Schütze, Hinrich, Hovy, Dirk |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
IssueBench: Millions of Realistic Prompts for Measuring Issue Bias in LLM Writing Assistance
par: Röttger, Paul, et autres
Publié: (2025)
par: Röttger, Paul, et autres
Publié: (2025)
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
par: Röttger, Paul, et autres
Publié: (2023)
par: Röttger, Paul, et autres
Publié: (2023)
The Pluralistic Moral Gap: Understanding Judgment and Value Differences between Humans and Large Language Models
par: Russo, Giuseppe, et autres
Publié: (2025)
par: Russo, Giuseppe, et autres
Publié: (2025)
Compromesso! Italian Many-Shot Jailbreaks Undermine the Safety of Large Language Models
par: Pernisi, Fabio, et autres
Publié: (2024)
par: Pernisi, Fabio, et autres
Publié: (2024)
SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety
par: Röttger, Paul, et autres
Publié: (2024)
par: Röttger, Paul, et autres
Publié: (2024)
The Ecological Fallacy in Annotation: Modelling Human Label Variation goes beyond Sociodemographics
par: Orlikowski, Matthias, et autres
Publié: (2023)
par: Orlikowski, Matthias, et autres
Publié: (2023)
Beyond Flesch-Kincaid: Prompt-based Metrics Improve Difficulty Classification of Educational Texts
par: Rooein, Donya, et autres
Publié: (2024)
par: Rooein, Donya, et autres
Publié: (2024)
Measuring and Mitigating Persona Distortions from AI Writing Assistance
par: Röttger, Paul, et autres
Publié: (2026)
par: Röttger, Paul, et autres
Publié: (2026)
AutoPersuade: A Framework for Evaluating and Explaining Persuasive Arguments
par: Saenger, Till Raphael, et autres
Publié: (2024)
par: Saenger, Till Raphael, et autres
Publié: (2024)
Using Imperfect Surrogates for Downstream Inference: Design-based Supervised Learning for Social Science Applications of Large Language Models
par: Egami, Naoki, et autres
Publié: (2023)
par: Egami, Naoki, et autres
Publié: (2023)
Principled Personas: Defining and Measuring the Intended Effects of Persona Prompting on Task Performance
par: de Araujo, Pedro Henrique Luz, et autres
Publié: (2025)
par: de Araujo, Pedro Henrique Luz, et autres
Publié: (2025)
Steering Large Language Models to Evaluate and Amplify Creativity
par: Olson, Matthew Lyle, et autres
Publié: (2024)
par: Olson, Matthew Lyle, et autres
Publié: (2024)
Diffusion Language Models Are Natively Length-Aware
par: Rossi, Vittorio, et autres
Publié: (2026)
par: Rossi, Vittorio, et autres
Publié: (2026)
SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors
par: Hu, Tiancheng, et autres
Publié: (2025)
par: Hu, Tiancheng, et autres
Publié: (2025)
Beyond Demographics: Fine-tuning Large Language Models to Predict Individuals' Subjective Text Perceptions
par: Orlikowski, Matthias, et autres
Publié: (2025)
par: Orlikowski, Matthias, et autres
Publié: (2025)
Derivational Morphology Reveals Analogical Generalization in Large Language Models
par: Hofmann, Valentin, et autres
Publié: (2024)
par: Hofmann, Valentin, et autres
Publié: (2024)
Evaluating Contextually Mediated Factual Recall in Multilingual Large Language Models
par: Liu, Yihong, et autres
Publié: (2026)
par: Liu, Yihong, et autres
Publié: (2026)
LLaVA-Gemma: Accelerating Multimodal Foundation Models with a Compact Language Model
par: Hinck, Musashi, et autres
Publié: (2024)
par: Hinck, Musashi, et autres
Publié: (2024)
Hybrid Human-LLM Corpus Construction and LLM Evaluation for Rare Linguistic Phenomena
par: Weissweiler, Leonie, et autres
Publié: (2024)
par: Weissweiler, Leonie, et autres
Publié: (2024)
GLUScope: A Tool for Analyzing GLU Neurons in Transformer Language Models
par: Gerstner, Sebastian, et autres
Publié: (2026)
par: Gerstner, Sebastian, et autres
Publié: (2026)
GKnow: Measuring the Entanglement of Gender Bias and Factual Gender
par: Veloso, Leonor, et autres
Publié: (2026)
par: Veloso, Leonor, et autres
Publié: (2026)
Understanding Gated Neurons in Transformers from Their Input-Output Functionality
par: Gerstner, Sebastian, et autres
Publié: (2025)
par: Gerstner, Sebastian, et autres
Publié: (2025)
Verbing Weirds Language (Models): Evaluation of English Zero-Derivation in Five LLMs
par: Mortensen, David R., et autres
Publié: (2024)
par: Mortensen, David R., et autres
Publié: (2024)
Geographic Adaptation of Pretrained Language Models
par: Hofmann, Valentin, et autres
Publié: (2022)
par: Hofmann, Valentin, et autres
Publié: (2022)
No for Some, Yes for Others: Persona Prompts and Other Sources of False Refusal in Language Models
par: Plaza-del-Arco, Flor Miriam, et autres
Publié: (2025)
par: Plaza-del-Arco, Flor Miriam, et autres
Publié: (2025)
SLAyiNG: Towards Queer Language Processing
par: Veloso, Leonor, et autres
Publié: (2025)
par: Veloso, Leonor, et autres
Publié: (2025)
SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models
par: Vidgen, Bertie, et autres
Publié: (2023)
par: Vidgen, Bertie, et autres
Publié: (2023)
Debiasing Large Vision-Language Models by Ablating Protected Attribute Representations
par: Ratzlaff, Neale, et autres
Publié: (2024)
par: Ratzlaff, Neale, et autres
Publié: (2024)
RET-LLM: Towards a General Read-Write Memory for Large Language Models
par: Modarressi, Ali, et autres
Publié: (2023)
par: Modarressi, Ali, et autres
Publié: (2023)
LieCraft: A Multi-Agent Framework for Evaluating Deceptive Capabilities in Language Models
par: Olson, Matthew Lyle, et autres
Publié: (2026)
par: Olson, Matthew Lyle, et autres
Publié: (2026)
Left, Right, or Center? Evaluating LLM Framing in News Classification and Generation
par: Kennedy, Molly, et autres
Publié: (2026)
par: Kennedy, Molly, et autres
Publié: (2026)
What Do Dialect Speakers Want? A Survey of Attitudes Towards Language Technology for German Dialects
par: Blaschke, Verena, et autres
Publié: (2024)
par: Blaschke, Verena, et autres
Publié: (2024)
Conversations as a Source for Teaching Scientific Concepts at Different Education Levels
par: Rooein, Donya, et autres
Publié: (2024)
par: Rooein, Donya, et autres
Publié: (2024)
Narratives at Conflict: Computational Analysis of News Framing in Multilingual Disinformation Campaigns
par: Sinelnik, Antonina, et autres
Publié: (2024)
par: Sinelnik, Antonina, et autres
Publié: (2024)
"My Answer is C": First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models
par: Wang, Xinpeng, et autres
Publié: (2024)
par: Wang, Xinpeng, et autres
Publié: (2024)
Multilingual != Multicultural: Evaluating Gaps Between Multilingual Capabilities and Cultural Alignment in LLMs
par: Rystrøm, Jonathan, et autres
Publié: (2025)
par: Rystrøm, Jonathan, et autres
Publié: (2025)
Probing Semantic Routing in Large Mixture-of-Expert Models
par: Olson, Matthew Lyle, et autres
Publié: (2025)
par: Olson, Matthew Lyle, et autres
Publié: (2025)
Evaluate What You Can't Evaluate: Unassessable Quality for Generated Response
par: Liu, Yongkang, et autres
Publié: (2023)
par: Liu, Yongkang, et autres
Publié: (2023)
Promptly Predicting Structures: The Return of Inference
par: Mehta, Maitrey, et autres
Publié: (2024)
par: Mehta, Maitrey, et autres
Publié: (2024)
Relational Linearity is a Predictor of Hallucinations
par: Lu, Yuetian, et autres
Publié: (2026)
par: Lu, Yuetian, et autres
Publié: (2026)
Documents similaires
-
IssueBench: Millions of Realistic Prompts for Measuring Issue Bias in LLM Writing Assistance
par: Röttger, Paul, et autres
Publié: (2025) -
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
par: Röttger, Paul, et autres
Publié: (2023) -
The Pluralistic Moral Gap: Understanding Judgment and Value Differences between Humans and Large Language Models
par: Russo, Giuseppe, et autres
Publié: (2025) -
Compromesso! Italian Many-Shot Jailbreaks Undermine the Safety of Large Language Models
par: Pernisi, Fabio, et autres
Publié: (2024) -
SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety
par: Röttger, Paul, et autres
Publié: (2024)