The Potential and Challenges of Evaluating Attitudes, Opinions, and Values in Large Language Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Ma, Bolei, Wang, Xinpeng, Hu, Tiancheng, Haensch, Anna-Carolina, Hedderich, Michael A., Plank, Barbara, Kreuter, Frauke |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Algorithmic Fidelity of Large Language Models in Generating Synthetic German Public Opinions: A Case Study
par: Ma, Bolei, et autres
Publié: (2024)
par: Ma, Bolei, et autres
Publié: (2024)
Too Open for Opinion? Embracing Open-Endedness in Large Language Models for Social Simulation
par: Ma, Bolei, et autres
Publié: (2025)
par: Ma, Bolei, et autres
Publié: (2025)
Capabilities and Evaluation Biases of Large Language Models in Classical Chinese Poetry Generation: A Case Study on Tang Poetry
par: Ma, Bolei, et autres
Publié: (2025)
par: Ma, Bolei, et autres
Publié: (2025)
"My Answer is C": First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models
par: Wang, Xinpeng, et autres
Publié: (2024)
par: Wang, Xinpeng, et autres
Publié: (2024)
Pragmatics in the Era of Large Language Models: A Survey on Datasets, Evaluation, Opportunities and Challenges
par: Ma, Bolei, et autres
Publié: (2025)
par: Ma, Bolei, et autres
Publié: (2025)
Aligning NLP Models with Target Population Perspectives using PAIR: Population-Aligned Instance Replication
par: Eckman, Stephanie, et autres
Publié: (2025)
par: Eckman, Stephanie, et autres
Publié: (2025)
Look at the Text: Instruction-Tuned Language Models are More Robust Multiple Choice Selectors than You Think
par: Wang, Xinpeng, et autres
Publié: (2024)
par: Wang, Xinpeng, et autres
Publié: (2024)
Survey Response Generation: Generating Closed-Ended Survey Responses In-Silico with Large Language Models
par: Ahnert, Georg, et autres
Publié: (2025)
par: Ahnert, Georg, et autres
Publié: (2025)
Understanding When Tree of Thoughts Succeeds: Larger Models Excel in Generation, Not Discrimination
par: Chen, Qiqi, et autres
Publié: (2024)
par: Chen, Qiqi, et autres
Publié: (2024)
Decomposed Prompting: Probing Multilingual Linguistic Structure Knowledge in Large Language Models
par: Nie, Ercong, et autres
Publié: (2024)
par: Nie, Ercong, et autres
Publié: (2024)
Surgical, Cheap, and Flexible: Mitigating False Refusal in Language Models via Single Vector Ablation
par: Wang, Xinpeng, et autres
Publié: (2024)
par: Wang, Xinpeng, et autres
Publié: (2024)
MAKIEval: A Multilingual Automatic WiKidata-based Framework for Cultural Awareness Evaluation for LLMs
par: Zhao, Raoyuan, et autres
Publié: (2025)
par: Zhao, Raoyuan, et autres
Publié: (2025)
Annotation Sensitivity: Training Data Collection Methods Affect Model Performance
par: Kern, Christoph, et autres
Publié: (2023)
par: Kern, Christoph, et autres
Publié: (2023)
Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models
par: Ball, Sarah, et autres
Publié: (2024)
par: Ball, Sarah, et autres
Publié: (2024)
Systematic Evaluation of Uncertainty Estimation Methods in Large Language Models
par: Hobelsberger, Christian, et autres
Publié: (2025)
par: Hobelsberger, Christian, et autres
Publié: (2025)
Moral Lenses, Political Coordinates: Towards Ideological Positioning of Morally Conditioned LLMs
par: Yuan, Chenchen, et autres
Publié: (2026)
par: Yuan, Chenchen, et autres
Publié: (2026)
ToPro: Token-Level Prompt Decomposition for Cross-Lingual Sequence Labeling Tasks
par: Ma, Bolei, et autres
Publié: (2024)
par: Ma, Bolei, et autres
Publié: (2024)
United in Diversity? Contextual Biases in LLM-Based Predictions of the 2024 European Parliament Elections
par: von der Heyde, Leah, et autres
Publié: (2024)
par: von der Heyde, Leah, et autres
Publié: (2024)
Position: Insights from Survey Methodology can Improve Training Data
par: Eckman, Stephanie, et autres
Publié: (2024)
par: Eckman, Stephanie, et autres
Publié: (2024)
The Imperfective Paradox in Large Language Models
par: Ma, Bolei, et autres
Publié: (2026)
par: Ma, Bolei, et autres
Publié: (2026)
Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
par: Mondorf, Philipp, et autres
Publié: (2024)
par: Mondorf, Philipp, et autres
Publié: (2024)
Probing LLMs for Multilingual Discourse Generalization Through a Unified Label Set
par: Eichin, Florian, et autres
Publié: (2025)
par: Eichin, Florian, et autres
Publié: (2025)
"It Listens Better Than My Therapist": Exploring Social Media Discourse on LLMs as Mental Health Tool
par: Haensch, Anna-Carolina
Publié: (2025)
par: Haensch, Anna-Carolina
Publié: (2025)
Refusal Direction is Universal Across Safety-Aligned Languages
par: Wang, Xinpeng, et autres
Publié: (2025)
par: Wang, Xinpeng, et autres
Publié: (2025)
Evaluating Robustness of Large Language Models Against Multilingual Typographical Errors
par: Zhao, Raoyuan, et autres
Publié: (2025)
par: Zhao, Raoyuan, et autres
Publié: (2025)
Exploring Large Language Models for Product Attribute Value Identification
par: Sabeh, Kassem, et autres
Publié: (2024)
par: Sabeh, Kassem, et autres
Publié: (2024)
Liar, Liar, Logical Mire: A Benchmark for Suppositional Reasoning in Large Language Models
par: Mondorf, Philipp, et autres
Publié: (2024)
par: Mondorf, Philipp, et autres
Publié: (2024)
What's the Difference? Supporting Users in Identifying the Effects of Prompt and Model Changes Through Token Patterns
par: Hedderich, Michael A., et autres
Publié: (2025)
par: Hedderich, Michael A., et autres
Publié: (2025)
Comparing Inferential Strategies of Humans and Large Language Models in Deductive Reasoning
par: Mondorf, Philipp, et autres
Publié: (2024)
par: Mondorf, Philipp, et autres
Publié: (2024)
Toward Understanding the Transferability of Adversarial Suffixes in Large Language Models
par: Ball, Sarah, et autres
Publié: (2025)
par: Ball, Sarah, et autres
Publié: (2025)
To share or not to share: What risks would laypeople accept to give sensitive data to differentially-private NLP systems?
par: Weiss, Christopher, et autres
Publié: (2023)
par: Weiss, Christopher, et autres
Publié: (2023)
Table Question Answering in the Era of Large Language Models: A Comprehensive Survey of Tasks, Methods, and Evaluation
par: Zhou, Wei, et autres
Publié: (2025)
par: Zhou, Wei, et autres
Publié: (2025)
Decoupling the Effect of Chain-of-Thought Reasoning: A Human Label Variation Perspective
par: Chen, Beiduo, et autres
Publié: (2026)
par: Chen, Beiduo, et autres
Publié: (2026)
Why Lift so Heavy? Slimming Large Language Models by Cutting Off the Layers
par: Yuan, Shuzhou, et autres
Publié: (2024)
par: Yuan, Shuzhou, et autres
Publié: (2024)
Copy First, Translate Later: Interpreting Translation Dynamics in Multilingual Pretraining
par: Körner, Felicia, et autres
Publié: (2026)
par: Körner, Felicia, et autres
Publié: (2026)
Evaluating Large Language Models for Cross-Lingual Retrieval
par: Zuo, Longfei, et autres
Publié: (2025)
par: Zuo, Longfei, et autres
Publié: (2025)
AIn't Nothing But a Survey? Using Large Language Models for Coding German Open-Ended Survey Responses on Survey Motivation
par: von der Heyde, Leah, et autres
Publié: (2025)
par: von der Heyde, Leah, et autres
Publié: (2025)
If Probable, Then Acceptable? Understanding Conditional Acceptability Judgments in Large Language Models
par: Orth, Jasmin, et autres
Publié: (2025)
par: Orth, Jasmin, et autres
Publié: (2025)
"Seeing the Big through the Small": Can LLMs Approximate Human Judgment Distributions on NLI from a Few Explanations?
par: Chen, Beiduo, et autres
Publié: (2024)
par: Chen, Beiduo, et autres
Publié: (2024)
What Do Dialect Speakers Want? A Survey of Attitudes Towards Language Technology for German Dialects
par: Blaschke, Verena, et autres
Publié: (2024)
par: Blaschke, Verena, et autres
Publié: (2024)
Documents similaires
-
Algorithmic Fidelity of Large Language Models in Generating Synthetic German Public Opinions: A Case Study
par: Ma, Bolei, et autres
Publié: (2024) -
Too Open for Opinion? Embracing Open-Endedness in Large Language Models for Social Simulation
par: Ma, Bolei, et autres
Publié: (2025) -
Capabilities and Evaluation Biases of Large Language Models in Classical Chinese Poetry Generation: A Case Study on Tang Poetry
par: Ma, Bolei, et autres
Publié: (2025) -
"My Answer is C": First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models
par: Wang, Xinpeng, et autres
Publié: (2024) -
Pragmatics in the Era of Large Language Models: A Survey on Datasets, Evaluation, Opportunities and Challenges
par: Ma, Bolei, et autres
Publié: (2025)