How Utilitarian Are OpenAI's Models Really? Replicating and Reinterpreting Pfeffer, Krügel, and Uhl (2025)
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Himmelreich, Johannes |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Large language models can replicate cross-cultural differences in personality
von: Niszczota, Paweł, et al.
Veröffentlicht: (2023)
von: Niszczota, Paweł, et al.
Veröffentlicht: (2023)
PromptAug: Fine-grained Conflict Classification Using Data Augmentation
von: Warke, Oliver, et al.
Veröffentlicht: (2025)
von: Warke, Oliver, et al.
Veröffentlicht: (2025)
The Company You Keep: How LLMs Respond to Dark Triad Traits
von: Lu, Zeyi, et al.
Veröffentlicht: (2026)
von: Lu, Zeyi, et al.
Veröffentlicht: (2026)
Textual Entailment is not a Better Bias Metric than Token Probability
von: Felkner, Virginia K., et al.
Veröffentlicht: (2025)
von: Felkner, Virginia K., et al.
Veröffentlicht: (2025)
GPT is Not an Annotator: The Necessity of Human Annotation in Fairness Benchmark Construction
von: Felkner, Virginia K., et al.
Veröffentlicht: (2024)
von: Felkner, Virginia K., et al.
Veröffentlicht: (2024)
Chatbot Deployment Considerations for Application-Agnostic Human-Machine Dialogues
von: Rivas, Pablo, et al.
Veröffentlicht: (2025)
von: Rivas, Pablo, et al.
Veröffentlicht: (2025)
SectEval: Evaluating the Latent Sectarian Preferences of Large Language Models
von: Maheshwari, Aditya, et al.
Veröffentlicht: (2026)
von: Maheshwari, Aditya, et al.
Veröffentlicht: (2026)
Rejected Dialects: Biases Against African American Language in Reward Models
von: Mire, Joel, et al.
Veröffentlicht: (2025)
von: Mire, Joel, et al.
Veröffentlicht: (2025)
Benchmarking Bengali Dialectal Bias: A Multi-Stage Framework Integrating RAG-Based Translation and Human-Augmented RLAIF
von: Sami, K. M. Jubair, et al.
Veröffentlicht: (2026)
von: Sami, K. M. Jubair, et al.
Veröffentlicht: (2026)
Balancing Innovation and Integrity: AI Integration in Liberal Arts College Administration
von: Read, Ian Olivo
Veröffentlicht: (2025)
von: Read, Ian Olivo
Veröffentlicht: (2025)
Using a cognitive architecture to consider antiBlackness in design and development of AI systems
von: Dancy, Christopher L.
Veröffentlicht: (2022)
von: Dancy, Christopher L.
Veröffentlicht: (2022)
Reward Model Interpretability via Optimal and Pessimal Tokens
von: Christian, Brian, et al.
Veröffentlicht: (2025)
von: Christian, Brian, et al.
Veröffentlicht: (2025)
What are People Talking about in #BlackLivesMatter and #StopAsianHate? Exploring and Categorizing Twitter Topics Emerging in Online Social Movements through the Latent Dirichlet Allocation Model
von: Tong, Xin, et al.
Veröffentlicht: (2022)
von: Tong, Xin, et al.
Veröffentlicht: (2022)
How Growing Toxicity Manifests: A Topic Trajectory Analysis of U.S. Immigration Discourse on Social Media
von: Joh, Una, et al.
Veröffentlicht: (2025)
von: Joh, Una, et al.
Veröffentlicht: (2025)
Analysis of LLM as a grammatical feature tagger for African American English
von: Porwal, Rahul, et al.
Veröffentlicht: (2025)
von: Porwal, Rahul, et al.
Veröffentlicht: (2025)
On Fact and Frequency: LLM Responses to Misinformation Expressed with Uncertainty
von: van de Sande, Yana, et al.
Veröffentlicht: (2025)
von: van de Sande, Yana, et al.
Veröffentlicht: (2025)
Understanding Gen Alpha Digital Language: Evaluation of LLM Safety Systems for Content Moderation
von: Mehta, Manisha, et al.
Veröffentlicht: (2025)
von: Mehta, Manisha, et al.
Veröffentlicht: (2025)
NLP Occupational Emergence Analysis: How Occupations Form and Evolve in Real Time -- A Zero-Assumption Method Demonstrated on AI in the US Technology Workforce, 2022-2026
von: Nordfors, David
Veröffentlicht: (2026)
von: Nordfors, David
Veröffentlicht: (2026)
Evaluation of Hate Speech Detection Using Large Language Models and Geographical Contextualization
von: Zahid, Anwar Hossain, et al.
Veröffentlicht: (2025)
von: Zahid, Anwar Hossain, et al.
Veröffentlicht: (2025)
Collective Constitutional AI: Aligning a Language Model with Public Input
von: Huang, Saffron, et al.
Veröffentlicht: (2024)
von: Huang, Saffron, et al.
Veröffentlicht: (2024)
VEAT Quantifies Implicit Associations in Text-to-Video Generator Sora and Reveals Challenges in Bias Mitigation
von: Sun, Yongxu, et al.
Veröffentlicht: (2026)
von: Sun, Yongxu, et al.
Veröffentlicht: (2026)
Transforming Computer Security and Public Trust Through the Exploration of Fine-Tuning Large Language Models
von: Crumrine, Garrett, et al.
Veröffentlicht: (2024)
von: Crumrine, Garrett, et al.
Veröffentlicht: (2024)
EQUITRIAGE: A Fairness Audit of Gender Bias in LLM-Based Emergency Department Triage
von: Young, Richard J., et al.
Veröffentlicht: (2026)
von: Young, Richard J., et al.
Veröffentlicht: (2026)
ChatGPT as Research Scientist: Probing GPT's Capabilities as a Research Librarian, Research Ethicist, Data Generator and Data Predictor
von: Lehr, Steven A., et al.
Veröffentlicht: (2024)
von: Lehr, Steven A., et al.
Veröffentlicht: (2024)
Towards the Terminator Economy: Assessing Job Exposure to AI through LLMs
von: Colombo, Emilio, et al.
Veröffentlicht: (2024)
von: Colombo, Emilio, et al.
Veröffentlicht: (2024)
The Invisible Coalition Partner: How LLMs Vote When Democracy Gets Concrete
von: Barmettler, Joel
Veröffentlicht: (2026)
von: Barmettler, Joel
Veröffentlicht: (2026)
Disaster Question Answering with LoRA Efficiency and Accurate End Position
von: Yasuno, Takato
Veröffentlicht: (2026)
von: Yasuno, Takato
Veröffentlicht: (2026)
The Journal of Prompt-Engineered (Moral) Philosophy Or: Why AI-Assisted Ethics Research Requires Process Transparency
von: Loi, Michele
Veröffentlicht: (2025)
von: Loi, Michele
Veröffentlicht: (2025)
Benchmarking Educational LLMs with Analytics: A Case Study on Gender Bias in Feedback
von: Du, Yishan, et al.
Veröffentlicht: (2025)
von: Du, Yishan, et al.
Veröffentlicht: (2025)
Big AI is accelerating the metacrisis: What can we do?
von: Bird, Steven
Veröffentlicht: (2025)
von: Bird, Steven
Veröffentlicht: (2025)
Detecting Effects of AI-Mediated Communication on Language Complexity and Sentiment
von: Sussman, Kristen, et al.
Veröffentlicht: (2025)
von: Sussman, Kristen, et al.
Veröffentlicht: (2025)
Discursive objection strategies in online comments: Developing a classification schema and validating its training
von: Shea, Ashley L., et al.
Veröffentlicht: (2024)
von: Shea, Ashley L., et al.
Veröffentlicht: (2024)
Do Political Opinions Transfer Between Western Languages? An Analysis of Unaligned and Aligned Multilingual LLMs
von: Weeber, Franziska, et al.
Veröffentlicht: (2025)
von: Weeber, Franziska, et al.
Veröffentlicht: (2025)
Towards Fairer Health Recommendations: finding informative unbiased samples via Word Sense Disambiguation
von: Butts, Gavin, et al.
Veröffentlicht: (2024)
von: Butts, Gavin, et al.
Veröffentlicht: (2024)
Replicating TEMPEST at Scale: Multi-Turn Adversarial Attacks Against Trillion-Parameter Frontier Models
von: Young, Richard
Veröffentlicht: (2025)
von: Young, Richard
Veröffentlicht: (2025)
Assessing Crime Disclosure Patterns in a Large-Scale Cybercrime Forum
von: Hoheisel, Raphael, et al.
Veröffentlicht: (2026)
von: Hoheisel, Raphael, et al.
Veröffentlicht: (2026)
WSC+: Enhancing The Winograd Schema Challenge Using Tree-of-Experts
von: Zahraei, Pardis Sadat, et al.
Veröffentlicht: (2024)
von: Zahraei, Pardis Sadat, et al.
Veröffentlicht: (2024)
Leveraging Multi-Source Textural UGC for Neighbourhood Housing Quality Assessment: A GPT-Enhanced Framework
von: Hong, Qiyuan, et al.
Veröffentlicht: (2025)
von: Hong, Qiyuan, et al.
Veröffentlicht: (2025)
Beyond the Cloud: Assessing the Benefits and Drawbacks of Local LLM Deployment for Translators
von: Sandrini, Peter
Veröffentlicht: (2025)
von: Sandrini, Peter
Veröffentlicht: (2025)
Can Humans Tell? A Dual-Axis Study of Human Perception of LLM-Generated News
von: Loth, Alexander, et al.
Veröffentlicht: (2026)
von: Loth, Alexander, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Large language models can replicate cross-cultural differences in personality
von: Niszczota, Paweł, et al.
Veröffentlicht: (2023) -
PromptAug: Fine-grained Conflict Classification Using Data Augmentation
von: Warke, Oliver, et al.
Veröffentlicht: (2025) -
The Company You Keep: How LLMs Respond to Dark Triad Traits
von: Lu, Zeyi, et al.
Veröffentlicht: (2026) -
Textual Entailment is not a Better Bias Metric than Token Probability
von: Felkner, Virginia K., et al.
Veröffentlicht: (2025) -
GPT is Not an Annotator: The Necessity of Human Annotation in Fairness Benchmark Construction
von: Felkner, Virginia K., et al.
Veröffentlicht: (2024)