RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity
Fuente:
arXiv
Saved in:
| Main Authors: | Shin, Jisu, Song, Hoyun, Oh, Juhyun, Ko, Changgeon, Kim, Eunsu, Jung, Chani, Oh, Alice |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Spotting Out-of-Character Behavior: Atomic-Level Evaluation of Persona Fidelity in Open-Ended Generation
by: Shin, Jisu, et al.
Published: (2025)
by: Shin, Jisu, et al.
Published: (2025)
Different Bias Under Different Criteria: Assessing Bias in LLMs with a Fact-Based Approach
by: Ko, Changgeon, et al.
Published: (2024)
by: Ko, Changgeon, et al.
Published: (2024)
Flex-TravelPlanner: A Benchmark for Flexible Planning with Language Agents
by: Oh, Juhyun, et al.
Published: (2025)
by: Oh, Juhyun, et al.
Published: (2025)
Beyond Static Benchmarks: Synthesizing Harmful Content via Persona-based Simulation for Robust Evaluation
by: Lee, Huije, et al.
Published: (2026)
by: Lee, Huije, et al.
Published: (2026)
The Generative AI Paradox on Evaluation: What It Can Solve, It May Not Evaluate
by: Oh, Juhyun, et al.
Published: (2024)
by: Oh, Juhyun, et al.
Published: (2024)
Multi-FAct: Assessing Factuality of Multilingual LLMs using FActScore
by: Shafayat, Sheikh, et al.
Published: (2024)
by: Shafayat, Sheikh, et al.
Published: (2024)
FINEST: Improving LLM Responses to Sensitive Topics Through Fine-Grained Evaluation
by: Oh, Juhyun, et al.
Published: (2026)
by: Oh, Juhyun, et al.
Published: (2026)
Ask LLMs Directly, "What shapes your bias?": Measuring Social Bias in Large Language Models
by: Shin, Jisu, et al.
Published: (2024)
by: Shin, Jisu, et al.
Published: (2024)
Are they lovers or friends? Evaluating LLMs' Social Reasoning in English and Korean Dialogues
by: Kim, Eunsu, et al.
Published: (2025)
by: Kim, Eunsu, et al.
Published: (2025)
Does Rationale Quality Matter? Enhancing Mental Disorder Detection via Selective Reasoning Distillation
by: Song, Hoyun, et al.
Published: (2025)
by: Song, Hoyun, et al.
Published: (2025)
MentalBench: A DSM-Grounded Benchmark for Evaluating Psychiatric Diagnostic Capability of Large Language Models
by: Song, Hoyun, et al.
Published: (2026)
by: Song, Hoyun, et al.
Published: (2026)
Social Dynamics as Critical Vulnerabilities that Undermine Objective Decision-Making in LLM Collectives
by: Ko, Changgeon, et al.
Published: (2026)
by: Ko, Changgeon, et al.
Published: (2026)
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation
by: Kim, Eunsu, et al.
Published: (2025)
by: Kim, Eunsu, et al.
Published: (2025)
Getting Bored of Cyberwar: Exploring the Role of Low-level Cybercrime Actors in the Russia-Ukraine Conflict
by: Vu, Anh V., et al.
Published: (2022)
by: Vu, Anh V., et al.
Published: (2022)
Uncovering Factor Level Preferences to Improve Human-Model Alignment
by: Oh, Juhyun, et al.
Published: (2024)
by: Oh, Juhyun, et al.
Published: (2024)
Measuring Interest Group Positions on Legislation: An AI-Driven Analysis of Lobbying Reports
by: Kim, Jiseon, et al.
Published: (2025)
by: Kim, Jiseon, et al.
Published: (2025)
CLIcK: A Benchmark Dataset of Cultural and Linguistic Intelligence in Korean
by: Kim, Eunsu, et al.
Published: (2024)
by: Kim, Eunsu, et al.
Published: (2024)
From Protest to Power Plant: Interpreting the Role of Escalatory Hacktivism in Cyber Conflict
by: Derbyshire, Richard, et al.
Published: (2025)
by: Derbyshire, Richard, et al.
Published: (2025)
A Multi-Task Benchmark for Abusive Language Detection in Low-Resource Settings
by: Gaim, Fitsum, et al.
Published: (2025)
by: Gaim, Fitsum, et al.
Published: (2025)
MUG-Eval: A Proxy Evaluation Framework for Multilingual Generation Capabilities in Any Language
by: Song, Seyoung, et al.
Published: (2025)
by: Song, Seyoung, et al.
Published: (2025)
"Make It Sound Like a Lawyer Wrote It": Scenarios of Potential Impacts of Generative AI for Legal Conflict Resolution
by: Kieslich, Kimon, et al.
Published: (2026)
by: Kieslich, Kimon, et al.
Published: (2026)
XCR-Bench: A Multi-Task Benchmark for Evaluating Cultural Reasoning in LLMs
by: Kabir, Mohsinul, et al.
Published: (2026)
by: Kabir, Mohsinul, et al.
Published: (2026)
OLA: Output Language Alignment in Code-Switched LLM Interactions
by: Oh, Juhyun, et al.
Published: (2026)
by: Oh, Juhyun, et al.
Published: (2026)
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs
by: Oh, Gyutaek, et al.
Published: (2025)
by: Oh, Gyutaek, et al.
Published: (2025)
Culture is Everywhere: A Call for Intentionally Cultural Evaluation
by: Oh, Juhyun, et al.
Published: (2025)
by: Oh, Juhyun, et al.
Published: (2025)
PLawBench: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal Practice
by: Shi, Yuzhen, et al.
Published: (2026)
by: Shi, Yuzhen, et al.
Published: (2026)
Robust Misinformation Detection by Visiting Potential Commonsense Conflict
by: Wang, Bing, et al.
Published: (2025)
by: Wang, Bing, et al.
Published: (2025)
Computational Sociology of Humans and Machines; Conflict and Collaboration
by: Yasseri, Taha
Published: (2024)
by: Yasseri, Taha
Published: (2024)
LLM-Driven Rubric-Based Assessment of Algebraic Competence in Multi-Stage Block Coding Tasks with Design and Field Evaluation
by: Lee, Yong Oh, et al.
Published: (2025)
by: Lee, Yong Oh, et al.
Published: (2025)
ClinBench-HPB: A Clinical Benchmark for Evaluating LLMs in Hepato-Pancreato-Biliary Diseases
by: Li, Yuchong, et al.
Published: (2025)
by: Li, Yuchong, et al.
Published: (2025)
Moral Sensitivity in LLMs: A Tiered Evaluation of Contextual Bias via Behavioral Profiling and Mechanistic Interpretability
by: Aggarwal, Yash, et al.
Published: (2026)
by: Aggarwal, Yash, et al.
Published: (2026)
LLM-as-an-Interviewer: Beyond Static Testing Through Dynamic LLM Evaluation
by: Kim, Eunsu, et al.
Published: (2024)
by: Kim, Eunsu, et al.
Published: (2024)
An AI-Based Framework for Assessing Sustainability Conflicts in Medical Device Development
by: Chakrabarti, Apala
Published: (2025)
by: Chakrabarti, Apala
Published: (2025)
JuICE: A Benchmark for Evaluating LLM-Judge in Identifying Cultural Errors
by: Jin, Jiho, et al.
Published: (2026)
by: Jin, Jiho, et al.
Published: (2026)
Quantifying the Risk of Pastoral Conflict in 4 Central African Countries
by: Solaa, Lirika, et al.
Published: (2024)
by: Solaa, Lirika, et al.
Published: (2024)
Exploring Persona-dependent LLM Alignment for the Moral Machine Experiment
by: Kim, Jiseon, et al.
Published: (2025)
by: Kim, Jiseon, et al.
Published: (2025)
Blockchain and Artificial Intelligence: Synergies and Conflicts
by: Witt, Leon, et al.
Published: (2024)
by: Witt, Leon, et al.
Published: (2024)
Perceptions to Beliefs: Exploring Precursory Inferences for Theory of Mind in Large Language Models
by: Jung, Chani, et al.
Published: (2024)
by: Jung, Chani, et al.
Published: (2024)
LoCar: Localization-Aware Evaluation of In-Vehicle Assistants through Fine-Grained Sociolinguistic Control
by: Jeong, Seogyeong, et al.
Published: (2026)
by: Jeong, Seogyeong, et al.
Published: (2026)
LocalBench: Benchmarking LLMs on County-Level Local Knowledge and Reasoning
by: Gao, Zihan, et al.
Published: (2025)
by: Gao, Zihan, et al.
Published: (2025)
Similar Items
-
Spotting Out-of-Character Behavior: Atomic-Level Evaluation of Persona Fidelity in Open-Ended Generation
by: Shin, Jisu, et al.
Published: (2025) -
Different Bias Under Different Criteria: Assessing Bias in LLMs with a Fact-Based Approach
by: Ko, Changgeon, et al.
Published: (2024) -
Flex-TravelPlanner: A Benchmark for Flexible Planning with Language Agents
by: Oh, Juhyun, et al.
Published: (2025) -
Beyond Static Benchmarks: Synthesizing Harmful Content via Persona-based Simulation for Robust Evaluation
by: Lee, Huije, et al.
Published: (2026) -
The Generative AI Paradox on Evaluation: What It Can Solve, It May Not Evaluate
by: Oh, Juhyun, et al.
Published: (2024)