Enregistré dans:
| Auteurs principaux: | Kim, Kyuhee, Lee, Sangah |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2507.04014 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
KoCoNovel: Annotated Dataset of Character Coreference in Korean Novels
par: Kim, Kyuhee, et autres
Publié: (2024)
par: Kim, Kyuhee, et autres
Publié: (2024)
K-Act2Emo: Korean Commonsense Knowledge Graph for Indirect Emotional Expression
par: Kim, Kyuhee, et autres
Publié: (2024)
par: Kim, Kyuhee, et autres
Publié: (2024)
XCR-Bench: A Multi-Task Benchmark for Evaluating Cultural Reasoning in LLMs
par: Kabir, Mohsinul, et autres
Publié: (2026)
par: Kabir, Mohsinul, et autres
Publié: (2026)
DarkBench: Benchmarking Dark Patterns in Large Language Models
par: Kran, Esben, et autres
Publié: (2025)
par: Kran, Esben, et autres
Publié: (2025)
WorldView-Bench: A Benchmark for Evaluating Global Cultural Perspectives in Large Language Models
par: Mushtaq, Abdullah, et autres
Publié: (2025)
par: Mushtaq, Abdullah, et autres
Publié: (2025)
LocalBench: Benchmarking LLMs on County-Level Local Knowledge and Reasoning
par: Gao, Zihan, et autres
Publié: (2025)
par: Gao, Zihan, et autres
Publié: (2025)
BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses
par: Xu, Xin, et autres
Publié: (2025)
par: Xu, Xin, et autres
Publié: (2025)
SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors
par: Hu, Tiancheng, et autres
Publié: (2025)
par: Hu, Tiancheng, et autres
Publié: (2025)
RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity
par: Shin, Jisu, et autres
Publié: (2025)
par: Shin, Jisu, et autres
Publié: (2025)
KVoiceBench, KOpenAudioBench, and KMMAU: Agent-Driven Korean Speech Benchmarks for Evaluating SpeechLMs
par: Kim, Haechan, et autres
Publié: (2026)
par: Kim, Haechan, et autres
Publié: (2026)
LocalValueBench: A Collaboratively Built and Extensible Benchmark for Evaluating Localized Value Alignment and Ethical Safety in Large Language Models
par: Meadows, Gwenyth Isobel, et autres
Publié: (2024)
par: Meadows, Gwenyth Isobel, et autres
Publié: (2024)
VideoNorms: Benchmarking Cultural Awareness of Video Language Models
par: Varimalla, Nikhil Reddy, et autres
Publié: (2025)
par: Varimalla, Nikhil Reddy, et autres
Publié: (2025)
QueerBench: Quantifying Discrimination in Language Models Toward Queer Identities
par: Sosto, Mae, et autres
Publié: (2024)
par: Sosto, Mae, et autres
Publié: (2024)
CBT-Bench: Evaluating Large Language Models on Assisting Cognitive Behavior Therapy
par: Zhang, Mian, et autres
Publié: (2024)
par: Zhang, Mian, et autres
Publié: (2024)
Can Large Language Models Replace Human Coders? Introducing ContentBench
par: Haman, Michael
Publié: (2026)
par: Haman, Michael
Publié: (2026)
Psychological Assessments with Large Language Models: A Privacy-Focused and Cost-Effective Approach
par: Blanco-Cuaresma, Sergi
Publié: (2024)
par: Blanco-Cuaresma, Sergi
Publié: (2024)
LAPIS: Language Model-Augmented Police Investigation System
par: Kim, Heedou, et autres
Publié: (2024)
par: Kim, Heedou, et autres
Publié: (2024)
ClinBench-HPB: A Clinical Benchmark for Evaluating LLMs in Hepato-Pancreato-Biliary Diseases
par: Li, Yuchong, et autres
Publié: (2025)
par: Li, Yuchong, et autres
Publié: (2025)
PLawBench: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal Practice
par: Shi, Yuzhen, et autres
Publié: (2026)
par: Shi, Yuzhen, et autres
Publié: (2026)
Quantifying Risk Propensities of Large Language Models: Ethical Focus and Bias Detection through Role-Play
par: Zeng, Yifan, et autres
Publié: (2024)
par: Zeng, Yifan, et autres
Publié: (2024)
MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes
par: Chiu, Yu Ying, et autres
Publié: (2025)
par: Chiu, Yu Ying, et autres
Publié: (2025)
OpenTuringBench: An Open-Model-based Benchmark and Framework for Machine-Generated Text Detection and Attribution
par: La Cava, Lucio, et autres
Publié: (2025)
par: La Cava, Lucio, et autres
Publié: (2025)
Language Models and Logic Programs for Trustworthy Tax Reasoning
par: Jurayj, William, et autres
Publié: (2025)
par: Jurayj, William, et autres
Publié: (2025)
Dual Traits in Probabilistic Reasoning of Large Language Models
par: Li, Shenxiong, et autres
Publié: (2024)
par: Li, Shenxiong, et autres
Publié: (2024)
Can LLMs Think Like Consumers? Benchmarking Crowd-Level Reaction Reconstruction with ConsumerSimBench
par: Wang, Tianyu, et autres
Publié: (2026)
par: Wang, Tianyu, et autres
Publié: (2026)
Invisible Filters: Cultural Bias in Hiring Evaluations Using Large Language Models
par: Rao, Pooja S. B., et autres
Publié: (2025)
par: Rao, Pooja S. B., et autres
Publié: (2025)
Training-Free Cultural Alignment of Large Language Models via Persona Disagreement
par: Kiet, Huynh Trung, et autres
Publié: (2026)
par: Kiet, Huynh Trung, et autres
Publié: (2026)
Navigating the Cultural Kaleidoscope: A Hitchhiker's Guide to Sensitivity in Large Language Models
par: Banerjee, Somnath, et autres
Publié: (2024)
par: Banerjee, Somnath, et autres
Publié: (2024)
Open Ko-LLM Leaderboard: Evaluating Large Language Models in Korean with Ko-H5 Benchmark
par: Park, Chanjun, et autres
Publié: (2024)
par: Park, Chanjun, et autres
Publié: (2024)
AccessEval: Benchmarking Disability Bias in Large Language Models
par: Panda, Srikant, et autres
Publié: (2025)
par: Panda, Srikant, et autres
Publié: (2025)
No Culture Left Behind: ArtELingo-28, a Benchmark of WikiArt with Captions in 28 Languages
par: Mohamed, Youssef, et autres
Publié: (2024)
par: Mohamed, Youssef, et autres
Publié: (2024)
ChatBench: From Static Benchmarks to Human-AI Evaluation
par: Chang, Serina, et autres
Publié: (2025)
par: Chang, Serina, et autres
Publié: (2025)
HypoBench: Towards Systematic and Principled Benchmarking for Hypothesis Generation
par: Liu, Haokun, et autres
Publié: (2025)
par: Liu, Haokun, et autres
Publié: (2025)
HugAgent: Benchmarking LLMs for Simulation of Individualized Human Reasoning
par: Li, Chance Jiajie, et autres
Publié: (2025)
par: Li, Chance Jiajie, et autres
Publié: (2025)
STOP! Benchmarking Large Language Models with Sensitivity Testing on Offensive Progressions
par: Morabito, Robert, et autres
Publié: (2024)
par: Morabito, Robert, et autres
Publié: (2024)
Beyond Instrumental and Substitutive Paradigms: Introducing Machine Culture as an Emergent Phenomenon in Large Language Models
par: Hu, Yueqing, et autres
Publié: (2026)
par: Hu, Yueqing, et autres
Publié: (2026)
Policy Frameworks for Transparent Chain-of-Thought Reasoning in Large Language Models
par: Chen, Yihang, et autres
Publié: (2025)
par: Chen, Yihang, et autres
Publié: (2025)
Do LLMs Game Formalization? Evaluating Faithfulness in Logical Reasoning
par: Kim, Kyuhee, et autres
Publié: (2026)
par: Kim, Kyuhee, et autres
Publié: (2026)
EthicsMH: A Pilot Benchmark for Ethical Reasoning in Mental Health AI
par: Kasu, Sai Kartheek Reddy
Publié: (2025)
par: Kasu, Sai Kartheek Reddy
Publié: (2025)
Dr.Academy: A Benchmark for Evaluating Questioning Capability in Education for Large Language Models
par: Chen, Yuyan, et autres
Publié: (2024)
par: Chen, Yuyan, et autres
Publié: (2024)
Documents similaires
-
KoCoNovel: Annotated Dataset of Character Coreference in Korean Novels
par: Kim, Kyuhee, et autres
Publié: (2024) -
K-Act2Emo: Korean Commonsense Knowledge Graph for Indirect Emotional Expression
par: Kim, Kyuhee, et autres
Publié: (2024) -
XCR-Bench: A Multi-Task Benchmark for Evaluating Cultural Reasoning in LLMs
par: Kabir, Mohsinul, et autres
Publié: (2026) -
DarkBench: Benchmarking Dark Patterns in Large Language Models
par: Kran, Esben, et autres
Publié: (2025) -
WorldView-Bench: A Benchmark for Evaluating Global Cultural Perspectives in Large Language Models
par: Mushtaq, Abdullah, et autres
Publié: (2025)