KorNAT: LLM Alignment Benchmark for Korean Social Values and Common Knowledge
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Jiyoung, Kim, Minwoo, Kim, Seungho, Kim, Junghwan, Won, Seunghyun, Lee, Hwaran, Choi, Edward |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
KorMedMCQA: Multi-Choice Question Answering Benchmark for Korean Healthcare Professional Licensing Examinations
by: Kweon, Sunjun, et al.
Published: (2024)
by: Kweon, Sunjun, et al.
Published: (2024)
KoBBQ: Korean Bias Benchmark for Question Answering
by: Jin, Jiho, et al.
Published: (2023)
by: Jin, Jiho, et al.
Published: (2023)
Alignment Data Map for Efficient Preference Data Selection and Diagnosis
by: Lee, Seohyeong, et al.
Published: (2025)
by: Lee, Seohyeong, et al.
Published: (2025)
Trans-EnV: A Framework for Evaluating the Linguistic Robustness of LLMs Against English Varieties
by: Lee, Jiyoung, et al.
Published: (2025)
by: Lee, Jiyoung, et al.
Published: (2025)
KorMedMCQA-V: A Multimodal Benchmark for Evaluating Vision-Language Models on the Korean Medical Licensing Examination
by: Choi, Byungjin, et al.
Published: (2026)
by: Choi, Byungjin, et al.
Published: (2026)
HAE-RAE Bench: Evaluation of Korean Knowledge in Language Models
by: Son, Guijin, et al.
Published: (2023)
by: Son, Guijin, et al.
Published: (2023)
STAR-Teaming: A Strategy-Response Multiplex Network Approach to Automated LLM Red Teaming
by: Jung, MinJae, et al.
Published: (2026)
by: Jung, MinJae, et al.
Published: (2026)
EHRNoteQA: An LLM Benchmark for Real-World Clinical Practice Using Discharge Summaries
by: Kweon, Sunjun, et al.
Published: (2024)
by: Kweon, Sunjun, et al.
Published: (2024)
Expanding Foundational Language Capabilities in Open-Source LLMs through a Korean Case Study
by: Lim, Junghwan, et al.
Published: (2025)
by: Lim, Junghwan, et al.
Published: (2025)
Open Ko-LLM Leaderboard: Evaluating Large Language Models in Korean with Ko-H5 Benchmark
by: Park, Chanjun, et al.
Published: (2024)
by: Park, Chanjun, et al.
Published: (2024)
Drift: Decoding-time Personalized Alignments with Implicit User Preferences
by: Kim, Minbeom, et al.
Published: (2025)
by: Kim, Minbeom, et al.
Published: (2025)
HanjaBridge: Resolving Semantic Ambiguity in Korean LLMs via Hanja-Augmented Pre-Training
by: Choi, Seungho
Published: (2025)
by: Choi, Seungho
Published: (2025)
K-Act2Emo: Korean Commonsense Knowledge Graph for Indirect Emotional Expression
by: Kim, Kyuhee, et al.
Published: (2024)
by: Kim, Kyuhee, et al.
Published: (2024)
K/DA: Automated Data Generation Pipeline for Detoxifying Implicitly Offensive Language in Korean
by: Jeon, Minkyeong, et al.
Published: (2025)
by: Jeon, Minkyeong, et al.
Published: (2025)
SDS KoPub VDR: A Benchmark Dataset for Visual Document Retrieval in Korean Public Documents
by: Lee, Jaehoon, et al.
Published: (2025)
by: Lee, Jaehoon, et al.
Published: (2025)
Nunchi-Bench: Benchmarking Language Models on Cultural Reasoning with a Focus on Korean Superstition
by: Kim, Kyuhee, et al.
Published: (2025)
by: Kim, Kyuhee, et al.
Published: (2025)
From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation
by: Hong, Seokhee, et al.
Published: (2025)
by: Hong, Seokhee, et al.
Published: (2025)
KoCoSa: Korean Context-aware Sarcasm Detection Dataset
by: Kim, Yumin, et al.
Published: (2024)
by: Kim, Yumin, et al.
Published: (2024)
Open Ko-LLM Leaderboard2: Bridging Foundational and Practical Evaluation for Korean LLMs
by: Kim, Hyeonwoo, et al.
Published: (2024)
by: Kim, Hyeonwoo, et al.
Published: (2024)
AdvisorQA: Towards Helpful and Harmless Advice-seeking Question Answering with Collective Intelligence
by: Kim, Minbeom, et al.
Published: (2024)
by: Kim, Minbeom, et al.
Published: (2024)
Thunder-LLM: Efficiently Adapting LLMs to Korean with Minimal Resources
by: Kim, Jinpyo, et al.
Published: (2025)
by: Kim, Jinpyo, et al.
Published: (2025)
KoBALT: Korean Benchmark For Advanced Linguistic Tasks
by: Shin, Hyopil, et al.
Published: (2025)
by: Shin, Hyopil, et al.
Published: (2025)
"Well, Keep Thinking": Enhancing LLM Reasoning with Adaptive Injection Decoding
by: Jin, Hyunbin, et al.
Published: (2025)
by: Jin, Hyunbin, et al.
Published: (2025)
KoCoNovel: Annotated Dataset of Character Coreference in Korean Novels
by: Kim, Kyuhee, et al.
Published: (2024)
by: Kim, Kyuhee, et al.
Published: (2024)
Thunder-KoNUBench: A Corpus-Aligned Benchmark for Korean Negation Understanding
by: Jung, Sungmok, et al.
Published: (2026)
by: Jung, Sungmok, et al.
Published: (2026)
TimeChara: Evaluating Point-in-Time Character Hallucination of Role-Playing Large Language Models
by: Ahn, Jaewoo, et al.
Published: (2024)
by: Ahn, Jaewoo, et al.
Published: (2024)
K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts
by: Lee, Nahyun, et al.
Published: (2026)
by: Lee, Nahyun, et al.
Published: (2026)
Can Code-Switched Texts Activate a Knowledge Switch in LLMs? A Case Study on English-Korean Code-Switching
by: Kim, Seoyeon, et al.
Published: (2024)
by: Kim, Seoyeon, et al.
Published: (2024)
KRETA: A Benchmark for Korean Reading and Reasoning in Text-Rich VQA Attuned to Diverse Visual Contexts
by: Hwang, Taebaek, et al.
Published: (2025)
by: Hwang, Taebaek, et al.
Published: (2025)
LifeTox: Unveiling Implicit Toxicity in Life Advice
by: Kim, Minbeom, et al.
Published: (2023)
by: Kim, Minbeom, et al.
Published: (2023)
Code-Switching Red-Teaming: LLM Evaluation for Safety and Multilingual Understanding
by: Yoo, Haneul, et al.
Published: (2024)
by: Yoo, Haneul, et al.
Published: (2024)
Framing Matters: Addressing Framing Sensitivity in Decision-Making through Behaviorally-Grounded Value Alignment
by: Hwang, Seojin, et al.
Published: (2026)
by: Hwang, Seojin, et al.
Published: (2026)
Making Qwen3 Think in Korean with Reinforcement Learning
by: Lee, Jungyup, et al.
Published: (2025)
by: Lee, Jungyup, et al.
Published: (2025)
CAGE: A Framework for Culturally Adaptive Red-Teaming Benchmark Generation
by: Kim, Chaeyun, et al.
Published: (2026)
by: Kim, Chaeyun, et al.
Published: (2026)
KVoiceBench, KOpenAudioBench, and KMMAU: Agent-Driven Korean Speech Benchmarks for Evaluating SpeechLMs
by: Kim, Haechan, et al.
Published: (2026)
by: Kim, Haechan, et al.
Published: (2026)
SAIF: A Comprehensive Framework for Evaluating the Risks of Generative AI in the Public Sector
by: Lee, Kyeongryul, et al.
Published: (2025)
by: Lee, Kyeongryul, et al.
Published: (2025)
KoSimpleQA: A Korean Factuality Benchmark with an Analysis of Reasoning LLMs
by: Ko, Donghyeon, et al.
Published: (2025)
by: Ko, Donghyeon, et al.
Published: (2025)
KMMLU: Measuring Massive Multitask Language Understanding in Korean
by: Son, Guijin, et al.
Published: (2024)
by: Son, Guijin, et al.
Published: (2024)
Assessing Socio-Cultural Alignment and Technical Safety of Sovereign LLMs
by: Chae, Kyubyung, et al.
Published: (2025)
by: Chae, Kyubyung, et al.
Published: (2025)
Guaranteed Generation from Large Language Models
by: Kim, Minbeom, et al.
Published: (2024)
by: Kim, Minbeom, et al.
Published: (2024)
Similar Items
-
KorMedMCQA: Multi-Choice Question Answering Benchmark for Korean Healthcare Professional Licensing Examinations
by: Kweon, Sunjun, et al.
Published: (2024) -
KoBBQ: Korean Bias Benchmark for Question Answering
by: Jin, Jiho, et al.
Published: (2023) -
Alignment Data Map for Efficient Preference Data Selection and Diagnosis
by: Lee, Seohyeong, et al.
Published: (2025) -
Trans-EnV: A Framework for Evaluating the Linguistic Robustness of LLMs Against English Varieties
by: Lee, Jiyoung, et al.
Published: (2025) -
KorMedMCQA-V: A Multimodal Benchmark for Evaluating Vision-Language Models on the Korean Medical Licensing Examination
by: Choi, Byungjin, et al.
Published: (2026)