KoSimpleQA: A Korean Factuality Benchmark with an Analysis of Reasoning LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Ko, Donghyeon, Jin, Yeguk, Chae, Kyubyung, Lee, Byungwook, Jo, Chansong, In, Sookyo, Lee, Jaehong, Kim, Taesup, Kwak, Donghyun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Threat to Tool: Leveraging Refusal-Aware Injection Attacks for Safety Alignment
by: Chae, Kyubyung, et al.
Published: (2025)
by: Chae, Kyubyung, et al.
Published: (2025)
What Models Know, How Well They Know It: Knowledge-Weighted Fine-Tuning for Learning When to Say "I Don't Know"
by: Lee, Joosung, et al.
Published: (2026)
by: Lee, Joosung, et al.
Published: (2026)
Mitigating Hallucination in Abstractive Summarization with Domain-Conditional Mutual Information
by: Chae, Kyubyung, et al.
Published: (2024)
by: Chae, Kyubyung, et al.
Published: (2024)
Assessing Socio-Cultural Alignment and Technical Safety of Sovereign LLMs
by: Chae, Kyubyung, et al.
Published: (2025)
by: Chae, Kyubyung, et al.
Published: (2025)
ReGUIDE: Data Efficient GUI Grounding via Spatial Reasoning and Search
by: Lee, Hyunseok, et al.
Published: (2025)
by: Lee, Hyunseok, et al.
Published: (2025)
What If TSF: A Benchmark for Reframing Forecasting as Scenario-Guided Multimodal Forecasting
by: Jang, Jinkwan, et al.
Published: (2026)
by: Jang, Jinkwan, et al.
Published: (2026)
Model-based Preference Optimization in Abstractive Summarization without Human Feedback
by: Choi, Jaepill, et al.
Published: (2024)
by: Choi, Jaepill, et al.
Published: (2024)
KoBBQ: Korean Bias Benchmark for Question Answering
by: Jin, Jiho, et al.
Published: (2023)
by: Jin, Jiho, et al.
Published: (2023)
Beyond Case Law: Evaluating Structure-Aware Retrieval and Safety in Statute-Centric Legal QA
by: Chae, Kyubyung, et al.
Published: (2026)
by: Chae, Kyubyung, et al.
Published: (2026)
KoBALT: Korean Benchmark For Advanced Linguistic Tasks
by: Shin, Hyopil, et al.
Published: (2025)
by: Shin, Hyopil, et al.
Published: (2025)
When Vision Models Meet Parameter Efficient Look-Aside Adapters Without Large-Scale Audio Pretraining
by: Yeo, Juan, et al.
Published: (2024)
by: Yeo, Juan, et al.
Published: (2024)
Open Ko-LLM Leaderboard: Evaluating Large Language Models in Korean with Ko-H5 Benchmark
by: Park, Chanjun, et al.
Published: (2024)
by: Park, Chanjun, et al.
Published: (2024)
Ko-MuSR: A Multistep Soft Reasoning Benchmark for LLMs Capable of Understanding Korean
by: Park, Chanwoo, et al.
Published: (2025)
by: Park, Chanwoo, et al.
Published: (2025)
Thunder-KoNUBench: A Corpus-Aligned Benchmark for Korean Negation Understanding
by: Jung, Sungmok, et al.
Published: (2026)
by: Jung, Sungmok, et al.
Published: (2026)
Robust Domain Generalization under Divergent Marginal and Conditional Distributions
by: Yeom, Jewon, et al.
Published: (2026)
by: Yeom, Jewon, et al.
Published: (2026)
SDS KoPub VDR: A Benchmark Dataset for Visual Document Retrieval in Korean Public Documents
by: Lee, Jaehoon, et al.
Published: (2025)
by: Lee, Jaehoon, et al.
Published: (2025)
Open Ko-LLM Leaderboard2: Bridging Foundational and Practical Evaluation for Korean LLMs
by: Kim, Hyeonwoo, et al.
Published: (2024)
by: Kim, Hyeonwoo, et al.
Published: (2024)
KoCoNovel: Annotated Dataset of Character Coreference in Korean Novels
by: Kim, Kyuhee, et al.
Published: (2024)
by: Kim, Kyuhee, et al.
Published: (2024)
X-PEFT: eXtremely Parameter-Efficient Fine-Tuning for Extreme Multi-Profile Scenarios
by: Kwak, Namju, et al.
Published: (2024)
by: Kwak, Namju, et al.
Published: (2024)
KoDialogBench: Evaluating Conversational Understanding of Language Models with Korean Dialogue Benchmark
by: Jang, Seongbo, et al.
Published: (2024)
by: Jang, Seongbo, et al.
Published: (2024)
KoCoSa: Korean Context-aware Sarcasm Detection Dataset
by: Kim, Yumin, et al.
Published: (2024)
by: Kim, Yumin, et al.
Published: (2024)
Factuality or Fiction? Benchmarking Modern LLMs on Ambiguous QA with Citations
by: Patel, Maya, et al.
Published: (2024)
by: Patel, Maya, et al.
Published: (2024)
SimpleQA Verified: A Reliable Factuality Benchmark to Measure Parametric Knowledge
by: Haas, Lukas, et al.
Published: (2025)
by: Haas, Lukas, et al.
Published: (2025)
Retaining and Enhancing Pre-trained Knowledge in Vision-Language Models with Prompt Ensembling
by: Kim, Donggeun, et al.
Published: (2024)
by: Kim, Donggeun, et al.
Published: (2024)
Ko-PIQA: A Korean Physical Commonsense Reasoning Dataset with Cultural Context
by: Choi, Dasol, et al.
Published: (2025)
by: Choi, Dasol, et al.
Published: (2025)
Memory-Free Continual Learning with Null Space Adaptation for Zero-Shot Vision-Language Models
by: Jo, Yujin, et al.
Published: (2025)
by: Jo, Yujin, et al.
Published: (2025)
Dynamical Billiard and a long-time behavior of the Boltzmann equation in general 3D toroidal domains
by: Ko, Gyounghun, et al.
Published: (2023)
by: Ko, Gyounghun, et al.
Published: (2023)
KoTaP: A Panel Dataset for Corporate Tax Avoidance, Performance, and Governance in Korea
by: Na, Hyungjong, et al.
Published: (2025)
by: Na, Hyungjong, et al.
Published: (2025)
Nunchi-Bench: Benchmarking Language Models on Cultural Reasoning with a Focus on Korean Superstition
by: Kim, Kyuhee, et al.
Published: (2025)
by: Kim, Kyuhee, et al.
Published: (2025)
Gated Low-rank Adaptation for personalized Code-Switching Automatic Speech Recognition on the low-spec devices
by: Kim, Gwantae, et al.
Published: (2024)
by: Kim, Gwantae, et al.
Published: (2024)
Making Sense of Korean Sentences: A Comprehensive Evaluation of LLMs through KoSEnd Dataset
by: Yu, Seunguk, et al.
Published: (2025)
by: Yu, Seunguk, et al.
Published: (2025)
KoBLEX: Open Legal Question Answering with Multi-hop Reasoning
by: Lee, Jihyung, et al.
Published: (2025)
by: Lee, Jihyung, et al.
Published: (2025)
TriBench-Ko: Evaluating LLM Risks in Judicial Workflows
by: Lee, Haesung, et al.
Published: (2026)
by: Lee, Haesung, et al.
Published: (2026)
ExpertGenQA: Open-ended QA generation in Specialized Domains
by: Shahgir, Haz Sameen, et al.
Published: (2025)
by: Shahgir, Haz Sameen, et al.
Published: (2025)
Factors Associated With Health‐Promoting Behaviors Among South Korean Adults: A Cross‐Sectional Study
by: Yoonjung Kim, et al.
Published: (2024)
by: Yoonjung Kim, et al.
Published: (2024)
CodeSimpleQA: Scaling Factuality in Code Large Language Models
by: Yang, Jian, et al.
Published: (2025)
by: Yang, Jian, et al.
Published: (2025)
CARBD-Ko: A Contextually Annotated Review Benchmark Dataset for Aspect-Level Sentiment Classification in Korean
by: Jang, Dongjun, et al.
Published: (2024)
by: Jang, Dongjun, et al.
Published: (2024)
KMMMU: Evaluation of Massive Multi-discipline Multimodal Understanding in Korean Language and Context
by: Lee, Nahyun, et al.
Published: (2026)
by: Lee, Nahyun, et al.
Published: (2026)
BankMathBench: A Benchmark for Numerical Reasoning in Banking Scenarios
by: Lee, Yunseung, et al.
Published: (2026)
by: Lee, Yunseung, et al.
Published: (2026)
Video SimpleQA: Towards Factuality Evaluation in Large Video Language Models
by: Cao, Meng, et al.
Published: (2025)
by: Cao, Meng, et al.
Published: (2025)
Similar Items
-
From Threat to Tool: Leveraging Refusal-Aware Injection Attacks for Safety Alignment
by: Chae, Kyubyung, et al.
Published: (2025) -
What Models Know, How Well They Know It: Knowledge-Weighted Fine-Tuning for Learning When to Say "I Don't Know"
by: Lee, Joosung, et al.
Published: (2026) -
Mitigating Hallucination in Abstractive Summarization with Domain-Conditional Mutual Information
by: Chae, Kyubyung, et al.
Published: (2024) -
Assessing Socio-Cultural Alignment and Technical Safety of Sovereign LLMs
by: Chae, Kyubyung, et al.
Published: (2025) -
ReGUIDE: Data Efficient GUI Grounding via Spatial Reasoning and Search
by: Lee, Hyunseok, et al.
Published: (2025)