Ko-PIQA: A Korean Physical Commonsense Reasoning Dataset with Cultural Context
Fuente:
arXiv
Saved in:
| Main Authors: | Choi, Dasol, Kim, Jungwhan, Son, Guijin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-Step Reasoning in Korean and the Emergent Mirage
by: Son, Guijin, et al.
Published: (2025)
by: Son, Guijin, et al.
Published: (2025)
Understand, Solve and Translate: Bridging the Multilingual Mathematical Reasoning Gap
by: Ko, Hyunwoo, et al.
Published: (2025)
by: Ko, Hyunwoo, et al.
Published: (2025)
Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures
by: Chang, Tyler A., et al.
Published: (2025)
by: Chang, Tyler A., et al.
Published: (2025)
Sinhala Physical Common Sense Reasoning Dataset for Global PIQA
by: de Silva, Nisansa, et al.
Published: (2026)
by: de Silva, Nisansa, et al.
Published: (2026)
Improving Fine-grained Visual Understanding in VLMs through Text-Only Training
by: Choi, Dasol, et al.
Published: (2024)
by: Choi, Dasol, et al.
Published: (2024)
KoCoSa: Korean Context-aware Sarcasm Detection Dataset
by: Kim, Yumin, et al.
Published: (2024)
by: Kim, Yumin, et al.
Published: (2024)
Redefining Evaluation Standards: A Unified Framework for Evaluating the Korean Capabilities of Language Models
by: Lee, Hanwool, et al.
Published: (2025)
by: Lee, Hanwool, et al.
Published: (2025)
KAIO: A Collection of More Challenging Korean Questions
by: Lee, Nahyun, et al.
Published: (2025)
by: Lee, Nahyun, et al.
Published: (2025)
Won: Establishing Best Practices for Korean Financial NLP
by: Son, Guijin, et al.
Published: (2025)
by: Son, Guijin, et al.
Published: (2025)
KoCoNovel: Annotated Dataset of Character Coreference in Korean Novels
by: Kim, Kyuhee, et al.
Published: (2024)
by: Kim, Kyuhee, et al.
Published: (2024)
Revisiting the UID Hypothesis in LLM Reasoning Traces
by: Gwak, Minju, et al.
Published: (2025)
by: Gwak, Minju, et al.
Published: (2025)
Revisiting the Uniform Information Density Hypothesis in LLM Reasoning
by: Gwak, Minju, et al.
Published: (2025)
by: Gwak, Minju, et al.
Published: (2025)
Making Sense of Korean Sentences: A Comprehensive Evaluation of LLMs through KoSEnd Dataset
by: Yu, Seunguk, et al.
Published: (2025)
by: Yu, Seunguk, et al.
Published: (2025)
KMMMU: Evaluation of Massive Multi-discipline Multimodal Understanding in Korean Language and Context
by: Lee, Nahyun, et al.
Published: (2026)
by: Lee, Nahyun, et al.
Published: (2026)
KMMLU: Measuring Massive Multitask Language Understanding in Korean
by: Son, Guijin, et al.
Published: (2024)
by: Son, Guijin, et al.
Published: (2024)
Everyday Physics in Korean Contexts: A Culturally Grounded Physical Reasoning Benchmark
by: Jeong, Jihae, et al.
Published: (2025)
by: Jeong, Jihae, et al.
Published: (2025)
KoSimpleQA: A Korean Factuality Benchmark with an Analysis of Reasoning LLMs
by: Ko, Donghyeon, et al.
Published: (2025)
by: Ko, Donghyeon, et al.
Published: (2025)
From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation
by: Hong, Seokhee, et al.
Published: (2025)
by: Hong, Seokhee, et al.
Published: (2025)
SDS KoPub VDR: A Benchmark Dataset for Visual Document Retrieval in Korean Public Documents
by: Lee, Jaehoon, et al.
Published: (2025)
by: Lee, Jaehoon, et al.
Published: (2025)
CARBD-Ko: A Contextually Annotated Review Benchmark Dataset for Aspect-Level Sentiment Classification in Korean
by: Jang, Dongjun, et al.
Published: (2024)
by: Jang, Dongjun, et al.
Published: (2024)
Commonsense Reasoning in Arab Culture
by: Sadallah, Abdelrahman, et al.
Published: (2025)
by: Sadallah, Abdelrahman, et al.
Published: (2025)
K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts
by: Lee, Nahyun, et al.
Published: (2026)
by: Lee, Nahyun, et al.
Published: (2026)
KoBALT: Korean Benchmark For Advanced Linguistic Tasks
by: Shin, Hyopil, et al.
Published: (2025)
by: Shin, Hyopil, et al.
Published: (2025)
Open Ko-LLM Leaderboard: Evaluating Large Language Models in Korean with Ko-H5 Benchmark
by: Park, Chanjun, et al.
Published: (2024)
by: Park, Chanjun, et al.
Published: (2024)
HAE-RAE Bench: Evaluation of Korean Knowledge in Language Models
by: Son, Guijin, et al.
Published: (2023)
by: Son, Guijin, et al.
Published: (2023)
Pushing the Boundaries of Multiple Choice Evaluation to One Hundred Options
by: Lee, Nahyun, et al.
Published: (2026)
by: Lee, Nahyun, et al.
Published: (2026)
Pushing on Multilingual Reasoning Models with Language-Mixed Chain-of-Thought
by: Son, Guijin, et al.
Published: (2025)
by: Son, Guijin, et al.
Published: (2025)
No Language Data Left Behind: A Comparative Study of CJK Language Datasets in the Hugging Face Ecosystem
by: Choi, Dasol, et al.
Published: (2025)
by: Choi, Dasol, et al.
Published: (2025)
Ko-MuSR: A Multistep Soft Reasoning Benchmark for LLMs Capable of Understanding Korean
by: Park, Chanwoo, et al.
Published: (2025)
by: Park, Chanwoo, et al.
Published: (2025)
KoBBQ: Korean Bias Benchmark for Question Answering
by: Jin, Jiho, et al.
Published: (2023)
by: Jin, Jiho, et al.
Published: (2023)
Linguistic Generalizability of Test-Time Scaling in Mathematical Reasoning
by: Son, Guijin, et al.
Published: (2025)
by: Son, Guijin, et al.
Published: (2025)
KoACD: The First Korean Adolescent Dataset for Cognitive Distortion Analysis via Role-Switching Multi-LLM Negotiation
by: Kim, JunSeo, et al.
Published: (2025)
by: Kim, JunSeo, et al.
Published: (2025)
Thunder-KoNUBench: A Corpus-Aligned Benchmark for Korean Negation Understanding
by: Jung, Sungmok, et al.
Published: (2026)
by: Jung, Sungmok, et al.
Published: (2026)
What Users Leave Unsaid: Under-Specified Queries Limit Vision-Language Models
by: Choi, Dasol, et al.
Published: (2026)
by: Choi, Dasol, et al.
Published: (2026)
Open Ko-LLM Leaderboard2: Bridging Foundational and Practical Evaluation for Korean LLMs
by: Kim, Hyeonwoo, et al.
Published: (2024)
by: Kim, Hyeonwoo, et al.
Published: (2024)
K-Act2Emo: Korean Commonsense Knowledge Graph for Indirect Emotional Expression
by: Kim, Kyuhee, et al.
Published: (2024)
by: Kim, Kyuhee, et al.
Published: (2024)
CLIcK: A Benchmark Dataset of Cultural and Linguistic Intelligence in Korean
by: Kim, Eunsu, et al.
Published: (2024)
by: Kim, Eunsu, et al.
Published: (2024)
KoGEC : Korean Grammatical Error Correction with Pre-trained Translation Models
by: Kim, Taeeun, et al.
Published: (2025)
by: Kim, Taeeun, et al.
Published: (2025)
KULTURE Bench: A Benchmark for Assessing Language Model in Korean Cultural Context
by: Wang, Xiaonan, et al.
Published: (2024)
by: Wang, Xiaonan, et al.
Published: (2024)
Bias Mitigation or Cultural Commonsense? Evaluating LLMs with a Japanese Dataset
by: Yamamoto, Taisei, et al.
Published: (2025)
by: Yamamoto, Taisei, et al.
Published: (2025)
Similar Items
-
Multi-Step Reasoning in Korean and the Emergent Mirage
by: Son, Guijin, et al.
Published: (2025) -
Understand, Solve and Translate: Bridging the Multilingual Mathematical Reasoning Gap
by: Ko, Hyunwoo, et al.
Published: (2025) -
Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures
by: Chang, Tyler A., et al.
Published: (2025) -
Sinhala Physical Common Sense Reasoning Dataset for Global PIQA
by: de Silva, Nisansa, et al.
Published: (2026) -
Improving Fine-grained Visual Understanding in VLMs through Text-Only Training
by: Choi, Dasol, et al.
Published: (2024)