Ko-MuSR: A Multistep Soft Reasoning Benchmark for LLMs Capable of Understanding Korean
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Park, Chanwoo, Park, Suyoung, Kang, JiA, Park, Jongyeon, Kim, Sangho, Park, Hyunji M., Bae, Sumin, Kang, Mingyu, Lee, Jaejin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Beyond Line-Level Filtering for the Pretraining Corpora of LLMs
von: Park, Chanwoo, et al.
Veröffentlicht: (2025)
von: Park, Chanwoo, et al.
Veröffentlicht: (2025)
MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning
von: Sprague, Zayne, et al.
Veröffentlicht: (2023)
von: Sprague, Zayne, et al.
Veröffentlicht: (2023)
Thunder-NUBench: A Benchmark for LLMs' Sentence-Level Negation Understanding
von: So, Yeonkyoung, et al.
Veröffentlicht: (2025)
von: So, Yeonkyoung, et al.
Veröffentlicht: (2025)
Thunder-KoNUBench: A Corpus-Aligned Benchmark for Korean Negation Understanding
von: Jung, Sungmok, et al.
Veröffentlicht: (2026)
von: Jung, Sungmok, et al.
Veröffentlicht: (2026)
Thunder-LLM: Efficiently Adapting LLMs to Korean with Minimal Resources
von: Kim, Jinpyo, et al.
Veröffentlicht: (2025)
von: Kim, Jinpyo, et al.
Veröffentlicht: (2025)
Models Know Models Best: Evaluation via Model-Preferred Formats
von: Lee, Joonhak, et al.
Veröffentlicht: (2026)
von: Lee, Joonhak, et al.
Veröffentlicht: (2026)
Thunder-DeID: Accurate and Efficient De-identification Framework for Korean Court Judgments
von: Hahm, Sungeun, et al.
Veröffentlicht: (2025)
von: Hahm, Sungeun, et al.
Veröffentlicht: (2025)
88‐4: Distinguished Student Paper: Enhancing Face Recognition Accuracy for Under‐Display Cameras via Image Restoration
von: Kyusu Ahn, et al.
Veröffentlicht: (2025)
von: Kyusu Ahn, et al.
Veröffentlicht: (2025)
Comprehensive Asset Pricing Tests in the Korean Stock Market
von: Jaewan Bae, et al.
Veröffentlicht: (2024)
von: Jaewan Bae, et al.
Veröffentlicht: (2024)
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models
von: Cho, Gyeongje, et al.
Veröffentlicht: (2025)
von: Cho, Gyeongje, et al.
Veröffentlicht: (2025)
An Effective Energy Mask-based Adversarial Evasion Attacks against Misclassification in Speaker Recognition Systems
von: Park, Chanwoo, et al.
Veröffentlicht: (2026)
von: Park, Chanwoo, et al.
Veröffentlicht: (2026)
How Many Experts Are Enough? Towards Optimal Semantic Specialization for Mixture-of-Experts
von: Park, Sumin, et al.
Veröffentlicht: (2025)
von: Park, Sumin, et al.
Veröffentlicht: (2025)
Open Ko-LLM Leaderboard: Evaluating Large Language Models in Korean with Ko-H5 Benchmark
von: Park, Chanjun, et al.
Veröffentlicht: (2024)
von: Park, Chanjun, et al.
Veröffentlicht: (2024)
Integrating Spatial and Frequency Information for Under-Display Camera Image Restoration
von: Ahn, Kyusu, et al.
Veröffentlicht: (2025)
von: Ahn, Kyusu, et al.
Veröffentlicht: (2025)
Are Graph Transformers Necessary? Efficient Long-Range Message Passing with Fractal Nodes in MPNNs
von: Choi, Jeongwhan, et al.
Veröffentlicht: (2025)
von: Choi, Jeongwhan, et al.
Veröffentlicht: (2025)
Open Ko-LLM Leaderboard2: Bridging Foundational and Practical Evaluation for Korean LLMs
von: Kim, Hyeonwoo, et al.
Veröffentlicht: (2024)
von: Kim, Hyeonwoo, et al.
Veröffentlicht: (2024)
The Korean integrative family therapy model
von: Tai‐Young Park, et al.
Veröffentlicht: (2025)
von: Tai‐Young Park, et al.
Veröffentlicht: (2025)
ToxReason: A Benchmark for Mechanistic Chemical Toxicity Reasoning via Adverse Outcome Pathway
von: Park, Jueon, et al.
Veröffentlicht: (2026)
von: Park, Jueon, et al.
Veröffentlicht: (2026)
Perspectives on Technology Transfer Strategies of Korean Companies in Point of Resource and Capability Based View
von: Seung-Ho Park
Veröffentlicht: (2011)
von: Seung-Ho Park
Veröffentlicht: (2011)
Constituency Structure over Eojeol in Korean Treebanks
von: Park, Jungyeul, et al.
Veröffentlicht: (2025)
von: Park, Jungyeul, et al.
Veröffentlicht: (2025)
SDS KoPub VDR: A Benchmark Dataset for Visual Document Retrieval in Korean Public Documents
von: Lee, Jaehoon, et al.
Veröffentlicht: (2025)
von: Lee, Jaehoon, et al.
Veröffentlicht: (2025)
PsyProbe: Proactive and Interpretable Dialogue through User State Modeling for Exploratory Counseling
von: Park, Sohhyung, et al.
Veröffentlicht: (2026)
von: Park, Sohhyung, et al.
Veröffentlicht: (2026)
Latent Profiles of Parenting Styles and Internalising and Externalising Behaviour Problems and Smartphone Dependency Among Children in Foster Care
von: Hyunji Lee, et al.
Veröffentlicht: (2026)
von: Hyunji Lee, et al.
Veröffentlicht: (2026)
Understanding LLM Development Through Longitudinal Study: Insights from the Open Ko-LLM Leaderboard
von: Park, Chanjun, et al.
Veröffentlicht: (2024)
von: Park, Chanjun, et al.
Veröffentlicht: (2024)
Are they lovers or friends? Evaluating LLMs' Social Reasoning in English and Korean Dialogues
von: Kim, Eunsu, et al.
Veröffentlicht: (2025)
von: Kim, Eunsu, et al.
Veröffentlicht: (2025)
KITE: A Benchmark for Evaluating Korean Instruction-Following Abilities in Large Language Models
von: Kim, Dongjun, et al.
Veröffentlicht: (2025)
von: Kim, Dongjun, et al.
Veröffentlicht: (2025)
KoSimpleQA: A Korean Factuality Benchmark with an Analysis of Reasoning LLMs
von: Ko, Donghyeon, et al.
Veröffentlicht: (2025)
von: Ko, Donghyeon, et al.
Veröffentlicht: (2025)
Evaluating Large language models on Understanding Korean indirect Speech acts
von: Koo, Youngeun, et al.
Veröffentlicht: (2025)
von: Koo, Youngeun, et al.
Veröffentlicht: (2025)
PANDA: Expanded Width-Aware Message Passing Beyond Rewiring
von: Choi, Jeongwhan, et al.
Veröffentlicht: (2024)
von: Choi, Jeongwhan, et al.
Veröffentlicht: (2024)
CaGR-RAG: Context-aware Query Grouping for Disk-based Vector Search in RAG Systems
von: Jeong, Yeonwoo, et al.
Veröffentlicht: (2025)
von: Jeong, Yeonwoo, et al.
Veröffentlicht: (2025)
The complete mitogenome of the Korean greater tube-nosed bat, Murina leucogaster (Chiroptera: Vespertilionidae)
von: Yoon, Gwang Bae, et al.
Veröffentlicht: (2016)
von: Yoon, Gwang Bae, et al.
Veröffentlicht: (2016)
LCIRC: A Recurrent Compression Approach for Efficient Long-form Context and Query Dependent Modeling in LLMs
von: An, Sumin, et al.
Veröffentlicht: (2025)
von: An, Sumin, et al.
Veröffentlicht: (2025)
Reasoning-Based Approach with Chain-of-Thought for Alzheimer's Detection Using Speech and Large Language Models
von: Park, Chanwoo, et al.
Veröffentlicht: (2025)
von: Park, Chanwoo, et al.
Veröffentlicht: (2025)
Exploring the genetic and morphological diversity of Pantala flavescens across different south Korean river basins
von: Da Som Park, et al.
Veröffentlicht: (2025)
von: Da Som Park, et al.
Veröffentlicht: (2025)
Analytical Investigation of Temperature‐Dependent Characteristics in 4.5 kV SiC Super‐Junction Devices
von: Sumin Park, et al.
Veröffentlicht: (2026)
von: Sumin Park, et al.
Veröffentlicht: (2026)
VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis
von: Park, Jinho, et al.
Veröffentlicht: (2026)
von: Park, Jinho, et al.
Veröffentlicht: (2026)
Optimal First-Order Algorithms as a Function of Inequalities
von: Park, Chanwoo, et al.
Veröffentlicht: (2021)
von: Park, Chanwoo, et al.
Veröffentlicht: (2021)
Understanding South Korean Immigrant Adolescents' Experiences of Staying Home Alone
von: Sol Park, et al.
Veröffentlicht: (2025)
von: Sol Park, et al.
Veröffentlicht: (2025)
Capture of Sulfur Dioxide in Ship Exhaust Gas by Tertiary Amine Absorbents
von: Kwanghwi Kim, et al.
Veröffentlicht: (2024)
von: Kwanghwi Kim, et al.
Veröffentlicht: (2024)
KFinEval-Pilot: A Comprehensive Benchmark Suite for Korean Financial Language Understanding
von: Hwang, Bokwang, et al.
Veröffentlicht: (2025)
von: Hwang, Bokwang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Beyond Line-Level Filtering for the Pretraining Corpora of LLMs
von: Park, Chanwoo, et al.
Veröffentlicht: (2025) -
MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning
von: Sprague, Zayne, et al.
Veröffentlicht: (2023) -
Thunder-NUBench: A Benchmark for LLMs' Sentence-Level Negation Understanding
von: So, Yeonkyoung, et al.
Veröffentlicht: (2025) -
Thunder-KoNUBench: A Corpus-Aligned Benchmark for Korean Negation Understanding
von: Jung, Sungmok, et al.
Veröffentlicht: (2026) -
Thunder-LLM: Efficiently Adapting LLMs to Korean with Minimal Resources
von: Kim, Jinpyo, et al.
Veröffentlicht: (2025)