KoBALT: Korean Benchmark For Advanced Linguistic Tasks
Fuente:
arXiv
Salvato in:
| Autori principali: | Shin, Hyopil, Lee, Sangah, Jang, Dongjun, Song, Wooseok, Kim, Jaeyoon, Oh, Chaeyoung, Jo, Hyemi, Ahn, Youngchae, Oh, Sihyun, Chang, Hyohyeong, Kim, Sunkyoung, Lee, Jinsik |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
KIT-19: A Comprehensive Korean Instruction Toolkit on 19 Tasks for Fine-Tuning Korean Large Language Models
di: Jang, Dongjun, et al.
Pubblicazione: (2024)
di: Jang, Dongjun, et al.
Pubblicazione: (2024)
P-CoT: A Pedagogically-motivated Participatory Chain-of-Thought Prompting for Phonological Reasoning in LLMs
di: Jang, Dongjun, et al.
Pubblicazione: (2025)
di: Jang, Dongjun, et al.
Pubblicazione: (2025)
RCScore: Quantifying Response Consistency in Large Language Models
di: Jang, Dongjun, et al.
Pubblicazione: (2025)
di: Jang, Dongjun, et al.
Pubblicazione: (2025)
Korean Bio-Medical Corpus (KBMC) for Medical Named Entity Recognition
di: Byun, Sungjoo, et al.
Pubblicazione: (2024)
di: Byun, Sungjoo, et al.
Pubblicazione: (2024)
CARBD-Ko: A Contextually Annotated Review Benchmark Dataset for Aspect-Level Sentiment Classification in Korean
di: Jang, Dongjun, et al.
Pubblicazione: (2024)
di: Jang, Dongjun, et al.
Pubblicazione: (2024)
KoCoNovel: Annotated Dataset of Character Coreference in Korean Novels
di: Kim, Kyuhee, et al.
Pubblicazione: (2024)
di: Kim, Kyuhee, et al.
Pubblicazione: (2024)
A Study on How Attention Scores in the BERT Model are Aware of Lexical Categories in Syntactic and Semantic Tasks on the GLUE Benchmark
di: Jang, Dongjun, et al.
Pubblicazione: (2024)
di: Jang, Dongjun, et al.
Pubblicazione: (2024)
From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation
di: Hong, Seokhee, et al.
Pubblicazione: (2025)
di: Hong, Seokhee, et al.
Pubblicazione: (2025)
MoFE: Mixture of Frozen Experts Architecture
di: Seo, Jean, et al.
Pubblicazione: (2025)
di: Seo, Jean, et al.
Pubblicazione: (2025)
Cross-lingual QA: A Key to Unlocking In-context Cross-lingual Performance
di: Kim, Sunkyoung, et al.
Pubblicazione: (2023)
di: Kim, Sunkyoung, et al.
Pubblicazione: (2023)
Nunchi-Bench: Benchmarking Language Models on Cultural Reasoning with a Focus on Korean Superstition
di: Kim, Kyuhee, et al.
Pubblicazione: (2025)
di: Kim, Kyuhee, et al.
Pubblicazione: (2025)
KoBBQ: Korean Bias Benchmark for Question Answering
di: Jin, Jiho, et al.
Pubblicazione: (2023)
di: Jin, Jiho, et al.
Pubblicazione: (2023)
K-Act2Emo: Korean Commonsense Knowledge Graph for Indirect Emotional Expression
di: Kim, Kyuhee, et al.
Pubblicazione: (2024)
di: Kim, Kyuhee, et al.
Pubblicazione: (2024)
How does a Language-Specific Tokenizer affect LLMs?
di: Seo, Jean, et al.
Pubblicazione: (2025)
di: Seo, Jean, et al.
Pubblicazione: (2025)
DAHL: Domain-specific Automated Hallucination Evaluation of Long-Form Text through a Benchmark Dataset in Biomedicine
di: Seo, Jean, et al.
Pubblicazione: (2024)
di: Seo, Jean, et al.
Pubblicazione: (2024)
Absence of the Lavrentiev phenomenon for degenerate parabolic double phase problems
di: Kim, Bogi, et al.
Pubblicazione: (2026)
di: Kim, Bogi, et al.
Pubblicazione: (2026)
Thunder-KoNUBench: A Corpus-Aligned Benchmark for Korean Negation Understanding
di: Jung, Sungmok, et al.
Pubblicazione: (2026)
di: Jung, Sungmok, et al.
Pubblicazione: (2026)
Open Ko-LLM Leaderboard: Evaluating Large Language Models in Korean with Ko-H5 Benchmark
di: Park, Chanjun, et al.
Pubblicazione: (2024)
di: Park, Chanjun, et al.
Pubblicazione: (2024)
2D Materials: From Design and Synthesis to Applications in Electrical and Electrochemical Biosensors
di: Masud, et al.
Pubblicazione: (2025)
di: Masud, et al.
Pubblicazione: (2025)
CLIcK: A Benchmark Dataset of Cultural and Linguistic Intelligence in Korean
di: Kim, Eunsu, et al.
Pubblicazione: (2024)
di: Kim, Eunsu, et al.
Pubblicazione: (2024)
KoSimpleQA: A Korean Factuality Benchmark with an Analysis of Reasoning LLMs
di: Ko, Donghyeon, et al.
Pubblicazione: (2025)
di: Ko, Donghyeon, et al.
Pubblicazione: (2025)
KoCoSa: Korean Context-aware Sarcasm Detection Dataset
di: Kim, Yumin, et al.
Pubblicazione: (2024)
di: Kim, Yumin, et al.
Pubblicazione: (2024)
Open Ko-LLM Leaderboard2: Bridging Foundational and Practical Evaluation for Korean LLMs
di: Kim, Hyeonwoo, et al.
Pubblicazione: (2024)
di: Kim, Hyeonwoo, et al.
Pubblicazione: (2024)
SDS KoPub VDR: A Benchmark Dataset for Visual Document Retrieval in Korean Public Documents
di: Lee, Jaehoon, et al.
Pubblicazione: (2025)
di: Lee, Jaehoon, et al.
Pubblicazione: (2025)
Molecular Mechanism of Ionic Conductivity Enhancement by Heterogeneous Interface in Composite Hydrogel Electrolytes
di: Hongdeok Kim, et al.
Pubblicazione: (2025)
di: Hongdeok Kim, et al.
Pubblicazione: (2025)
Exploring the Spectrum of Visio-Linguistic Compositionality and Recognition
di: Oh, Youngtaek, et al.
Pubblicazione: (2024)
di: Oh, Youngtaek, et al.
Pubblicazione: (2024)
KoDialogBench: Evaluating Conversational Understanding of Language Models with Korean Dialogue Benchmark
di: Jang, Seongbo, et al.
Pubblicazione: (2024)
di: Jang, Seongbo, et al.
Pubblicazione: (2024)
KOFFVQA: An Objectively Evaluated Free-form VQA Benchmark for Large Vision-Language Models in the Korean Language
di: Kim, Yoonshik, et al.
Pubblicazione: (2025)
di: Kim, Yoonshik, et al.
Pubblicazione: (2025)
GECKO: Generative Language Model for English, Code and Korean
di: Oh, Sungwoo, et al.
Pubblicazione: (2024)
di: Oh, Sungwoo, et al.
Pubblicazione: (2024)
KoALa-Bench: Evaluating Large Audio Language Models on Korean Speech Understanding and Faithfulness
di: Kim, Jinyoung, et al.
Pubblicazione: (2026)
di: Kim, Jinyoung, et al.
Pubblicazione: (2026)
Mergen: The First Manchu-Korean Machine Translation Model Trained on Augmented Data
di: Seo, Jean, et al.
Pubblicazione: (2023)
di: Seo, Jean, et al.
Pubblicazione: (2023)
Association Between Meat Intake and Metabolic Dysfunction‐Associated Steatotic Liver Disease Incidence in a Korean Population From the Health Examinees Study
di: Uyangamaa Nyamsuren, et al.
Pubblicazione: (2025)
di: Uyangamaa Nyamsuren, et al.
Pubblicazione: (2025)
Do LLMs Need Inherent Reasoning Before Reinforcement Learning? A Study in Korean Self-Correction
di: Kim, Hongjin, et al.
Pubblicazione: (2026)
di: Kim, Hongjin, et al.
Pubblicazione: (2026)
Tidiness Score-Guided Monte Carlo Tree Search for Visual Tabletop Rearrangement
di: Kee, Hogun, et al.
Pubblicazione: (2025)
di: Kee, Hogun, et al.
Pubblicazione: (2025)
Multi-View Attention Multiple-Instance Learning Enhanced by LLM Reasoning for Cognitive Distortion Detection
di: Kim, Jun Seo, et al.
Pubblicazione: (2025)
di: Kim, Jun Seo, et al.
Pubblicazione: (2025)
Advanced Fabrication of Ultrathin Ruthenium Films Using Synergistic Atomic Layer Deposition and Etching
di: Jeongbin Lee, et al.
Pubblicazione: (2025)
di: Jeongbin Lee, et al.
Pubblicazione: (2025)
Skin‐Adhesive Conductive Hydrogel With Potential for Flexible Wearable Sensor Associated With Sarcopenia
di: Hye‐Won Kim, et al.
Pubblicazione: (2025)
di: Hye‐Won Kim, et al.
Pubblicazione: (2025)
LookaheadKV: Fast and Accurate KV Cache Eviction by Glimpsing into the Future without Generation
di: Ahn, Jinwoo, et al.
Pubblicazione: (2026)
di: Ahn, Jinwoo, et al.
Pubblicazione: (2026)
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information
di: Park, Seungcheol, et al.
Pubblicazione: (2025)
di: Park, Seungcheol, et al.
Pubblicazione: (2025)
Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean
di: Kim, SungHo, et al.
Pubblicazione: (2025)
di: Kim, SungHo, et al.
Pubblicazione: (2025)
Documenti analoghi
-
KIT-19: A Comprehensive Korean Instruction Toolkit on 19 Tasks for Fine-Tuning Korean Large Language Models
di: Jang, Dongjun, et al.
Pubblicazione: (2024) -
P-CoT: A Pedagogically-motivated Participatory Chain-of-Thought Prompting for Phonological Reasoning in LLMs
di: Jang, Dongjun, et al.
Pubblicazione: (2025) -
RCScore: Quantifying Response Consistency in Large Language Models
di: Jang, Dongjun, et al.
Pubblicazione: (2025) -
Korean Bio-Medical Corpus (KBMC) for Medical Named Entity Recognition
di: Byun, Sungjoo, et al.
Pubblicazione: (2024) -
CARBD-Ko: A Contextually Annotated Review Benchmark Dataset for Aspect-Level Sentiment Classification in Korean
di: Jang, Dongjun, et al.
Pubblicazione: (2024)