Redefining Evaluation Standards: A Unified Framework for Evaluating the Korean Capabilities of Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lee, Hanwool, Choi, Dasol, Kim, Sooyong, Jeong, Ilgyun, Baek, Sangwon, Son, Guijin, Hwang, Inseon, Lee, Naeun, Hong, Seunghyeok |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multi-Task Inference: Can Large Language Models Follow Multiple Instructions at Once?
von: Son, Guijin, et al.
Veröffentlicht: (2024)
von: Son, Guijin, et al.
Veröffentlicht: (2024)
Multi-Step Reasoning in Korean and the Emergent Mirage
von: Son, Guijin, et al.
Veröffentlicht: (2025)
von: Son, Guijin, et al.
Veröffentlicht: (2025)
IVE: Enhanced Probabilistic Forecasting of Intraday Volume Ratio with Transformers
von: Lee, Hanwool, et al.
Veröffentlicht: (2024)
von: Lee, Hanwool, et al.
Veröffentlicht: (2024)
Ko-PIQA: A Korean Physical Commonsense Reasoning Dataset with Cultural Context
von: Choi, Dasol, et al.
Veröffentlicht: (2025)
von: Choi, Dasol, et al.
Veröffentlicht: (2025)
(G)I-DLE: Generative Inference via Distribution-preserving Logit Exclusion with KL Divergence Minimization for Constrained Decoding
von: Lee, Hanwool
Veröffentlicht: (2025)
von: Lee, Hanwool
Veröffentlicht: (2025)
TWICE: What Advantages Can Low-Resource Domain-Specific Embedding Model Bring? -- A Case Study on Korea Financial Texts
von: Hwang, Yewon, et al.
Veröffentlicht: (2025)
von: Hwang, Yewon, et al.
Veröffentlicht: (2025)
Improving Fine-grained Visual Understanding in VLMs through Text-Only Training
von: Choi, Dasol, et al.
Veröffentlicht: (2024)
von: Choi, Dasol, et al.
Veröffentlicht: (2024)
What Users Leave Unsaid: Under-Specified Queries Limit Vision-Language Models
von: Choi, Dasol, et al.
Veröffentlicht: (2026)
von: Choi, Dasol, et al.
Veröffentlicht: (2026)
NMIXX: Domain-Adapted Neural Embeddings for Cross-Lingual eXploration of Finance
von: Lee, Hanwool, et al.
Veröffentlicht: (2025)
von: Lee, Hanwool, et al.
Veröffentlicht: (2025)
HAE-RAE Bench: Evaluation of Korean Knowledge in Language Models
von: Son, Guijin, et al.
Veröffentlicht: (2023)
von: Son, Guijin, et al.
Veröffentlicht: (2023)
Understand, Solve and Translate: Bridging the Multilingual Mathematical Reasoning Gap
von: Ko, Hyunwoo, et al.
Veröffentlicht: (2025)
von: Ko, Hyunwoo, et al.
Veröffentlicht: (2025)
Beyond Sentiment Classification: A Generative Framework for Emotion Intensity Evaluation in Text
von: Fabozzi, Francesco A., et al.
Veröffentlicht: (2026)
von: Fabozzi, Francesco A., et al.
Veröffentlicht: (2026)
LLM-as-a-Judge & Reward Model: What They Can and Cannot Do
von: Son, Guijin, et al.
Veröffentlicht: (2024)
von: Son, Guijin, et al.
Veröffentlicht: (2024)
Pushing the Boundaries of Multiple Choice Evaluation to One Hundred Options
von: Lee, Nahyun, et al.
Veröffentlicht: (2026)
von: Lee, Nahyun, et al.
Veröffentlicht: (2026)
From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation
von: Hong, Seokhee, et al.
Veröffentlicht: (2025)
von: Hong, Seokhee, et al.
Veröffentlicht: (2025)
AI PB: A Grounded Generative Agent for Personalized Investment Insights
von: Park, Daewoo, et al.
Veröffentlicht: (2025)
von: Park, Daewoo, et al.
Veröffentlicht: (2025)
KMMLU: Measuring Massive Multitask Language Understanding in Korean
von: Son, Guijin, et al.
Veröffentlicht: (2024)
von: Son, Guijin, et al.
Veröffentlicht: (2024)
Evaluating LLMs in Finance Requires Explicit Bias Consideration
von: Kong, Yaxuan, et al.
Veröffentlicht: (2026)
von: Kong, Yaxuan, et al.
Veröffentlicht: (2026)
Won: Establishing Best Practices for Korean Financial NLP
von: Son, Guijin, et al.
Veröffentlicht: (2025)
von: Son, Guijin, et al.
Veröffentlicht: (2025)
KAIO: A Collection of More Challenging Korean Questions
von: Lee, Nahyun, et al.
Veröffentlicht: (2025)
von: Lee, Nahyun, et al.
Veröffentlicht: (2025)
Transfer Learning-Based Surrogate Modeling for Nonlinear Time-History Response Analysis of High-Fidelity Structural Models
von: Ishikawa, Keiichi, et al.
Veröffentlicht: (2025)
von: Ishikawa, Keiichi, et al.
Veröffentlicht: (2025)
Words That Unite The World: A Unified Framework for Deciphering Central Bank Communications Globally
von: Shah, Agam, et al.
Veröffentlicht: (2025)
von: Shah, Agam, et al.
Veröffentlicht: (2025)
An Explainable Approach to Document-level Translation Evaluation with Topic Modeling
von: Lee, Hyeokmin, et al.
Veröffentlicht: (2026)
von: Lee, Hyeokmin, et al.
Veröffentlicht: (2026)
EducationQ: Evaluating LLMs' Teaching Capabilities Through Multi-Agent Dialogue Framework
von: Shi, Yao, et al.
Veröffentlicht: (2025)
von: Shi, Yao, et al.
Veröffentlicht: (2025)
Evaluation of Thermal Control Based on Spatial Thermal Comfort with Reconstructed Environmental Data
von: Kim, Youngkyu, et al.
Veröffentlicht: (2025)
von: Kim, Youngkyu, et al.
Veröffentlicht: (2025)
RAISE: A Unified Framework for Responsible AI Scoring and Evaluation
von: Nguyen, Loc Phuc Truong, et al.
Veröffentlicht: (2025)
von: Nguyen, Loc Phuc Truong, et al.
Veröffentlicht: (2025)
Decision by Supervised Learning with Deep Ensembles: A Practical Framework for Robust Portfolio Optimization
von: Kim, Juhyeong, et al.
Veröffentlicht: (2025)
von: Kim, Juhyeong, et al.
Veröffentlicht: (2025)
Temporal Representation Learning for Stock Similarities and Its Applications in Investment Management
von: Hwang, Yoontae, et al.
Veröffentlicht: (2024)
von: Hwang, Yoontae, et al.
Veröffentlicht: (2024)
Counterfactual Voting Adjustment for Quality Assessment and Fairer Voting in Online Platforms with Helpfulness Evaluation
von: Liu, Chang, et al.
Veröffentlicht: (2025)
von: Liu, Chang, et al.
Veröffentlicht: (2025)
SECQUE: A Benchmark for Evaluating Real-World Financial Analysis Capabilities
von: Yoash, Noga Ben, et al.
Veröffentlicht: (2025)
von: Yoash, Noga Ben, et al.
Veröffentlicht: (2025)
KMMMU: Evaluation of Massive Multi-discipline Multimodal Understanding in Korean Language and Context
von: Lee, Nahyun, et al.
Veröffentlicht: (2026)
von: Lee, Nahyun, et al.
Veröffentlicht: (2026)
Quantifying Qualitative Insights: Leveraging LLMs to Market Predict
von: Lee, Hoyoung, et al.
Veröffentlicht: (2024)
von: Lee, Hoyoung, et al.
Veröffentlicht: (2024)
Evaluating the Efficiency and Cost-effectiveness of RPB-based CO2 Capture: A Comprehensive Approach to Simultaneous Design and Operating Condition Optimization
von: Jung, Howoun, et al.
Veröffentlicht: (2024)
von: Jung, Howoun, et al.
Veröffentlicht: (2024)
Redefining Computing: Rise of ARM from consumer to Cloud for energy efficiency
von: Rahman, Tahmid Noor, et al.
Veröffentlicht: (2024)
von: Rahman, Tahmid Noor, et al.
Veröffentlicht: (2024)
Dominant Design Prediction with Phylogenetic Networks
von: He, Youwei, et al.
Veröffentlicht: (2024)
von: He, Youwei, et al.
Veröffentlicht: (2024)
Can MLLMs Reason About Visual Persuasion? Evaluating the Efficacy and Faithfulness of Reasoning
von: Lee, Naeun, et al.
Veröffentlicht: (2026)
von: Lee, Naeun, et al.
Veröffentlicht: (2026)
Toward Black Scholes for Prediction Markets: A Unified Kernel and Market Maker's Handbook
von: Dalen, Shaw
Veröffentlicht: (2025)
von: Dalen, Shaw
Veröffentlicht: (2025)
Common Task Framework For a Critical Evaluation of Scientific Machine Learning Algorithms
von: Wyder, Philippe Martin, et al.
Veröffentlicht: (2025)
von: Wyder, Philippe Martin, et al.
Veröffentlicht: (2025)
Attention-Based Reading, Highlighting, and Forecasting of the Limit Order Book
von: Jung, Jiwon, et al.
Veröffentlicht: (2024)
von: Jung, Jiwon, et al.
Veröffentlicht: (2024)
Can GANs Learn the Stylized Facts of Financial Time Series?
von: Kwon, Sohyeon, et al.
Veröffentlicht: (2024)
von: Kwon, Sohyeon, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Multi-Task Inference: Can Large Language Models Follow Multiple Instructions at Once?
von: Son, Guijin, et al.
Veröffentlicht: (2024) -
Multi-Step Reasoning in Korean and the Emergent Mirage
von: Son, Guijin, et al.
Veröffentlicht: (2025) -
IVE: Enhanced Probabilistic Forecasting of Intraday Volume Ratio with Transformers
von: Lee, Hanwool, et al.
Veröffentlicht: (2024) -
Ko-PIQA: A Korean Physical Commonsense Reasoning Dataset with Cultural Context
von: Choi, Dasol, et al.
Veröffentlicht: (2025) -
(G)I-DLE: Generative Inference via Distribution-preserving Logit Exclusion with KL Divergence Minimization for Constrained Decoding
von: Lee, Hanwool
Veröffentlicht: (2025)