Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean
Fuente:
arXiv
Salvato in:
| Autori principali: | Kim, SungHo, Kim, Nayeon, Jeon, Taehee, Lee, SangKeun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
KOMBO: Korean Character Representations Based on the Combination Rules of Subcharacters
di: Kim, SungHo, et al.
Pubblicazione: (2026)
di: Kim, SungHo, et al.
Pubblicazione: (2026)
SCRIPT: A Subcharacter Compositional Representation Injection Module for Korean Pre-Trained Language Models
di: Kim, SungHo, et al.
Pubblicazione: (2026)
di: Kim, SungHo, et al.
Pubblicazione: (2026)
Incorporating Domain Knowledge into Materials Tokenization
di: Oh, Yerim, et al.
Pubblicazione: (2025)
di: Oh, Yerim, et al.
Pubblicazione: (2025)
Handling Korean Out-of-Vocabulary Words with Phoneme Representation Learning
di: Kim, Nayeon, et al.
Pubblicazione: (2025)
di: Kim, Nayeon, et al.
Pubblicazione: (2025)
Mentor-KD: Making Small Language Models Better Multi-step Reasoners
di: Lee, Hojae, et al.
Pubblicazione: (2024)
di: Lee, Hojae, et al.
Pubblicazione: (2024)
Zero-shot Commonsense Reasoning over Machine Imagination
di: Park, Hyuntae, et al.
Pubblicazione: (2024)
di: Park, Hyuntae, et al.
Pubblicazione: (2024)
MELT: Materials-aware Continued Pre-training for Language Model Adaptation to Materials Science
di: Kim, Junho, et al.
Pubblicazione: (2024)
di: Kim, Junho, et al.
Pubblicazione: (2024)
KoBBQ: Korean Bias Benchmark for Question Answering
di: Jin, Jiho, et al.
Pubblicazione: (2023)
di: Jin, Jiho, et al.
Pubblicazione: (2023)
CleaR: Towards Robust and Generalized Parameter-Efficient Fine-Tuning for Noisy Label Learning
di: Kim, Yeachan, et al.
Pubblicazione: (2024)
di: Kim, Yeachan, et al.
Pubblicazione: (2024)
Enhancing Zero-shot Commonsense Reasoning by Integrating Visual Knowledge via Machine Imagination
di: Park, Hyuntae, et al.
Pubblicazione: (2026)
di: Park, Hyuntae, et al.
Pubblicazione: (2026)
Open Ko-LLM Leaderboard2: Bridging Foundational and Practical Evaluation for Korean LLMs
di: Kim, Hyeonwoo, et al.
Pubblicazione: (2024)
di: Kim, Hyeonwoo, et al.
Pubblicazione: (2024)
Bridging the Gap Between Molecule and Textual Descriptions via Substructure-aware Alignment
di: Park, Hyuntae, et al.
Pubblicazione: (2025)
di: Park, Hyuntae, et al.
Pubblicazione: (2025)
KatFishNet: Detecting LLM-Generated Korean Text through Linguistic Feature Analysis
di: Park, Shinwoo, et al.
Pubblicazione: (2025)
di: Park, Shinwoo, et al.
Pubblicazione: (2025)
SLM-Based Agentic AI with P-C-G: Optimized for Korean Tool Use
di: Jeon, Changhyun, et al.
Pubblicazione: (2025)
di: Jeon, Changhyun, et al.
Pubblicazione: (2025)
Enhanced Facet Generation with LLM Editing
di: Lee, Joosung, et al.
Pubblicazione: (2024)
di: Lee, Joosung, et al.
Pubblicazione: (2024)
Trans-EnV: A Framework for Evaluating the Linguistic Robustness of LLMs Against English Varieties
di: Lee, Jiyoung, et al.
Pubblicazione: (2025)
di: Lee, Jiyoung, et al.
Pubblicazione: (2025)
Multi-Facet Blending for Faceted Query-by-Example Retrieval
di: Do, Heejin, et al.
Pubblicazione: (2024)
di: Do, Heejin, et al.
Pubblicazione: (2024)
DIVE: Towards Descriptive and Diverse Visual Commonsense Generation
di: Park, Jun-Hyung, et al.
Pubblicazione: (2024)
di: Park, Jun-Hyung, et al.
Pubblicazione: (2024)
Do LLMs Need Inherent Reasoning Before Reinforcement Learning? A Study in Korean Self-Correction
di: Kim, Hongjin, et al.
Pubblicazione: (2026)
di: Kim, Hongjin, et al.
Pubblicazione: (2026)
Expanding Foundational Language Capabilities in Open-Source LLMs through a Korean Case Study
di: Lim, Junghwan, et al.
Pubblicazione: (2025)
di: Lim, Junghwan, et al.
Pubblicazione: (2025)
Open Ko-LLM Leaderboard: Evaluating Large Language Models in Korean with Ko-H5 Benchmark
di: Park, Chanjun, et al.
Pubblicazione: (2024)
di: Park, Chanjun, et al.
Pubblicazione: (2024)
KoCoSa: Korean Context-aware Sarcasm Detection Dataset
di: Kim, Yumin, et al.
Pubblicazione: (2024)
di: Kim, Yumin, et al.
Pubblicazione: (2024)
Optimizing Language Augmentation for Multilingual Large Language Models: A Case Study on Korean
di: Choi, ChangSu, et al.
Pubblicazione: (2024)
di: Choi, ChangSu, et al.
Pubblicazione: (2024)
Retrieve Only Relevant Tables Whether Few or Many: Adaptive Table Retrieval Method
di: Kim, Taehee, et al.
Pubblicazione: (2026)
di: Kim, Taehee, et al.
Pubblicazione: (2026)
GEM: A Gym for Agentic LLMs
di: Liu, Zichen, et al.
Pubblicazione: (2025)
di: Liu, Zichen, et al.
Pubblicazione: (2025)
From Words to Proverbs: Evaluating LLMs Linguistic and Cultural Competence in Saudi Dialects with Absher
di: Al-Monef, Renad, et al.
Pubblicazione: (2025)
di: Al-Monef, Renad, et al.
Pubblicazione: (2025)
CORAL: Adaptive Retrieval Loop for Culturally-Aligned Multilingual RAG
di: Lee, Nayeon, et al.
Pubblicazione: (2026)
di: Lee, Nayeon, et al.
Pubblicazione: (2026)
Nunchi-Bench: Benchmarking Language Models on Cultural Reasoning with a Focus on Korean Superstition
di: Kim, Kyuhee, et al.
Pubblicazione: (2025)
di: Kim, Kyuhee, et al.
Pubblicazione: (2025)
When Thoughts Meet Facts: Reusable Reasoning for Long-Context LMs
di: Jeong, Soyeong, et al.
Pubblicazione: (2025)
di: Jeong, Soyeong, et al.
Pubblicazione: (2025)
Ko-MuSR: A Multistep Soft Reasoning Benchmark for LLMs Capable of Understanding Korean
di: Park, Chanwoo, et al.
Pubblicazione: (2025)
di: Park, Chanwoo, et al.
Pubblicazione: (2025)
From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation
di: Hong, Seokhee, et al.
Pubblicazione: (2025)
di: Hong, Seokhee, et al.
Pubblicazione: (2025)
GECKO: Generative Language Model for English, Code and Korean
di: Oh, Sungwoo, et al.
Pubblicazione: (2024)
di: Oh, Sungwoo, et al.
Pubblicazione: (2024)
Leveraging What's Overfixed: Post-Correction via LLM Grammatical Error Overcorrection
di: Park, Taehee, et al.
Pubblicazione: (2025)
di: Park, Taehee, et al.
Pubblicazione: (2025)
CUPID: Evaluating Personalized and Contextualized Alignment of LLMs from Interactions
di: Kim, Tae Soo, et al.
Pubblicazione: (2025)
di: Kim, Tae Soo, et al.
Pubblicazione: (2025)
Efficient Latent Semantic Clustering for Scaling Test-Time Computation of LLMs
di: Lee, Sungjae, et al.
Pubblicazione: (2025)
di: Lee, Sungjae, et al.
Pubblicazione: (2025)
UKTA: Unified Korean Text Analyzer
di: Ahn, Seokho, et al.
Pubblicazione: (2025)
di: Ahn, Seokho, et al.
Pubblicazione: (2025)
Language-Agnostic Suicidal Risk Detection Using Large Language Models
di: Kim, June-Woo, et al.
Pubblicazione: (2025)
di: Kim, June-Woo, et al.
Pubblicazione: (2025)
KVoiceBench, KOpenAudioBench, and KMMAU: Agent-Driven Korean Speech Benchmarks for Evaluating SpeechLMs
di: Kim, Haechan, et al.
Pubblicazione: (2026)
di: Kim, Haechan, et al.
Pubblicazione: (2026)
Measuring Political Bias in Large Language Models: What Is Said and How It Is Said
di: Bang, Yejin, et al.
Pubblicazione: (2024)
di: Bang, Yejin, et al.
Pubblicazione: (2024)
The Collective Turing Test: Large Language Models Can Generate Realistic Multi-User Discussions
di: Bouleimen, Azza, et al.
Pubblicazione: (2025)
di: Bouleimen, Azza, et al.
Pubblicazione: (2025)
Documenti analoghi
-
KOMBO: Korean Character Representations Based on the Combination Rules of Subcharacters
di: Kim, SungHo, et al.
Pubblicazione: (2026) -
SCRIPT: A Subcharacter Compositional Representation Injection Module for Korean Pre-Trained Language Models
di: Kim, SungHo, et al.
Pubblicazione: (2026) -
Incorporating Domain Knowledge into Materials Tokenization
di: Oh, Yerim, et al.
Pubblicazione: (2025) -
Handling Korean Out-of-Vocabulary Words with Phoneme Representation Learning
di: Kim, Nayeon, et al.
Pubblicazione: (2025) -
Mentor-KD: Making Small Language Models Better Multi-step Reasoners
di: Lee, Hojae, et al.
Pubblicazione: (2024)