Thunder-LLM: Efficiently Adapting LLMs to Korean with Minimal Resources
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Jinpyo, Cho, Gyeongje, Park, Chanwoo, Park, Jongwon, Kim, Jongmin, So, Yeonkyoun, Lee, Jaejin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models
by: Cho, Gyeongje, et al.
Published: (2025)
by: Cho, Gyeongje, et al.
Published: (2025)
Thunder-DeID: Accurate and Efficient De-identification Framework for Korean Court Judgments
by: Hahm, Sungeun, et al.
Published: (2025)
by: Hahm, Sungeun, et al.
Published: (2025)
Choices Speak Louder than Questions
by: Cho, Gyeongje, et al.
Published: (2025)
by: Cho, Gyeongje, et al.
Published: (2025)
Beyond Line-Level Filtering for the Pretraining Corpora of LLMs
by: Park, Chanwoo, et al.
Published: (2025)
by: Park, Chanwoo, et al.
Published: (2025)
Thunder-KoNUBench: A Corpus-Aligned Benchmark for Korean Negation Understanding
by: Jung, Sungmok, et al.
Published: (2026)
by: Jung, Sungmok, et al.
Published: (2026)
Thunder-NUBench: A Benchmark for LLMs' Sentence-Level Negation Understanding
by: So, Yeonkyoung, et al.
Published: (2025)
by: So, Yeonkyoung, et al.
Published: (2025)
Ko-MuSR: A Multistep Soft Reasoning Benchmark for LLMs Capable of Understanding Korean
by: Park, Chanwoo, et al.
Published: (2025)
by: Park, Chanwoo, et al.
Published: (2025)
Integrating Spatial and Frequency Information for Under-Display Camera Image Restoration
by: Ahn, Kyusu, et al.
Published: (2025)
by: Ahn, Kyusu, et al.
Published: (2025)
HetCCL: Accelerating LLM Training with Heterogeneous GPUs
by: Kim, Heehoon, et al.
Published: (2026)
by: Kim, Heehoon, et al.
Published: (2026)
Open Ko-LLM Leaderboard2: Bridging Foundational and Practical Evaluation for Korean LLMs
by: Kim, Hyeonwoo, et al.
Published: (2024)
by: Kim, Hyeonwoo, et al.
Published: (2024)
Open Ko-LLM Leaderboard: Evaluating Large Language Models in Korean with Ko-H5 Benchmark
by: Park, Chanjun, et al.
Published: (2024)
by: Park, Chanjun, et al.
Published: (2024)
How language models extrapolate outside the training data: A case study in Textualized Gridworld
by: Kim, Doyoung, et al.
Published: (2024)
by: Kim, Doyoung, et al.
Published: (2024)
Are they lovers or friends? Evaluating LLMs' Social Reasoning in English and Korean Dialogues
by: Kim, Eunsu, et al.
Published: (2025)
by: Kim, Eunsu, et al.
Published: (2025)
Can Code-Switched Texts Activate a Knowledge Switch in LLMs? A Case Study on English-Korean Code-Switching
by: Kim, Seoyeon, et al.
Published: (2024)
by: Kim, Seoyeon, et al.
Published: (2024)
SEDD: Scalable and Efficient Dataset Deduplication with GPUs
by: Son, Youngjun, et al.
Published: (2025)
by: Son, Youngjun, et al.
Published: (2025)
Steering LLMs toward Korean Local Speech: Iterative Refinement Framework for Faithful Dialect Translation
by: Park, Keunhyeung, et al.
Published: (2025)
by: Park, Keunhyeung, et al.
Published: (2025)
Unsupervised Extractive Dialogue Summarization in Hyperdimensional Space
by: Park, Seongmin, et al.
Published: (2024)
by: Park, Seongmin, et al.
Published: (2024)
RedWhale: An Adapted Korean LLM Through Efficient Continual Pretraining
by: Vo, Anh-Dung, et al.
Published: (2024)
by: Vo, Anh-Dung, et al.
Published: (2024)
K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts
by: Lee, Nahyun, et al.
Published: (2026)
by: Lee, Nahyun, et al.
Published: (2026)
Modeling Layered Consciousness with Multi-Agent Large Language Models
by: Kim, Sang Hun, et al.
Published: (2025)
by: Kim, Sang Hun, et al.
Published: (2025)
Assessing Socio-Cultural Alignment and Technical Safety of Sovereign LLMs
by: Chae, Kyubyung, et al.
Published: (2025)
by: Chae, Kyubyung, et al.
Published: (2025)
HiKE: Hierarchical Evaluation Framework for Korean-English Code-Switching Speech Recognition
by: Paik, Gio, et al.
Published: (2025)
by: Paik, Gio, et al.
Published: (2025)
Making Qwen3 Think in Korean with Reinforcement Learning
by: Lee, Jungyup, et al.
Published: (2025)
by: Lee, Jungyup, et al.
Published: (2025)
PHISH in MESH: Korean Adversarial Phonetic Substitution and Phonetic-Semantic Feature Integration Defense
by: Kim, Byungjun, et al.
Published: (2025)
by: Kim, Byungjun, et al.
Published: (2025)
Expanding Search Space with Diverse Prompting Agents: An Efficient Sampling Approach for LLM Mathematical Reasoning
by: Lee, Gisang, et al.
Published: (2024)
by: Lee, Gisang, et al.
Published: (2024)
SDS KoPub VDR: A Benchmark Dataset for Visual Document Retrieval in Korean Public Documents
by: Lee, Jaehoon, et al.
Published: (2025)
by: Lee, Jaehoon, et al.
Published: (2025)
Enhancing Korean Dependency Parsing with Morphosyntactic Features
by: Park, Jungyeul, et al.
Published: (2025)
by: Park, Jungyeul, et al.
Published: (2025)
KOMBO: Korean Character Representations Based on the Combination Rules of Subcharacters
by: Kim, SungHo, et al.
Published: (2026)
by: Kim, SungHo, et al.
Published: (2026)
Evaluating Multimodal Generative AI with Korean Educational Standards
by: Park, Sanghee, et al.
Published: (2025)
by: Park, Sanghee, et al.
Published: (2025)
KatFishNet: Detecting LLM-Generated Korean Text through Linguistic Feature Analysis
by: Park, Shinwoo, et al.
Published: (2025)
by: Park, Shinwoo, et al.
Published: (2025)
KLAAD: Refining Attention Mechanisms to Reduce Societal Bias in Generative Language Models
by: Kim, Seorin, et al.
Published: (2025)
by: Kim, Seorin, et al.
Published: (2025)
Optimizing Korean-Centric LLMs via Token Pruning
by: Kim, Hoyeol, et al.
Published: (2026)
by: Kim, Hoyeol, et al.
Published: (2026)
Handling Korean Out-of-Vocabulary Words with Phoneme Representation Learning
by: Kim, Nayeon, et al.
Published: (2025)
by: Kim, Nayeon, et al.
Published: (2025)
Models Know Models Best: Evaluation via Model-Preferred Formats
by: Lee, Joonhak, et al.
Published: (2026)
by: Lee, Joonhak, et al.
Published: (2026)
Multi-Dimensional Machine Translation Evaluation: Model Evaluation and Resource for Korean
by: Park, Dojun, et al.
Published: (2024)
by: Park, Dojun, et al.
Published: (2024)
Enhancing Document-Level Machine Translation via Filtered Synthetic Corpora and Two-Stage LLM Adaptation
by: Kim, Ireh, et al.
Published: (2026)
by: Kim, Ireh, et al.
Published: (2026)
Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models
by: Kim, Zae Myung, et al.
Published: (2025)
by: Kim, Zae Myung, et al.
Published: (2025)
KorNAT: LLM Alignment Benchmark for Korean Social Values and Common Knowledge
by: Lee, Jiyoung, et al.
Published: (2024)
by: Lee, Jiyoung, et al.
Published: (2024)
KITE: A Benchmark for Evaluating Korean Instruction-Following Abilities in Large Language Models
by: Kim, Dongjun, et al.
Published: (2025)
by: Kim, Dongjun, et al.
Published: (2025)
Building Resource-Constrained Language Agents: A Korean Case Study on Chemical Toxicity Information
by: Cho, Hojun, et al.
Published: (2025)
by: Cho, Hojun, et al.
Published: (2025)
Similar Items
-
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models
by: Cho, Gyeongje, et al.
Published: (2025) -
Thunder-DeID: Accurate and Efficient De-identification Framework for Korean Court Judgments
by: Hahm, Sungeun, et al.
Published: (2025) -
Choices Speak Louder than Questions
by: Cho, Gyeongje, et al.
Published: (2025) -
Beyond Line-Level Filtering for the Pretraining Corpora of LLMs
by: Park, Chanwoo, et al.
Published: (2025) -
Thunder-KoNUBench: A Corpus-Aligned Benchmark for Korean Negation Understanding
by: Jung, Sungmok, et al.
Published: (2026)