RedWhale: An Adapted Korean LLM Through Efficient Continual Pretraining
Fuente:
arXiv
Saved in:
| Main Authors: | Vo, Anh-Dung, Jung, Minseong, Lee, Wonbeen, Choi, Daewoo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Thunder-LLM: Efficiently Adapting LLMs to Korean with Minimal Resources
by: Kim, Jinpyo, et al.
Published: (2025)
by: Kim, Jinpyo, et al.
Published: (2025)
Analyzing Error Propagation in Korean Spoken QA with ASR-LLM Cascades
by: Jung, Donghyuk, et al.
Published: (2026)
by: Jung, Donghyuk, et al.
Published: (2026)
PICKT: Practical Interlinked Concept Knowledge Tracing for Personalized Learning using Knowledge Map Concept Relations
by: Lee, Wonbeen, et al.
Published: (2025)
by: Lee, Wonbeen, et al.
Published: (2025)
DiffAdapt: Difficulty-Adaptive Reasoning for Token-Efficient LLM Inference
by: Liu, Xiang, et al.
Published: (2025)
by: Liu, Xiang, et al.
Published: (2025)
KorNAT: LLM Alignment Benchmark for Korean Social Values and Common Knowledge
by: Lee, Jiyoung, et al.
Published: (2024)
by: Lee, Jiyoung, et al.
Published: (2024)
LLM Pretraining with Continuous Concepts
by: Tack, Jihoon, et al.
Published: (2025)
by: Tack, Jihoon, et al.
Published: (2025)
Rethinking Multilingual Continual Pretraining: Data Mixing for Adapting LLMs Across Languages and Resources
by: Li, Zihao, et al.
Published: (2025)
by: Li, Zihao, et al.
Published: (2025)
ELO: Efficient Layer-Specific Optimization for Continual Pretraining of Multilingual LLMs
by: Yoo, HanGyeol, et al.
Published: (2026)
by: Yoo, HanGyeol, et al.
Published: (2026)
ManufactuBERT: Efficient Continual Pretraining for Manufacturing
by: Armingaud, Robin, et al.
Published: (2025)
by: Armingaud, Robin, et al.
Published: (2025)
Tracing Persona Vectors Through LLM Pretraining
by: Moskvoretskii, Viktor, et al.
Published: (2026)
by: Moskvoretskii, Viktor, et al.
Published: (2026)
Craw4LLM: Efficient Web Crawling for LLM Pretraining
by: Yu, Shi, et al.
Published: (2025)
by: Yu, Shi, et al.
Published: (2025)
Does Incomplete Syntax Influence Korean Language Model? Focusing on Word Order and Case Markers
by: Kim, Jong Myoung, et al.
Published: (2024)
by: Kim, Jong Myoung, et al.
Published: (2024)
HAE-RAE Bench: Evaluation of Korean Knowledge in Language Models
by: Son, Guijin, et al.
Published: (2023)
by: Son, Guijin, et al.
Published: (2023)
Multilingual LLM Prompting Strategies for Medical English-Vietnamese Machine Translation
by: Vo, Nhu, et al.
Published: (2025)
by: Vo, Nhu, et al.
Published: (2025)
Does Continued Pretraining on a Learner Corpus Improve Automated Essay Scoring on English Proficiency Tests? Evidence from EFCAMDAT
by: Nguyen, Duy Anh
Published: (2026)
by: Nguyen, Duy Anh
Published: (2026)
PonderLM-2: Pretraining LLM with Latent Thoughts in Continuous Space
by: Zeng, Boyi, et al.
Published: (2025)
by: Zeng, Boyi, et al.
Published: (2025)
Hear Both Sides: Efficient Multi-Agent Debate via Diversity-Aware Message Retention
by: Nguyen, Manh, et al.
Published: (2026)
by: Nguyen, Manh, et al.
Published: (2026)
Thunder-KoNUBench: A Corpus-Aligned Benchmark for Korean Negation Understanding
by: Jung, Sungmok, et al.
Published: (2026)
by: Jung, Sungmok, et al.
Published: (2026)
EVOKE: Emotion Vocabulary Of Korean and English
by: Jung, Yoonwon, et al.
Published: (2026)
by: Jung, Yoonwon, et al.
Published: (2026)
Efficient Domain-adaptive Continual Pretraining for the Process Industry in the German Language
by: Zhukova, Anastasia, et al.
Published: (2025)
by: Zhukova, Anastasia, et al.
Published: (2025)
Multi-Step Reasoning in Korean and the Emergent Mirage
by: Son, Guijin, et al.
Published: (2025)
by: Son, Guijin, et al.
Published: (2025)
Thunder-DeID: Accurate and Efficient De-identification Framework for Korean Court Judgments
by: Hahm, Sungeun, et al.
Published: (2025)
by: Hahm, Sungeun, et al.
Published: (2025)
Register Always Matters: Analysis of LLM Pretraining Data Through the Lens of Language Variation
by: Myntti, Amanda, et al.
Published: (2025)
by: Myntti, Amanda, et al.
Published: (2025)
Vi-Mistral-X: Building a Vietnamese Language Model with Advanced Continual Pre-training
by: Vo, James
Published: (2024)
by: Vo, James
Published: (2024)
Model Swarms: Collaborative Search to Adapt LLM Experts via Swarm Intelligence
by: Feng, Shangbin, et al.
Published: (2024)
by: Feng, Shangbin, et al.
Published: (2024)
NMIXX: Domain-Adapted Neural Embeddings for Cross-Lingual eXploration of Finance
by: Lee, Hanwool, et al.
Published: (2025)
by: Lee, Hanwool, et al.
Published: (2025)
Vision-and-Language Pretraining
by: Nguyen, Thong, et al.
Published: (2022)
by: Nguyen, Thong, et al.
Published: (2022)
Curió-Edu 7B: Examining Data Selection Impacts in LLM Continued Pretraining
by: Almeida, Thales Sales, et al.
Published: (2025)
by: Almeida, Thales Sales, et al.
Published: (2025)
SparseAccelerate: Efficient Long-Context Inference for Mid-Range GPUs
by: Vo, James
Published: (2024)
by: Vo, James
Published: (2024)
A Continued Pretrained LLM Approach for Automatic Medical Note Generation
by: Yuan, Dong, et al.
Published: (2024)
by: Yuan, Dong, et al.
Published: (2024)
QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining
by: Liu, Fengze, et al.
Published: (2025)
by: Liu, Fengze, et al.
Published: (2025)
Won: Establishing Best Practices for Korean Financial NLP
by: Son, Guijin, et al.
Published: (2025)
by: Son, Guijin, et al.
Published: (2025)
KMMLU: Measuring Massive Multitask Language Understanding in Korean
by: Son, Guijin, et al.
Published: (2024)
by: Son, Guijin, et al.
Published: (2024)
Transformer Layer Injection: A Novel Approach for Efficient Upscaling of Large Language Models
by: Vo, James
Published: (2024)
by: Vo, James
Published: (2024)
Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length
by: Ma, Xuezhe, et al.
Published: (2024)
by: Ma, Xuezhe, et al.
Published: (2024)
Style over Story: Measuring LLM Narrative Preferences via Structured Selection
by: Jung, Donghoon, et al.
Published: (2025)
by: Jung, Donghoon, et al.
Published: (2025)
Adapting Pretrained Language Models for Citation Classification via Self-Supervised Contrastive Learning
by: Li, Tong, et al.
Published: (2025)
by: Li, Tong, et al.
Published: (2025)
FINEST: Improving LLM Responses to Sensitive Topics Through Fine-Grained Evaluation
by: Oh, Juhyun, et al.
Published: (2026)
by: Oh, Juhyun, et al.
Published: (2026)
T-MAP: Red-Teaming LLM Agents with Trajectory-aware Evolutionary Search
by: Lee, Hyomin, et al.
Published: (2026)
by: Lee, Hyomin, et al.
Published: (2026)
KORMo: Korean Open Reasoning Model for Everyone
by: Kim, Minjun, et al.
Published: (2025)
by: Kim, Minjun, et al.
Published: (2025)
Similar Items
-
Thunder-LLM: Efficiently Adapting LLMs to Korean with Minimal Resources
by: Kim, Jinpyo, et al.
Published: (2025) -
Analyzing Error Propagation in Korean Spoken QA with ASR-LLM Cascades
by: Jung, Donghyuk, et al.
Published: (2026) -
PICKT: Practical Interlinked Concept Knowledge Tracing for Personalized Learning using Knowledge Map Concept Relations
by: Lee, Wonbeen, et al.
Published: (2025) -
DiffAdapt: Difficulty-Adaptive Reasoning for Token-Efficient LLM Inference
by: Liu, Xiang, et al.
Published: (2025) -
KorNAT: LLM Alignment Benchmark for Korean Social Values and Common Knowledge
by: Lee, Jiyoung, et al.
Published: (2024)