Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cho, Gyeongje, So, Yeonkyoun, Park, Chanwoo, Lee, Sangmin, Jung, Sungmok, Lee, Jaejin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Thunder-LLM: Efficiently Adapting LLMs to Korean with Minimal Resources
von: Kim, Jinpyo, et al.
Veröffentlicht: (2025)
von: Kim, Jinpyo, et al.
Veröffentlicht: (2025)
Choices Speak Louder than Questions
von: Cho, Gyeongje, et al.
Veröffentlicht: (2025)
von: Cho, Gyeongje, et al.
Veröffentlicht: (2025)
Thunder-KoNUBench: A Corpus-Aligned Benchmark for Korean Negation Understanding
von: Jung, Sungmok, et al.
Veröffentlicht: (2026)
von: Jung, Sungmok, et al.
Veröffentlicht: (2026)
Thunder-NUBench: A Benchmark for LLMs' Sentence-Level Negation Understanding
von: So, Yeonkyoung, et al.
Veröffentlicht: (2025)
von: So, Yeonkyoung, et al.
Veröffentlicht: (2025)
Models Know Models Best: Evaluation via Model-Preferred Formats
von: Lee, Joonhak, et al.
Veröffentlicht: (2026)
von: Lee, Joonhak, et al.
Veröffentlicht: (2026)
Accelerating Multilingual Language Model for Excessively Tokenized Languages
von: Hong, Jimin, et al.
Veröffentlicht: (2024)
von: Hong, Jimin, et al.
Veröffentlicht: (2024)
Ko-MuSR: A Multistep Soft Reasoning Benchmark for LLMs Capable of Understanding Korean
von: Park, Chanwoo, et al.
Veröffentlicht: (2025)
von: Park, Chanwoo, et al.
Veröffentlicht: (2025)
Beyond Line-Level Filtering for the Pretraining Corpora of LLMs
von: Park, Chanwoo, et al.
Veröffentlicht: (2025)
von: Park, Chanwoo, et al.
Veröffentlicht: (2025)
Thunder-DeID: Accurate and Efficient De-identification Framework for Korean Court Judgments
von: Hahm, Sungeun, et al.
Veröffentlicht: (2025)
von: Hahm, Sungeun, et al.
Veröffentlicht: (2025)
TiTok: Transfer Token-level Knowledge via Contrastive Excess to Transplant LoRA
von: Jung, Chanjoo, et al.
Veröffentlicht: (2025)
von: Jung, Chanjoo, et al.
Veröffentlicht: (2025)
SupraTok: Cross-Boundary Tokenization for Enhanced Language Model Performance
von: Tănase, Andrei-Valentin, et al.
Veröffentlicht: (2025)
von: Tănase, Andrei-Valentin, et al.
Veröffentlicht: (2025)
TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning
von: Zhang, Tunyu, et al.
Veröffentlicht: (2025)
von: Zhang, Tunyu, et al.
Veröffentlicht: (2025)
FunctionChat-Bench: Comprehensive Evaluation of Language Models' Generative Capabilities in Korean Tool-use Dialogs
von: Lee, Shinbok, et al.
Veröffentlicht: (2024)
von: Lee, Shinbok, et al.
Veröffentlicht: (2024)
Speed and Conversational Large Language Models: Not All Is About Tokens per Second
von: Conde, Javier, et al.
Veröffentlicht: (2025)
von: Conde, Javier, et al.
Veröffentlicht: (2025)
Open Ko-LLM Leaderboard: Evaluating Large Language Models in Korean with Ko-H5 Benchmark
von: Park, Chanjun, et al.
Veröffentlicht: (2024)
von: Park, Chanjun, et al.
Veröffentlicht: (2024)
DyLLM: Efficient Diffusion LLM Inference via Saliency-based Token Selection and Partial Attention
von: Lee, Younjoo, et al.
Veröffentlicht: (2026)
von: Lee, Younjoo, et al.
Veröffentlicht: (2026)
Token-Supervised Value Models for Enhancing Mathematical Problem-Solving Capabilities of Large Language Models
von: Lee, Jung Hyun, et al.
Veröffentlicht: (2024)
von: Lee, Jung Hyun, et al.
Veröffentlicht: (2024)
From Tokens to Words: On the Inner Lexicon of LLMs
von: Kaplan, Guy, et al.
Veröffentlicht: (2024)
von: Kaplan, Guy, et al.
Veröffentlicht: (2024)
Text Generation Beyond Discrete Token Sampling
von: Zhuang, Yufan, et al.
Veröffentlicht: (2025)
von: Zhuang, Yufan, et al.
Veröffentlicht: (2025)
TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
von: Lin, Haokun, et al.
Veröffentlicht: (2025)
von: Lin, Haokun, et al.
Veröffentlicht: (2025)
Incorporating Domain Knowledge into Materials Tokenization
von: Oh, Yerim, et al.
Veröffentlicht: (2025)
von: Oh, Yerim, et al.
Veröffentlicht: (2025)
Theme-Explanation Structure for Table Summarization using Large Language Models: A Case Study on Korean Tabular Data
von: Kwack, TaeYoon, et al.
Veröffentlicht: (2025)
von: Kwack, TaeYoon, et al.
Veröffentlicht: (2025)
Control Token with Dense Passage Retrieval
von: Lee, Juhwan, et al.
Veröffentlicht: (2024)
von: Lee, Juhwan, et al.
Veröffentlicht: (2024)
Unsupervised Extractive Dialogue Summarization in Hyperdimensional Space
von: Park, Seongmin, et al.
Veröffentlicht: (2024)
von: Park, Seongmin, et al.
Veröffentlicht: (2024)
TimeTok: Granularity-Controllable Time-Series Generation via Hierarchical Tokenization
von: Lee, Seokhyun, et al.
Veröffentlicht: (2026)
von: Lee, Seokhyun, et al.
Veröffentlicht: (2026)
The Art of Breaking Words: Rethinking Multilingual Tokenizer Design
von: Thakur, Aamod, et al.
Veröffentlicht: (2025)
von: Thakur, Aamod, et al.
Veröffentlicht: (2025)
UKTA: Unified Korean Text Analyzer
von: Ahn, Seokho, et al.
Veröffentlicht: (2025)
von: Ahn, Seokho, et al.
Veröffentlicht: (2025)
Thinking Tokens for Language Modeling
von: Herel, David, et al.
Veröffentlicht: (2024)
von: Herel, David, et al.
Veröffentlicht: (2024)
Tokenization Matters! Degrading Large Language Models through Challenging Their Tokenization
von: Wang, Dixuan, et al.
Veröffentlicht: (2024)
von: Wang, Dixuan, et al.
Veröffentlicht: (2024)
Explainability-Based Token Replacement on LLM-Generated Text
von: Mohammadi, Hadi, et al.
Veröffentlicht: (2025)
von: Mohammadi, Hadi, et al.
Veröffentlicht: (2025)
On Epistemic Uncertainty of Visual Tokens for Object Hallucinations in Large Vision-Language Models
von: Seo, Hoigi, et al.
Veröffentlicht: (2025)
von: Seo, Hoigi, et al.
Veröffentlicht: (2025)
Building Resource-Constrained Language Agents: A Korean Case Study on Chemical Toxicity Information
von: Cho, Hojun, et al.
Veröffentlicht: (2025)
von: Cho, Hojun, et al.
Veröffentlicht: (2025)
On the Fallacy of Global Token Perplexity in Spoken Language Model Evaluation
von: Hsu, Chan-Jan, et al.
Veröffentlicht: (2026)
von: Hsu, Chan-Jan, et al.
Veröffentlicht: (2026)
SemToken: Semantic-Aware Tokenization for Efficient Long-Context Language Modeling
von: Liu, Dong, et al.
Veröffentlicht: (2025)
von: Liu, Dong, et al.
Veröffentlicht: (2025)
One Token Is Enough: Improving Diffusion Language Models with a Sink Token
von: Zhang, Zihou, et al.
Veröffentlicht: (2026)
von: Zhang, Zihou, et al.
Veröffentlicht: (2026)
Token Assorted: Mixing Latent and Text Tokens for Improved Language Model Reasoning
von: Su, DiJia, et al.
Veröffentlicht: (2025)
von: Su, DiJia, et al.
Veröffentlicht: (2025)
Training Text-to-Molecule Models with Context-Aware Tokenization
von: Kim, Seojin, et al.
Veröffentlicht: (2025)
von: Kim, Seojin, et al.
Veröffentlicht: (2025)
SCRIPT: A Subcharacter Compositional Representation Injection Module for Korean Pre-Trained Language Models
von: Kim, SungHo, et al.
Veröffentlicht: (2026)
von: Kim, SungHo, et al.
Veröffentlicht: (2026)
Alternatives To Next Token Prediction In Text Generation -- A Survey
von: Wyatt, Charlie, et al.
Veröffentlicht: (2025)
von: Wyatt, Charlie, et al.
Veröffentlicht: (2025)
Reinforcement Learning with Token-level Feedback for Controllable Text Generation
von: Li, Wendi, et al.
Veröffentlicht: (2024)
von: Li, Wendi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Thunder-LLM: Efficiently Adapting LLMs to Korean with Minimal Resources
von: Kim, Jinpyo, et al.
Veröffentlicht: (2025) -
Choices Speak Louder than Questions
von: Cho, Gyeongje, et al.
Veröffentlicht: (2025) -
Thunder-KoNUBench: A Corpus-Aligned Benchmark for Korean Negation Understanding
von: Jung, Sungmok, et al.
Veröffentlicht: (2026) -
Thunder-NUBench: A Benchmark for LLMs' Sentence-Level Negation Understanding
von: So, Yeonkyoung, et al.
Veröffentlicht: (2025) -
Models Know Models Best: Evaluation via Model-Preferred Formats
von: Lee, Joonhak, et al.
Veröffentlicht: (2026)