Beyond Line-Level Filtering for the Pretraining Corpora of LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Park, Chanwoo, Park, Suyoung, Ahn, Yelim, Kim, Jongmin, Park, Jongyeon, Lee, Jaejin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Ko-MuSR: A Multistep Soft Reasoning Benchmark for LLMs Capable of Understanding Korean
by: Park, Chanwoo, et al.
Published: (2025)
by: Park, Chanwoo, et al.
Published: (2025)
Enhancing Document-Level Machine Translation via Filtered Synthetic Corpora and Two-Stage LLM Adaptation
by: Kim, Ireh, et al.
Published: (2026)
by: Kim, Ireh, et al.
Published: (2026)
PUZZLED: Jailbreaking LLMs through Word-Based Puzzles
by: Ahn, Yelim, et al.
Published: (2025)
by: Ahn, Yelim, et al.
Published: (2025)
Thunder-LLM: Efficiently Adapting LLMs to Korean with Minimal Resources
by: Kim, Jinpyo, et al.
Published: (2025)
by: Kim, Jinpyo, et al.
Published: (2025)
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models
by: Cho, Gyeongje, et al.
Published: (2025)
by: Cho, Gyeongje, et al.
Published: (2025)
Rethinking KenLM: Good and Bad Model Ensembles for Efficient Text Quality Filtering in Large Web Corpora
by: Kim, Yungi, et al.
Published: (2024)
by: Kim, Yungi, et al.
Published: (2024)
Unsupervised Extractive Dialogue Summarization in Hyperdimensional Space
by: Park, Seongmin, et al.
Published: (2024)
by: Park, Seongmin, et al.
Published: (2024)
Models Know Models Best: Evaluation via Model-Preferred Formats
by: Lee, Joonhak, et al.
Published: (2026)
by: Lee, Joonhak, et al.
Published: (2026)
Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents
by: Kim, Wonjoong, et al.
Published: (2025)
by: Kim, Wonjoong, et al.
Published: (2025)
RSCF: Relation-Semantics Consistent Filter for Entity Embedding of Knowledge Graph
by: Kim, Junsik, et al.
Published: (2025)
by: Kim, Junsik, et al.
Published: (2025)
MDAgents: An Adaptive Collaboration of LLMs for Medical Decision-Making
by: Kim, Yubin, et al.
Published: (2024)
by: Kim, Yubin, et al.
Published: (2024)
Attributing Culture-Conditioned Generations to Pretraining Corpora
by: Li, Huihan, et al.
Published: (2024)
by: Li, Huihan, et al.
Published: (2024)
Modeling Layered Consciousness with Multi-Agent Large Language Models
by: Kim, Sang Hun, et al.
Published: (2025)
by: Kim, Sang Hun, et al.
Published: (2025)
Neuro-RIT: Neuron-Guided Instruction Tuning for Robust Retrieval-Augmented Language Model
by: Kim, Jaemin, et al.
Published: (2026)
by: Kim, Jaemin, et al.
Published: (2026)
M2S: Multi-turn to Single-turn jailbreak in Red Teaming for LLMs
by: Ha, Junwoo, et al.
Published: (2025)
by: Ha, Junwoo, et al.
Published: (2025)
QPaug: Question and Passage Augmentation for Open-Domain Question Answering of LLMs
by: Kim, Minsang, et al.
Published: (2024)
by: Kim, Minsang, et al.
Published: (2024)
Open Ko-LLM Leaderboard2: Bridging Foundational and Practical Evaluation for Korean LLMs
by: Kim, Hyeonwoo, et al.
Published: (2024)
by: Kim, Hyeonwoo, et al.
Published: (2024)
Aligning Reasoning LLMs for Materials Discovery with Physics-aware Rejection Sampling
by: Hyun, Lee, et al.
Published: (2025)
by: Hyun, Lee, et al.
Published: (2025)
Expanding Search Space with Diverse Prompting Agents: An Efficient Sampling Approach for LLM Mathematical Reasoning
by: Lee, Gisang, et al.
Published: (2024)
by: Lee, Gisang, et al.
Published: (2024)
Efficient Latent Semantic Clustering for Scaling Test-Time Computation of LLMs
by: Lee, Sungjae, et al.
Published: (2025)
by: Lee, Sungjae, et al.
Published: (2025)
Zero-shot Commonsense Reasoning over Machine Imagination
by: Park, Hyuntae, et al.
Published: (2024)
by: Park, Hyuntae, et al.
Published: (2024)
Reward-Weighted Sampling: Enhancing Non-Autoregressive Characteristics in Masked Diffusion LLMs
by: Gwak, Daehoon, et al.
Published: (2025)
by: Gwak, Daehoon, et al.
Published: (2025)
R1-ACT: Efficient Reasoning Model Safety Alignment by Activating Safety Knowledge
by: In, Yeonjun, et al.
Published: (2025)
by: In, Yeonjun, et al.
Published: (2025)
Dataverse: Open-Source ETL (Extract, Transform, Load) Pipeline for Large Language Models
by: Park, Hyunbyung, et al.
Published: (2024)
by: Park, Hyunbyung, et al.
Published: (2024)
ixi-GEN: Efficient Industrial sLLMs through Domain Adaptive Continual Pretraining
by: Kim, Seonwu, et al.
Published: (2025)
by: Kim, Seonwu, et al.
Published: (2025)
GRAIL: Gradient-Based Adaptive Unlearning for Privacy and Copyright in LLMs
by: Kim, Kun-Woo, et al.
Published: (2025)
by: Kim, Kun-Woo, et al.
Published: (2025)
KITE: A Benchmark for Evaluating Korean Instruction-Following Abilities in Large Language Models
by: Kim, Dongjun, et al.
Published: (2025)
by: Kim, Dongjun, et al.
Published: (2025)
SaudiBERT: A Large Language Model Pretrained on Saudi Dialect Corpora
by: Qarah, Faisal
Published: (2024)
by: Qarah, Faisal
Published: (2024)
ProgRAG: Hallucination-Resistant Progressive Retrieval and Reasoning over Knowledge Graphs
by: Park, Minbae, et al.
Published: (2025)
by: Park, Minbae, et al.
Published: (2025)
EconCausal: A Context-Aware Economic Reasoning Benchmark for Large Language Models
by: Lee, Donggyu, et al.
Published: (2025)
by: Lee, Donggyu, et al.
Published: (2025)
Nudging Beyond the Comfort Zone: Efficient Strategy-Guided Exploration for RLVR
by: Lee, Chanuk, et al.
Published: (2026)
by: Lee, Chanuk, et al.
Published: (2026)
SyncVSR: Data-Efficient Visual Speech Recognition with End-to-End Crossmodal Audio Token Synchronization
by: Ahn, Young Jin, et al.
Published: (2024)
by: Ahn, Young Jin, et al.
Published: (2024)
InstaTrans: An Instruction-Aware Translation Framework for Non-English Instruction Datasets
by: Kim, Yungi, et al.
Published: (2024)
by: Kim, Yungi, et al.
Published: (2024)
Understanding LLM Development Through Longitudinal Study: Insights from the Open Ko-LLM Leaderboard
by: Park, Chanjun, et al.
Published: (2024)
by: Park, Chanjun, et al.
Published: (2024)
Your AI, Not Your View: The Bias of LLMs in Investment Analysis
by: Lee, Hoyoung, et al.
Published: (2025)
by: Lee, Hoyoung, et al.
Published: (2025)
InvThink: Premortem Reasoning for Safer Language Models
by: Kim, Yubin, et al.
Published: (2025)
by: Kim, Yubin, et al.
Published: (2025)
Do not think about pink elephant!
by: Hwang, Kyomin, et al.
Published: (2024)
by: Hwang, Kyomin, et al.
Published: (2024)
Enhancing Hallucination Detection via Future Context
by: Lee, Joosung, et al.
Published: (2025)
by: Lee, Joosung, et al.
Published: (2025)
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs
by: Oh, Gyutaek, et al.
Published: (2025)
by: Oh, Gyutaek, et al.
Published: (2025)
Reinforcement Learning from Reflective Feedback (RLRF): Aligning and Improving LLMs via Fine-Grained Self-Reflection
by: Lee, Kyungjae, et al.
Published: (2024)
by: Lee, Kyungjae, et al.
Published: (2024)
Similar Items
-
Ko-MuSR: A Multistep Soft Reasoning Benchmark for LLMs Capable of Understanding Korean
by: Park, Chanwoo, et al.
Published: (2025) -
Enhancing Document-Level Machine Translation via Filtered Synthetic Corpora and Two-Stage LLM Adaptation
by: Kim, Ireh, et al.
Published: (2026) -
PUZZLED: Jailbreaking LLMs through Word-Based Puzzles
by: Ahn, Yelim, et al.
Published: (2025) -
Thunder-LLM: Efficiently Adapting LLMs to Korean with Minimal Resources
by: Kim, Jinpyo, et al.
Published: (2025) -
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models
by: Cho, Gyeongje, et al.
Published: (2025)