SEDD: Scalable and Efficient Dataset Deduplication with GPUs
Fuente:
arXiv
Saved in:
| Main Authors: | Son, Youngjun, Kim, Chaewon, Lee, Jaejin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
KLAAD: Refining Attention Mechanisms to Reduce Societal Bias in Generative Language Models
by: Kim, Seorin, et al.
Published: (2025)
by: Kim, Seorin, et al.
Published: (2025)
Thunder-DeID: Accurate and Efficient De-identification Framework for Korean Court Judgments
by: Hahm, Sungeun, et al.
Published: (2025)
by: Hahm, Sungeun, et al.
Published: (2025)
FENCE: A Financial and Multimodal Jailbreak Detection Dataset
by: Kim, Mirae, et al.
Published: (2026)
by: Kim, Mirae, et al.
Published: (2026)
Thunder-LLM: Efficiently Adapting LLMs to Korean with Minimal Resources
by: Kim, Jinpyo, et al.
Published: (2025)
by: Kim, Jinpyo, et al.
Published: (2025)
Generative Deduplication For Socia Media Data Selection
by: Li, Xianming, et al.
Published: (2024)
by: Li, Xianming, et al.
Published: (2024)
Assessing Socio-Cultural Alignment and Technical Safety of Sovereign LLMs
by: Chae, Kyubyung, et al.
Published: (2025)
by: Chae, Kyubyung, et al.
Published: (2025)
BankMathBench: A Benchmark for Numerical Reasoning in Banking Scenarios
by: Lee, Yunseung, et al.
Published: (2026)
by: Lee, Yunseung, et al.
Published: (2026)
Deduplicating and Ranking Solution Programs for Suggesting Reference Solutions
by: Shirafuji, Atsushi, et al.
Published: (2023)
by: Shirafuji, Atsushi, et al.
Published: (2023)
Reasoning by Commented Code for Table Question Answering
by: Pyo, Seho, et al.
Published: (2026)
by: Pyo, Seho, et al.
Published: (2026)
Unsupervised Extractive Dialogue Summarization in Hyperdimensional Space
by: Park, Seongmin, et al.
Published: (2024)
by: Park, Seongmin, et al.
Published: (2024)
Models Know Models Best: Evaluation via Model-Preferred Formats
by: Lee, Joonhak, et al.
Published: (2026)
by: Lee, Joonhak, et al.
Published: (2026)
FairDeDup: Detecting and Mitigating Vision-Language Fairness Disparities in Semantic Dataset Deduplication
by: Slyman, Eric, et al.
Published: (2024)
by: Slyman, Eric, et al.
Published: (2024)
LaDiMo: Layer-wise Distillation Inspired MoEfier
by: Kim, Sungyoon, et al.
Published: (2024)
by: Kim, Sungyoon, et al.
Published: (2024)
Thunder-KoNUBench: A Corpus-Aligned Benchmark for Korean Negation Understanding
by: Jung, Sungmok, et al.
Published: (2026)
by: Jung, Sungmok, et al.
Published: (2026)
ExpGuard: LLM Content Moderation in Specialized Domains
by: Choi, Minseok, et al.
Published: (2026)
by: Choi, Minseok, et al.
Published: (2026)
Choices Speak Louder than Questions
by: Cho, Gyeongje, et al.
Published: (2025)
by: Cho, Gyeongje, et al.
Published: (2025)
Stress-Testing Emotional Support Models: Moving from Homogeneous to Diverse Help Seekers
by: Heo, Chaewon, et al.
Published: (2026)
by: Heo, Chaewon, et al.
Published: (2026)
Bounded Hyperbolic Tangent: A Stable and Efficient Alternative to Pre-Layer Normalization in Large Language Models
by: Byun, Hoyoon, et al.
Published: (2025)
by: Byun, Hoyoon, et al.
Published: (2025)
Ko-PIQA: A Korean Physical Commonsense Reasoning Dataset with Cultural Context
by: Choi, Dasol, et al.
Published: (2025)
by: Choi, Dasol, et al.
Published: (2025)
Thunder-NUBench: A Benchmark for LLMs' Sentence-Level Negation Understanding
by: So, Yeonkyoung, et al.
Published: (2025)
by: So, Yeonkyoung, et al.
Published: (2025)
DuET: Dual Execution for Test Output Prediction with Generated Code and Pseudocode
by: Han, Hojae, et al.
Published: (2026)
by: Han, Hojae, et al.
Published: (2026)
Merlin: Deterministic Byte-Exact Deduplication for Lossless Context Optimization in Large Language Model Inference
by: Schelpe, Sietse
Published: (2026)
by: Schelpe, Sietse
Published: (2026)
PILOT-Bench: A Benchmark for Legal Reasoning in the Patent Domain with IRAC-Aligned Classification Tasks
by: Jang, Yehoon, et al.
Published: (2026)
by: Jang, Yehoon, et al.
Published: (2026)
SparseAccelerate: Efficient Long-Context Inference for Mid-Range GPUs
by: Vo, James
Published: (2024)
by: Vo, James
Published: (2024)
Beyond Line-Level Filtering for the Pretraining Corpora of LLMs
by: Park, Chanwoo, et al.
Published: (2025)
by: Park, Chanwoo, et al.
Published: (2025)
References Indeed Matter? Reference-Free Preference Optimization for Conversational Query Reformulation
by: Kim, Doyoung, et al.
Published: (2025)
by: Kim, Doyoung, et al.
Published: (2025)
Byte-Exact Deduplication in Retrieval-Augmented Generation: A Three-Regime Empirical Analysis Across Public Benchmarks
by: Schelpe, Sietse
Published: (2026)
by: Schelpe, Sietse
Published: (2026)
PERC: Plan-As-Query Example Retrieval for Underrepresented Code Generation
by: Yoo, Jaeseok, et al.
Published: (2024)
by: Yoo, Jaeseok, et al.
Published: (2024)
ArchCode: Incorporating Software Requirements in Code Generation with Large Language Models
by: Han, Hojae, et al.
Published: (2024)
by: Han, Hojae, et al.
Published: (2024)
HetCCL: Accelerating LLM Training with Heterogeneous GPUs
by: Kim, Heehoon, et al.
Published: (2026)
by: Kim, Heehoon, et al.
Published: (2026)
Not All Adapters Matter: Selective Adapter Freezing for Memory-Efficient Fine-Tuning of Language Models
by: Son, Hyegang, et al.
Published: (2024)
by: Son, Hyegang, et al.
Published: (2024)
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models
by: Cho, Gyeongje, et al.
Published: (2025)
by: Cho, Gyeongje, et al.
Published: (2025)
Pretraining A Large Language Model using Distributed GPUs: A Memory-Efficient Decentralized Paradigm
by: Zhang, Jinrui, et al.
Published: (2026)
by: Zhang, Jinrui, et al.
Published: (2026)
Evaluating Deduplication Techniques for Economic Research Paper Titles with a Focus on Semantic Similarity using NLP and LLMs
by: You, Doohee, et al.
Published: (2024)
by: You, Doohee, et al.
Published: (2024)
KoCoNovel: Annotated Dataset of Character Coreference in Korean Novels
by: Kim, Kyuhee, et al.
Published: (2024)
by: Kim, Kyuhee, et al.
Published: (2024)
Locate&Edit: Energy-based Text Editing for Efficient, Flexible, and Faithful Controlled Text Generation
by: Son, Hye Ryung, et al.
Published: (2024)
by: Son, Hye Ryung, et al.
Published: (2024)
Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents
by: Jang, Yunseok, et al.
Published: (2025)
by: Jang, Yunseok, et al.
Published: (2025)
Pushing the Boundaries of Multiple Choice Evaluation to One Hundred Options
by: Lee, Nahyun, et al.
Published: (2026)
by: Lee, Nahyun, et al.
Published: (2026)
SDS KoPub VDR: A Benchmark Dataset for Visual Document Retrieval in Korean Public Documents
by: Lee, Jaehoon, et al.
Published: (2025)
by: Lee, Jaehoon, et al.
Published: (2025)
Privacy-Preserving Data Deduplication for Enhancing Federated Learning of Language Models (Extended Version)
by: Abadi, Aydin, et al.
Published: (2024)
by: Abadi, Aydin, et al.
Published: (2024)
Similar Items
-
KLAAD: Refining Attention Mechanisms to Reduce Societal Bias in Generative Language Models
by: Kim, Seorin, et al.
Published: (2025) -
Thunder-DeID: Accurate and Efficient De-identification Framework for Korean Court Judgments
by: Hahm, Sungeun, et al.
Published: (2025) -
FENCE: A Financial and Multimodal Jailbreak Detection Dataset
by: Kim, Mirae, et al.
Published: (2026) -
Thunder-LLM: Efficiently Adapting LLMs to Korean with Minimal Resources
by: Kim, Jinpyo, et al.
Published: (2025) -
Generative Deduplication For Socia Media Data Selection
by: Li, Xianming, et al.
Published: (2024)