SoftDedup: an Efficient Data Reweighting Method for Speeding Up Language Model Pre-training
Fuente:
arXiv
Saved in:
| Main Authors: | He, Nan, Xiong, Weichen, Liu, Hanwen, Liao, Yi, Ding, Lei, Zhang, Kai, Tang, Guohua, Han, Xiao, Yang, Wei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RaphaelRibes/FastDedup: FastDedup v1.1.0
by: Raphaël Ribes
Published: (2026)
by: Raphaël Ribes
Published: (2026)
Velocitune: A Velocity-based Dynamic Domain Reweighting Method for Continual Pre-training
by: Luo, Zheheng, et al.
Published: (2024)
by: Luo, Zheheng, et al.
Published: (2024)
ToReMi: Topic-Aware Data Reweighting for Dynamic Pre-Training Data Selection
by: Zhu, Xiaoxuan, et al.
Published: (2025)
by: Zhu, Xiaoxuan, et al.
Published: (2025)
PM-Dedup: Secure Deduplication with Partial Migration from Cloud to Edge Servers
by: Ke, Zhaokang, et al.
Published: (2025)
by: Ke, Zhaokang, et al.
Published: (2025)
10 Percent Wrong for 90% Done: A Practical Approach to Collection Deduping
by: Hamby, Rogan
Published: (2012)
by: Hamby, Rogan
Published: (2012)
Efficient Data Learning for Open Information Extraction with Pre-trained Language Models
by: Fan, Zhiyuan, et al.
Published: (2023)
by: Fan, Zhiyuan, et al.
Published: (2023)
Efficient Tactile Perception with Soft Electrical Impedance Tomography and Pre-trained Transformer
by: Dong, Huazhi, et al.
Published: (2025)
by: Dong, Huazhi, et al.
Published: (2025)
GPT4Image: Large Pre-trained Models Help Vision Models Learn Better on Perception Task
by: Ding, Ning, et al.
Published: (2023)
by: Ding, Ning, et al.
Published: (2023)
Hadamard Adapter: An Extreme Parameter-Efficient Adapter Tuning Method for Pre-trained Language Models
by: Chen, Yuyan, et al.
Published: (2024)
by: Chen, Yuyan, et al.
Published: (2024)
Enhancing Dense Retrievers' Robustness with Group-level Reweighting
by: Han, Peixuan, et al.
Published: (2023)
by: Han, Peixuan, et al.
Published: (2023)
Evaluating Discourse Cohesion in Pre-trained Language Models
by: He, Jie, et al.
Published: (2025)
by: He, Jie, et al.
Published: (2025)
Pre-training on Synthetic Driving Data for Trajectory Prediction
by: Li, Yiheng, et al.
Published: (2023)
by: Li, Yiheng, et al.
Published: (2023)
Identifying Pre-training Data in LLMs: A Neuron Activation-Based Detection Framework
by: Tang, Hongyi, et al.
Published: (2025)
by: Tang, Hongyi, et al.
Published: (2025)
Pre‐Embedded Multisite Chiral Molecules Realize Bottom‐Up Multilayer Manipulation toward Stable and Efficient Perovskite Solar Cells
by: Qian Zhou, et al.
Published: (2024)
by: Qian Zhou, et al.
Published: (2024)
BUS:Efficient and Effective Vision-language Pre-training with Bottom-Up Patch Summarization
by: Jiang, Chaoya, et al.
Published: (2023)
by: Jiang, Chaoya, et al.
Published: (2023)
Feedback-based Modal Mutual Search for Attacking Vision-Language Pre-training Models
by: Ding, Renhua, et al.
Published: (2024)
by: Ding, Renhua, et al.
Published: (2024)
Scaling Speech-Text Pre-training with Synthetic Interleaved Data
by: Zeng, Aohan, et al.
Published: (2024)
by: Zeng, Aohan, et al.
Published: (2024)
Contrastive Pre-training for Deep Session Data Understanding
by: Li, Zixuan, et al.
Published: (2024)
by: Li, Zixuan, et al.
Published: (2024)
A Reinforcement Learning-Based Automatic Video Editing Method Using Pre-trained Vision-Language Model
by: Hu, Panwen, et al.
Published: (2024)
by: Hu, Panwen, et al.
Published: (2024)
A General and Efficient FFT Computation Method
by: Wan, Yi, et al.
Published: (2026)
by: Wan, Yi, et al.
Published: (2026)
EndoMamba: An Efficient Foundation Model for Endoscopic Videos via Hierarchical Pre-training
by: Tian, Qingyao, et al.
Published: (2025)
by: Tian, Qingyao, et al.
Published: (2025)
Mind the Interference: Retaining Pre-trained Knowledge in Parameter Efficient Continual Learning of Vision-Language Models
by: Tang, Longxiang, et al.
Published: (2024)
by: Tang, Longxiang, et al.
Published: (2024)
UEPS: Robust and Efficient MRI Reconstruction (Pre-trained Model and Demo Data)
by: Zhou, Xiang, et al.
Published: (2026)
by: Zhou, Xiang, et al.
Published: (2026)
PLM4NDV: Minimizing Data Access for Number of Distinct Values Estimation with Pre-trained Language Models
by: Xu, Xianghong, et al.
Published: (2025)
by: Xu, Xianghong, et al.
Published: (2025)
Pruning then Reweighting: Towards Data-Efficient Training of Diffusion Models
by: Li, Yize, et al.
Published: (2024)
by: Li, Yize, et al.
Published: (2024)
Towards Efficient Pre-training: Exploring FP4 Precision in Large Language Models
by: Zhou, Jiecheng, et al.
Published: (2025)
by: Zhou, Jiecheng, et al.
Published: (2025)
Leveraging Pre-trained Models for FF-to-FFPE Histopathological Image Translation
by: Zhang, Qilai, et al.
Published: (2024)
by: Zhang, Qilai, et al.
Published: (2024)
Data Mixing Agent: Learning to Re-weight Domains for Continual Pre-training
by: Yang, Kailai, et al.
Published: (2025)
by: Yang, Kailai, et al.
Published: (2025)
Up to Speed on DVD.
by: Crawford, Walt
Published: (1999)
by: Crawford, Walt
Published: (1999)
E-RayZer: Self-supervised 3D Reconstruction as Spatial Visual Pre-training
by: Zhao, Qitao, et al.
Published: (2025)
by: Zhao, Qitao, et al.
Published: (2025)
DanQing: An Up-to-Date Large-Scale Chinese Vision-Language Pre-training Dataset
by: Shen, Hengyu, et al.
Published: (2026)
by: Shen, Hengyu, et al.
Published: (2026)
Efficient Unsupervised Community Search with Pre-trained Graph Transformer
by: Wang, Jianwei, et al.
Published: (2024)
by: Wang, Jianwei, et al.
Published: (2024)
Slight Corruption in Pre-training Data Makes Better Diffusion Models
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
Stationary Reweighting Yields Local Convergence of Soft Fitted Q-Iteration
by: van der Laan, Lars, et al.
Published: (2025)
by: van der Laan, Lars, et al.
Published: (2025)
Thinking Augmented Pre-training
by: Wang, Liang, et al.
Published: (2025)
by: Wang, Liang, et al.
Published: (2025)
Parallel Structures in Pre-training Data Yield In-Context Learning
by: Chen, Yanda, et al.
Published: (2024)
by: Chen, Yanda, et al.
Published: (2024)
From Curated Data to Scalable Models: Continual Pre-training of Dense and MoE Large Language Models for Tibetan
by: Yang, Lei, et al.
Published: (2025)
by: Yang, Lei, et al.
Published: (2025)
Rethinking Data Curation in LLM Training: Online Reweighting Offers Better Generalization than Offline Methods
by: Zhao, Wanru, et al.
Published: (2026)
by: Zhao, Wanru, et al.
Published: (2026)
Pre-training Everywhere: Parameter-Efficient Fine-Tuning for Medical Image Analysis via Target Parameter Pre-training
by: Lei, Xingliang, et al.
Published: (2024)
by: Lei, Xingliang, et al.
Published: (2024)
SoftMCL: Soft Momentum Contrastive Learning for Fine-grained Sentiment-aware Pre-training
by: Wang, Jin, et al.
Published: (2024)
by: Wang, Jin, et al.
Published: (2024)
Similar Items
-
RaphaelRibes/FastDedup: FastDedup v1.1.0
by: Raphaël Ribes
Published: (2026) -
Velocitune: A Velocity-based Dynamic Domain Reweighting Method for Continual Pre-training
by: Luo, Zheheng, et al.
Published: (2024) -
ToReMi: Topic-Aware Data Reweighting for Dynamic Pre-Training Data Selection
by: Zhu, Xiaoxuan, et al.
Published: (2025) -
PM-Dedup: Secure Deduplication with Partial Migration from Cloud to Edge Servers
by: Ke, Zhaokang, et al.
Published: (2025) -
10 Percent Wrong for 90% Done: A Practical Approach to Collection Deduping
by: Hamby, Rogan
Published: (2012)