Harnessing Diversity for Important Data Selection in Pretraining Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Chi, Zhong, Huaping, Zhang, Kuan, Chai, Chengliang, Wang, Rui, Zhuang, Xinlin, Bai, Tianyi, Qiu, Jiantao, Cao, Lei, Fan, Ju, Yuan, Ye, Wang, Guoren, He, Conghui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration
von: Bai, Tianyi, et al.
Veröffentlicht: (2024)
von: Bai, Tianyi, et al.
Veröffentlicht: (2024)
Not All Documents Are What You Need for Extracting Instruction Tuning Data
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models
von: Zhuang, Xinlin, et al.
Veröffentlicht: (2025)
von: Zhuang, Xinlin, et al.
Veröffentlicht: (2025)
Handling Label Noise via Instance-Level Difficulty Modeling and Dynamic Optimization
von: Zhang, Kuan, et al.
Veröffentlicht: (2025)
von: Zhang, Kuan, et al.
Veröffentlicht: (2025)
Topic Over Source: The Key to Effective Data Mixing for Language Models Pre-training
von: Peng, Jiahui, et al.
Veröffentlicht: (2025)
von: Peng, Jiahui, et al.
Veröffentlicht: (2025)
VADE: Variance-Aware Dynamic Sampling via Online Sample-Level Difficulty Estimation for Multimodal RL
von: Hu, Zengjie, et al.
Veröffentlicht: (2025)
von: Hu, Zengjie, et al.
Veröffentlicht: (2025)
QUEST: Query Optimization in Unstructured Document Analysis
von: Sun, Zhaoze, et al.
Veröffentlicht: (2025)
von: Sun, Zhaoze, et al.
Veröffentlicht: (2025)
AgenticOCR: Parsing Only What You Need for Efficient Retrieval-Augmented Generation
von: Wang, Zhengren, et al.
Veröffentlicht: (2026)
von: Wang, Zhengren, et al.
Veröffentlicht: (2026)
Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning
von: Bai, Tianyi, et al.
Veröffentlicht: (2025)
von: Bai, Tianyi, et al.
Veröffentlicht: (2025)
Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification
von: Bai, Tianyi, et al.
Veröffentlicht: (2025)
von: Bai, Tianyi, et al.
Veröffentlicht: (2025)
QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining
von: Liu, Fengze, et al.
Veröffentlicht: (2025)
von: Liu, Fengze, et al.
Veröffentlicht: (2025)
Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More
von: Wen, Zichen, et al.
Veröffentlicht: (2025)
von: Wen, Zichen, et al.
Veröffentlicht: (2025)
SG-BEV: Satellite-Guided BEV Fusion for Cross-View Semantic Segmentation
von: Ye, Junyan, et al.
Veröffentlicht: (2024)
von: Ye, Junyan, et al.
Veröffentlicht: (2024)
Cross-view image geo-localization with Panorama-BEV Co-Retrieval Network
von: Ye, Junyan, et al.
Veröffentlicht: (2024)
von: Ye, Junyan, et al.
Veröffentlicht: (2024)
PACE: Poisoning Attacks on Learned Cardinality Estimation
von: Zhang, Jintao, et al.
Veröffentlicht: (2024)
von: Zhang, Jintao, et al.
Veröffentlicht: (2024)
AutoCE: An Accurate and Efficient Model Advisor for Learned Cardinality Estimation
von: Zhang, Jintao, et al.
Veröffentlicht: (2024)
von: Zhang, Jintao, et al.
Veröffentlicht: (2024)
Influence Maximization in Hypergraphs by Stratified Sampling for Efficient Generation of Reverse Reachable Sets
von: Zhang, Lingling, et al.
Veröffentlicht: (2024)
von: Zhang, Lingling, et al.
Veröffentlicht: (2024)
Unstructured Data Analysis using LLMs: A Comprehensive Benchmark
von: Deng, Qiyan, et al.
Veröffentlicht: (2025)
von: Deng, Qiyan, et al.
Veröffentlicht: (2025)
KeyVideoLLM: Towards Large-scale Video Keyframe Selection
von: Liang, Hao, et al.
Veröffentlicht: (2024)
von: Liang, Hao, et al.
Veröffentlicht: (2024)
StreamTGN: A GPU-Efficient Serving System for Streaming Temporal Graph Neural Networks
von: Zhang, Lingling, et al.
Veröffentlicht: (2026)
von: Zhang, Lingling, et al.
Veröffentlicht: (2026)
CrossViewDiff: A Cross-View Diffusion Model for Satellite-to-Street View Synthesis
von: Li, Weijia, et al.
Veröffentlicht: (2024)
von: Li, Weijia, et al.
Veröffentlicht: (2024)
Graph-Based Feature Augmentation for Predictive Tasks on Relational Datasets
von: Qiao, Lianpeng, et al.
Veröffentlicht: (2025)
von: Qiao, Lianpeng, et al.
Veröffentlicht: (2025)
An Empirical Study on Influence-Based Pretraining Data Selection for Code Large Language Models
von: Xing, Chengli, et al.
Veröffentlicht: (2026)
von: Xing, Chengli, et al.
Veröffentlicht: (2026)
DIffSteISR: Harnessing Diffusion Prior for Superior Real-world Stereo Image Super-Resolution
von: Zhou, Yuanbo, et al.
Veröffentlicht: (2024)
von: Zhou, Yuanbo, et al.
Veröffentlicht: (2024)
Not All Instances Are Equally Valuable: Towards Influence-Weighted Dataset Distillation
von: Deng, Qiyan, et al.
Veröffentlicht: (2025)
von: Deng, Qiyan, et al.
Veröffentlicht: (2025)
P‐122: Dependence of Aerial image quality on Display parameters in aerial Display system
von: Xinlin Ye, et al.
Veröffentlicht: (2024)
von: Xinlin Ye, et al.
Veröffentlicht: (2024)
Multivariate Time Series Cleaning under Speed Constraints
von: Zhang, Aoqian, et al.
Veröffentlicht: (2024)
von: Zhang, Aoqian, et al.
Veröffentlicht: (2024)
Matryoshka: Optimization of Dynamic Diverse Quantum Chemistry Systems via Elastic Parallelism Transformation
von: Wang, Tuowei, et al.
Veröffentlicht: (2024)
von: Wang, Tuowei, et al.
Veröffentlicht: (2024)
Multiomics Analyses Demonstrate the Attenuation of Metabolic Cardiac Disorders Associated With Type 2 Diabetes by Stachydrine in Relation With the Transition of Gastrointestinal Microbiota
von: Chaoxing Yang, et al.
Veröffentlicht: (2025)
von: Chaoxing Yang, et al.
Veröffentlicht: (2025)
An Aggregation‐Induced Emission Active Peptide‐Based Fluorescent Probe for Highly Selective and Sensitive Detection of Hg(II) Ions and Its Multifield Applications
von: Shiyi Xiong, et al.
Veröffentlicht: (2025)
von: Shiyi Xiong, et al.
Veröffentlicht: (2025)
Diverse polymorphs and phase transitions in van der Waals In$_2$Se$_3$
von: Liu, Mingfeng, et al.
Veröffentlicht: (2025)
von: Liu, Mingfeng, et al.
Veröffentlicht: (2025)
Class-Imbalanced-Aware Adaptive Dataset Distillation for Scalable Pretrained Model on Credit Scoring
von: Li, Xia, et al.
Veröffentlicht: (2025)
von: Li, Xia, et al.
Veröffentlicht: (2025)
BLISS: A Lightweight Bilevel Influence Scoring Method for Data Selection in Language Model Pretraining
von: Hao, Jie, et al.
Veröffentlicht: (2025)
von: Hao, Jie, et al.
Veröffentlicht: (2025)
The Ad‐Micellar Preparation of Para‐Aramid Nanofiber/Polyacrylate Nanocomposite
von: Mengyu Zhang, et al.
Veröffentlicht: (2025)
von: Mengyu Zhang, et al.
Veröffentlicht: (2025)
VIGC: Visual Instruction Generation and Correction
von: Wang, Bin, et al.
Veröffentlicht: (2023)
von: Wang, Bin, et al.
Veröffentlicht: (2023)
IKnow: Instruction-Knowledge-Aware Continual Pretraining for Effective Domain Adaptation
von: Zhang, Tianyi, et al.
Veröffentlicht: (2025)
von: Zhang, Tianyi, et al.
Veröffentlicht: (2025)
Macformer: Transformer with Random Maclaurin Feature Attention
von: Guo, Yuhan, et al.
Veröffentlicht: (2024)
von: Guo, Yuhan, et al.
Veröffentlicht: (2024)
EvoGymCM: Harnessing Continuous Material Stiffness for Soft Robot Co-Design
von: Shen, Le, et al.
Veröffentlicht: (2026)
von: Shen, Le, et al.
Veröffentlicht: (2026)
Scalable $k$-clique Densest Subgraph Search
von: Ye, Xiaowei, et al.
Veröffentlicht: (2024)
von: Ye, Xiaowei, et al.
Veröffentlicht: (2024)
Harnessing Inherent Noises for Privacy Preservation in Quantum Machine Learning
von: Ju, Keyi, et al.
Veröffentlicht: (2023)
von: Ju, Keyi, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration
von: Bai, Tianyi, et al.
Veröffentlicht: (2024) -
Not All Documents Are What You Need for Extracting Instruction Tuning Data
von: Zhang, Chi, et al.
Veröffentlicht: (2025) -
Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models
von: Zhuang, Xinlin, et al.
Veröffentlicht: (2025) -
Handling Label Noise via Instance-Level Difficulty Modeling and Dynamic Optimization
von: Zhang, Kuan, et al.
Veröffentlicht: (2025) -
Topic Over Source: The Key to Effective Data Mixing for Language Models Pre-training
von: Peng, Jiahui, et al.
Veröffentlicht: (2025)