Combatting Dimensional Collapse in LLM Pre-Training Data via Diversified File Selection
Fuente:
arXiv
Saved in:
| Main Authors: | Fan, Ziqing, Du, Siyuan, Hu, Shengchao, Wang, Pingjie, Shen, Li, Zhang, Ya, Tao, Dacheng, Wang, Yanfeng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HarmoDT: Harmony Multi-Task Decision Transformer for Offline Reinforcement Learning
by: Hu, Shengchao, et al.
Published: (2024)
by: Hu, Shengchao, et al.
Published: (2024)
Reconstruct the Pruned Model without Any Retraining
by: Wang, Pingjie, et al.
Published: (2024)
by: Wang, Pingjie, et al.
Published: (2024)
Q-value Regularized Transformer for Offline Reinforcement Learning
by: Hu, Shengchao, et al.
Published: (2024)
by: Hu, Shengchao, et al.
Published: (2024)
Task-Aware Harmony Multi-Task Decision Transformer for Offline Reinforcement Learning
by: Fan, Ziqing, et al.
Published: (2024)
by: Fan, Ziqing, et al.
Published: (2024)
Diversified Batch Selection for Training Acceleration
by: Hong, Feng, et al.
Published: (2024)
by: Hong, Feng, et al.
Published: (2024)
Continual Task Learning through Adaptive Policy Self-Composition
by: Hu, Shengchao, et al.
Published: (2024)
by: Hu, Shengchao, et al.
Published: (2024)
Learning Multi-Agent Communication from Graph Modeling Perspective
by: Hu, Shengchao, et al.
Published: (2024)
by: Hu, Shengchao, et al.
Published: (2024)
Communication Learning in Multi-Agent Systems from Graph Modeling Perspective
by: Hu, Shengchao, et al.
Published: (2024)
by: Hu, Shengchao, et al.
Published: (2024)
Prompt Tuning with Diffusion for Few-Shot Pre-trained Policy Generalization
by: Hu, Shengchao, et al.
Published: (2024)
by: Hu, Shengchao, et al.
Published: (2024)
Selecting Auxiliary Data via Neural Tangent Kernels for Low-Resource Domains
by: Wang, Pingjie, et al.
Published: (2025)
by: Wang, Pingjie, et al.
Published: (2025)
Joint Selection for Large-Scale Pre-Training Data via Policy Gradient-based Mask Learning
by: Fan, Ziqing, et al.
Published: (2025)
by: Fan, Ziqing, et al.
Published: (2025)
SeWA: Selective Weight Average via Probabilistic Masking
by: Wang, Peng, et al.
Published: (2025)
by: Wang, Peng, et al.
Published: (2025)
Locally Estimated Global Perturbations are Better than Local Perturbations for Federated Sharpness-aware Minimization
by: Fan, Ziqing, et al.
Published: (2024)
by: Fan, Ziqing, et al.
Published: (2024)
Federated Learning under Partially Class-Disjoint Data via Manifold Reshaping
by: Fan, Ziqing, et al.
Published: (2024)
by: Fan, Ziqing, et al.
Published: (2024)
Rethinking the Role of Dynamic Sparse Training for Scalable Deep Reinforcement Learning
by: Ma, Guozheng, et al.
Published: (2025)
by: Ma, Guozheng, et al.
Published: (2025)
Federated Learning with Bilateral Curation for Partially Class-Disjoint Data
by: Fan, Ziqing, et al.
Published: (2024)
by: Fan, Ziqing, et al.
Published: (2024)
Domain-Inspired Sharpness-Aware Minimization Under Domain Shifts
by: Zhang, Ruipeng, et al.
Published: (2024)
by: Zhang, Ruipeng, et al.
Published: (2024)
LLM Data Selection and Utilization via Dynamic Bi-level Optimization
by: Yu, Yang, et al.
Published: (2025)
by: Yu, Yang, et al.
Published: (2025)
A Theoretical Perspective: How to Prevent Model Collapse in Self-consuming Training Loops
by: Fu, Shi, et al.
Published: (2025)
by: Fu, Shi, et al.
Published: (2025)
Task Groupings Regularization: Data-Free Meta-Learning with Heterogeneous Pre-trained Models
by: Wei, Yongxian, et al.
Published: (2024)
by: Wei, Yongxian, et al.
Published: (2024)
Solving Continual Offline RL through Selective Weights Activation on Aligned Spaces
by: Hu, Jifeng, et al.
Published: (2024)
by: Hu, Jifeng, et al.
Published: (2024)
SMILE: Zero-Shot Sparse Mixture of Low-Rank Experts Construction From Pre-Trained Foundation Models
by: Tang, Anke, et al.
Published: (2024)
by: Tang, Anke, et al.
Published: (2024)
Adaptive Defense against Harmful Fine-Tuning for Large Language Models via Bayesian Data Scheduler
by: Hu, Zixuan, et al.
Published: (2025)
by: Hu, Zixuan, et al.
Published: (2025)
Offline Behavioral Data Selection
by: Lei, Shiye, et al.
Published: (2025)
by: Lei, Shiye, et al.
Published: (2025)
FOAM: Blocked State Folding for Memory-Efficient LLM Training
by: Wen, Ziqing, et al.
Published: (2025)
by: Wen, Ziqing, et al.
Published: (2025)
Near-Oracle KV Selection via Pre-hoc Sparsity for Long-Context Inference
by: Gao, Yifei, et al.
Published: (2026)
by: Gao, Yifei, et al.
Published: (2026)
From Data to Action: Charting A Data-Driven Path to Combat Antimicrobial Resistance
by: Fu, Qian, et al.
Published: (2025)
by: Fu, Qian, et al.
Published: (2025)
OpenFedLLM: Training Large Language Models on Decentralized Private Data via Federated Learning
by: Ye, Rui, et al.
Published: (2024)
by: Ye, Rui, et al.
Published: (2024)
M2K-VDG: Model-Adaptive Multimodal Knowledge Anchor Enhanced Video-grounded Dialogue Generation
by: Liu, Hongcheng, et al.
Published: (2024)
by: Liu, Hongcheng, et al.
Published: (2024)
RAD: Towards Trustworthy Retrieval-Augmented Multi-modal Clinical Diagnosis
by: Li, Haolin, et al.
Published: (2025)
by: Li, Haolin, et al.
Published: (2025)
Analytic Energy-Guided Policy Optimization for Offline Reinforcement Learning
by: Hu, Jifeng, et al.
Published: (2025)
by: Hu, Jifeng, et al.
Published: (2025)
On exploring the potential of quantum auto-encoder for learning quantum systems
by: Du, Yuxuan, et al.
Published: (2021)
by: Du, Yuxuan, et al.
Published: (2021)
Unlocking Tuning-Free Few-Shot Adaptability in Visual Foundation Models by Recycling Pre-Tuned LoRAs
by: Hu, Zixuan, et al.
Published: (2024)
by: Hu, Zixuan, et al.
Published: (2024)
Revisiting Plasticity in Visual Reinforcement Learning: Data, Modules and Training Stages
by: Ma, Guozheng, et al.
Published: (2023)
by: Ma, Guozheng, et al.
Published: (2023)
Preventing Dimensional Collapse in Self-Supervised Learning via Orthogonality Regularization
by: He, Junlin, et al.
Published: (2024)
by: He, Junlin, et al.
Published: (2024)
FREE: Faster and Better Data-Free Meta-Learning
by: Wei, Yongxian, et al.
Published: (2024)
by: Wei, Yongxian, et al.
Published: (2024)
Accelerating LLM Pre-Training through Flat-Direction Dynamics Enhancement
by: Zhu, Shuchen, et al.
Published: (2026)
by: Zhu, Shuchen, et al.
Published: (2026)
Scaling Adversarial Training via Data Selection
by: Ye, Youran, et al.
Published: (2025)
by: Ye, Youran, et al.
Published: (2025)
Rethinking Data Curation in LLM Training: Online Reweighting Offers Better Generalization than Offline Methods
by: Zhao, Wanru, et al.
Published: (2026)
by: Zhao, Wanru, et al.
Published: (2026)
TimeGuard: Channel-wise Pool Training for Backdoor Defense in Time Series Forecasting
by: Nguyen, Quang Duc, et al.
Published: (2026)
by: Nguyen, Quang Duc, et al.
Published: (2026)
Similar Items
-
HarmoDT: Harmony Multi-Task Decision Transformer for Offline Reinforcement Learning
by: Hu, Shengchao, et al.
Published: (2024) -
Reconstruct the Pruned Model without Any Retraining
by: Wang, Pingjie, et al.
Published: (2024) -
Q-value Regularized Transformer for Offline Reinforcement Learning
by: Hu, Shengchao, et al.
Published: (2024) -
Task-Aware Harmony Multi-Task Decision Transformer for Offline Reinforcement Learning
by: Fan, Ziqing, et al.
Published: (2024) -
Diversified Batch Selection for Training Acceleration
by: Hong, Feng, et al.
Published: (2024)