Improving Continual Pre-training Through Seamless Data Packing
Fuente:
arXiv
Guardado en:
| Autores principales: | Yin, Ruicheng, Gao, Xuan, Lv, Changze, Wang, Xiaohua, Zheng, Xiaoqing, Huang, Xuanjing |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Decoding Continuous Character-based Language from Non-invasive Brain Recordings
por: Zhang, Cenyuan, et al.
Publicado: (2024)
por: Zhang, Cenyuan, et al.
Publicado: (2024)
Explainable Synthetic Image Detection through Diffusion Timestep Ensembling
por: Wu, Yixin, et al.
Publicado: (2025)
por: Wu, Yixin, et al.
Publicado: (2025)
Searching for Best Practices in Retrieval-Augmented Generation
por: Wang, Xiaohua, et al.
Publicado: (2024)
por: Wang, Xiaohua, et al.
Publicado: (2024)
UPLex: Fine-Grained Personality Control in Large Language Models via Unsupervised Lexical Modulation
por: Li, Tianlong, et al.
Publicado: (2023)
por: Li, Tianlong, et al.
Publicado: (2023)
Revisiting Jailbreaking for Large Language Models: A Representation Engineering Perspective
por: Li, Tianlong, et al.
Publicado: (2024)
por: Li, Tianlong, et al.
Publicado: (2024)
Spiking Convolutional Neural Networks for Text Classification
por: Lv, Changze, et al.
Publicado: (2024)
por: Lv, Changze, et al.
Publicado: (2024)
Aligning Large Language Models with Human Preferences through Representation Engineering
por: Liu, Wenhao, et al.
Publicado: (2023)
por: Liu, Wenhao, et al.
Publicado: (2023)
Promoting Data and Model Privacy in Federated Learning through Quantized LoRA
por: Zhu, JianHao, et al.
Publicado: (2024)
por: Zhu, JianHao, et al.
Publicado: (2024)
Layer-Specific Scaling of Positional Encodings for Superior Long-Context Modeling
por: Wang, Zhenghua, et al.
Publicado: (2025)
por: Wang, Zhenghua, et al.
Publicado: (2025)
Advancing Parameter Efficiency in Fine-tuning via Representation Editing
por: Wu, Muling, et al.
Publicado: (2024)
por: Wu, Muling, et al.
Publicado: (2024)
SpikeBERT: A Language Spikformer Learned from BERT with Knowledge Distillation
por: Lv, Changze, et al.
Publicado: (2023)
por: Lv, Changze, et al.
Publicado: (2023)
Beyond Single Labels: Improving Conversational Recommendation through LLM-Powered Data Augmentation
por: Xu, Haozhe, et al.
Publicado: (2025)
por: Xu, Haozhe, et al.
Publicado: (2025)
VIB-Probe: Detecting and Mitigating Hallucinations in Vision-Language Models via Variational Information Bottleneck
por: Zhang, Feiran, et al.
Publicado: (2026)
por: Zhang, Feiran, et al.
Publicado: (2026)
Progressive Mastery: Customized Curriculum Learning with Guided Prompting for Mathematical Reasoning
por: Wu, Muling, et al.
Publicado: (2025)
por: Wu, Muling, et al.
Publicado: (2025)
SpikeCLIP: A Contrastive Language-Image Pretrained Spiking Neural Network
por: Lv, Changze, et al.
Publicado: (2023)
por: Lv, Changze, et al.
Publicado: (2023)
Enhancing the Capability and Robustness of Large Language Models through Reinforcement Learning-Driven Query Refinement
por: Wang, Xiaohua, et al.
Publicado: (2024)
por: Wang, Xiaohua, et al.
Publicado: (2024)
Benchmark^2: Systematic Evaluation of LLM Benchmarks
por: Qian, Qi, et al.
Publicado: (2026)
por: Qian, Qi, et al.
Publicado: (2026)
Learning Query-Specific Rubrics from Human Preferences for DeepResearch Report Generation
por: Lv, Changze, et al.
Publicado: (2026)
por: Lv, Changze, et al.
Publicado: (2026)
Advancing Spiking Neural Networks for Sequential Modeling with Central Pattern Generators
por: Lv, Changze, et al.
Publicado: (2024)
por: Lv, Changze, et al.
Publicado: (2024)
Efficient and Effective Time-Series Forecasting with Spiking Neural Networks
por: Lv, Changze, et al.
Publicado: (2024)
por: Lv, Changze, et al.
Publicado: (2024)
Scaling Agents via Continual Pre-training
por: Su, Liangcai, et al.
Publicado: (2025)
por: Su, Liangcai, et al.
Publicado: (2025)
CSSG: Measuring Code Similarity with Semantic Graphs
por: Lu, Yiyang, et al.
Publicado: (2026)
por: Lu, Yiyang, et al.
Publicado: (2026)
Revealing the Learning Dynamics of Long-Context Continual Pre-training
por: Liang, Yupu, et al.
Publicado: (2026)
por: Liang, Yupu, et al.
Publicado: (2026)
Toward Relative Positional Encoding in Spiking Transformers
por: Lv, Changze, et al.
Publicado: (2025)
por: Lv, Changze, et al.
Publicado: (2025)
SATER: A Self-Aware and Token-Efficient Approach to Routing and Cascading
por: Shen, Yuanzhe, et al.
Publicado: (2025)
por: Shen, Yuanzhe, et al.
Publicado: (2025)
TripTailor: A Real-World Benchmark for Personalized Travel Planning
por: Shen, Yuanzhe, et al.
Publicado: (2025)
por: Shen, Yuanzhe, et al.
Publicado: (2025)
Adaptive Layer-skipping in Pre-trained LLMs
por: Luo, Xuan, et al.
Publicado: (2025)
por: Luo, Xuan, et al.
Publicado: (2025)
Dendritic Localized Learning: Toward Biologically Plausible Algorithm
por: Lv, Changze, et al.
Publicado: (2025)
por: Lv, Changze, et al.
Publicado: (2025)
Biologically Plausible Learning via Bidirectional Spike-Based Distillation
por: Lv, Changze, et al.
Publicado: (2025)
por: Lv, Changze, et al.
Publicado: (2025)
Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data
por: Guo, Xu, et al.
Publicado: (2026)
por: Guo, Xu, et al.
Publicado: (2026)
OPUS: Towards Efficient and Principled Data Selection in Large Language Model Pre-training in Every Iteration
por: Wang, Shaobo, et al.
Publicado: (2026)
por: Wang, Shaobo, et al.
Publicado: (2026)
Efficient Continual Pre-training by Mitigating the Stability Gap
por: Guo, Yiduo, et al.
Publicado: (2024)
por: Guo, Yiduo, et al.
Publicado: (2024)
Bag of Lies: Robustness in Continuous Pre-training BERT
por: Gevers, Ine, et al.
Publicado: (2024)
por: Gevers, Ine, et al.
Publicado: (2024)
CURP: Codebook-based Continuous User Representation for Personalized Generation with LLMs
por: Wang, Liang, et al.
Publicado: (2026)
por: Wang, Liang, et al.
Publicado: (2026)
Design as Desired: Utilizing Visual Question Answering for Multimodal Pre-training
por: Su, Tongkun, et al.
Publicado: (2024)
por: Su, Tongkun, et al.
Publicado: (2024)
Perplexity-Aware Data Scaling Law: Perplexity Landscapes Predict Performance for Continual Pre-training
por: Liu, Lei, et al.
Publicado: (2025)
por: Liu, Lei, et al.
Publicado: (2025)
AfriqueLLM: How Data Mixing and Model Architecture Impact Continued Pre-training for African Languages
por: Yu, Hao, et al.
Publicado: (2026)
por: Yu, Hao, et al.
Publicado: (2026)
From Unfamiliar to Familiar: Detecting Pre-training Data via Gradient Deviations in Large Language Models
por: Zhang, Ruiqi, et al.
Publicado: (2026)
por: Zhang, Ruiqi, et al.
Publicado: (2026)
Thinking Augmented Pre-training
por: Wang, Liang, et al.
Publicado: (2025)
por: Wang, Liang, et al.
Publicado: (2025)
Enhancing Question Answering on Charts Through Effective Pre-training Tasks
por: Gupta, Ashim, et al.
Publicado: (2024)
por: Gupta, Ashim, et al.
Publicado: (2024)
Ejemplares similares
-
Decoding Continuous Character-based Language from Non-invasive Brain Recordings
por: Zhang, Cenyuan, et al.
Publicado: (2024) -
Explainable Synthetic Image Detection through Diffusion Timestep Ensembling
por: Wu, Yixin, et al.
Publicado: (2025) -
Searching for Best Practices in Retrieval-Augmented Generation
por: Wang, Xiaohua, et al.
Publicado: (2024) -
UPLex: Fine-Grained Personality Control in Large Language Models via Unsupervised Lexical Modulation
por: Li, Tianlong, et al.
Publicado: (2023) -
Revisiting Jailbreaking for Large Language Models: A Representation Engineering Perspective
por: Li, Tianlong, et al.
Publicado: (2024)