A Comparative Study of Pre-training and Self-training
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yiheng, Lin, Jiayu, Lin, Zuoquan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An Empirical Study on Context Length for Open-Domain Dialog Generation
by: Shen, Xinyi, et al.
Published: (2024)
by: Shen, Xinyi, et al.
Published: (2024)
Local and Global Contexts for Conversation
by: Lin, Zuoquan, et al.
Published: (2024)
by: Lin, Zuoquan, et al.
Published: (2024)
Incorporating Exponential Smoothing into MLP: A Simple but Effective Sequence Model
by: Chu, Jiqun, et al.
Published: (2024)
by: Chu, Jiqun, et al.
Published: (2024)
Pre-trained Transformer-Based Approach for Arabic Question Answering : A Comparative Study
by: Alsubhi, Kholoud, et al.
Published: (2021)
by: Alsubhi, Kholoud, et al.
Published: (2021)
On Predicting the Post-training Potential of Pre-trained LLMs
by: Li, Xiaoyuan, et al.
Published: (2026)
by: Li, Xiaoyuan, et al.
Published: (2026)
Making Pre-trained Language Models Great on Tabular Prediction
by: Yan, Jiahuan, et al.
Published: (2024)
by: Yan, Jiahuan, et al.
Published: (2024)
Thinking Augmented Pre-training
by: Wang, Liang, et al.
Published: (2025)
by: Wang, Liang, et al.
Published: (2025)
Exploring the Benefit of Activation Sparsity in Pre-training
by: Zhang, Zhengyan, et al.
Published: (2024)
by: Zhang, Zhengyan, et al.
Published: (2024)
ColBERT-Zero: To Pre-train Or Not To Pre-train ColBERT models
by: Chaffin, Antoine, et al.
Published: (2026)
by: Chaffin, Antoine, et al.
Published: (2026)
Domain Pre-training Impact on Representations
by: Gonzalez-Gutierrez, Cesar, et al.
Published: (2025)
by: Gonzalez-Gutierrez, Cesar, et al.
Published: (2025)
Rethinking the Chain-of-Thought: The Roles of In-Context Learning and Pre-trained Priors
by: Yang, Hao, et al.
Published: (2025)
by: Yang, Hao, et al.
Published: (2025)
Enhancing LLM Language Adaption through Cross-lingual In-Context Pre-training
by: Wu, Linjuan, et al.
Published: (2025)
by: Wu, Linjuan, et al.
Published: (2025)
Fine-tuning Pre-trained Language Models for Few-shot Intent Detection: Supervised Pre-training and Isotropization
by: Zhang, Haode, et al.
Published: (2022)
by: Zhang, Haode, et al.
Published: (2022)
Scaling Agents via Continual Pre-training
by: Su, Liangcai, et al.
Published: (2025)
by: Su, Liangcai, et al.
Published: (2025)
Variator: Accelerating Pre-trained Models with Plug-and-Play Compression Modules
by: Xiao, Chaojun, et al.
Published: (2023)
by: Xiao, Chaojun, et al.
Published: (2023)
Pre-trained Language Models for Keyphrase Generation: A Thorough Empirical Study
by: Wu, Di, et al.
Published: (2022)
by: Wu, Di, et al.
Published: (2022)
DataMan: Data Manager for Pre-training Large Language Models
by: Peng, Ru, et al.
Published: (2025)
by: Peng, Ru, et al.
Published: (2025)
TriCon-Fair: Triplet Contrastive Learning for Mitigating Social Bias in Pre-trained Language Models
by: Lyu, Chong, et al.
Published: (2025)
by: Lyu, Chong, et al.
Published: (2025)
Self-training Strategies for Sentiment Analysis: An Empirical Study
by: Liu, Haochen, et al.
Published: (2023)
by: Liu, Haochen, et al.
Published: (2023)
Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text
by: Huang, Chengyu, et al.
Published: (2026)
by: Huang, Chengyu, et al.
Published: (2026)
Synergistic Anchored Contrastive Pre-training for Few-Shot Relation Extraction
by: Luo, Da, et al.
Published: (2023)
by: Luo, Da, et al.
Published: (2023)
Drop your Decoder: Pre-training with Bag-of-Word Prediction for Dense Passage Retrieval
by: Ma, Guangyuan, et al.
Published: (2024)
by: Ma, Guangyuan, et al.
Published: (2024)
Adaptive Layer-skipping in Pre-trained LLMs
by: Luo, Xuan, et al.
Published: (2025)
by: Luo, Xuan, et al.
Published: (2025)
PhoGPT: Generative Pre-training for Vietnamese
by: Nguyen, Dat Quoc, et al.
Published: (2023)
by: Nguyen, Dat Quoc, et al.
Published: (2023)
PART: Pre-trained Authorship Representation Transformer
by: Huertas-Tato, Javier, et al.
Published: (2022)
by: Huertas-Tato, Javier, et al.
Published: (2022)
RegMix: Data Mixture as Regression for Language Model Pre-training
by: Liu, Qian, et al.
Published: (2024)
by: Liu, Qian, et al.
Published: (2024)
InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training
by: Wang, Dingdong, et al.
Published: (2025)
by: Wang, Dingdong, et al.
Published: (2025)
Revealing the Learning Dynamics of Long-Context Continual Pre-training
by: Liang, Yupu, et al.
Published: (2026)
by: Liang, Yupu, et al.
Published: (2026)
MADLLM: Multivariate Anomaly Detection via Pre-trained LLMs
by: Tao, Wei, et al.
Published: (2025)
by: Tao, Wei, et al.
Published: (2025)
Improving Continual Pre-training Through Seamless Data Packing
by: Yin, Ruicheng, et al.
Published: (2025)
by: Yin, Ruicheng, et al.
Published: (2025)
RADAR: Revealing Asymmetric Development of Abilities in MLLM Pre-training
by: Nie, Yunshuang, et al.
Published: (2026)
by: Nie, Yunshuang, et al.
Published: (2026)
Pre-training vs. Fine-tuning: A Reproducibility Study on Dense Retrieval Knowledge Acquisition
by: Yao, Zheng, et al.
Published: (2025)
by: Yao, Zheng, et al.
Published: (2025)
Efficient Continual Pre-training by Mitigating the Stability Gap
by: Guo, Yiduo, et al.
Published: (2024)
by: Guo, Yiduo, et al.
Published: (2024)
FaBERT: Pre-training BERT on Persian Blogs
by: Masumi, Mostafa, et al.
Published: (2024)
by: Masumi, Mostafa, et al.
Published: (2024)
Pre-training and Diagnosing Knowledge Base Completion Models
by: Kocijan, Vid, et al.
Published: (2024)
by: Kocijan, Vid, et al.
Published: (2024)
To Code, or Not To Code? Exploring Impact of Code in Pre-training
by: Aryabumi, Viraat, et al.
Published: (2024)
by: Aryabumi, Viraat, et al.
Published: (2024)
Bag of Lies: Robustness in Continuous Pre-training BERT
by: Gevers, Ine, et al.
Published: (2024)
by: Gevers, Ine, et al.
Published: (2024)
Development of Cognitive Intelligence in Pre-trained Language Models
by: Shah, Raj Sanjay, et al.
Published: (2024)
by: Shah, Raj Sanjay, et al.
Published: (2024)
Summarization of Investment Reports Using Pre-trained Model
by: Sakaji, Hiroki, et al.
Published: (2024)
by: Sakaji, Hiroki, et al.
Published: (2024)
Text-to-Code Generation with Modality-relative Pre-training
by: Christopoulou, Fenia, et al.
Published: (2024)
by: Christopoulou, Fenia, et al.
Published: (2024)
Similar Items
-
An Empirical Study on Context Length for Open-Domain Dialog Generation
by: Shen, Xinyi, et al.
Published: (2024) -
Local and Global Contexts for Conversation
by: Lin, Zuoquan, et al.
Published: (2024) -
Incorporating Exponential Smoothing into MLP: A Simple but Effective Sequence Model
by: Chu, Jiqun, et al.
Published: (2024) -
Pre-trained Transformer-Based Approach for Arabic Question Answering : A Comparative Study
by: Alsubhi, Kholoud, et al.
Published: (2021) -
On Predicting the Post-training Potential of Pre-trained LLMs
by: Li, Xiaoyuan, et al.
Published: (2026)