Early-stopping for Transformer model training
Fuente:
arXiv
Saved in:
| Main Authors: | He, Jing, Jiang, Hua, Li, Cheng, Xin, Siqian, Yang, Shuzhen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Why pre-training is beneficial for downstream classification tasks?
by: Jiang, Xin, et al.
Published: (2024)
by: Jiang, Xin, et al.
Published: (2024)
Graph Generative Pre-trained Transformer
by: Chen, Xiaohui, et al.
Published: (2025)
by: Chen, Xiaohui, et al.
Published: (2025)
The Scaling Law for LoRA Base on Mutual Information Upper Bound
by: Zhang, Jing, et al.
Published: (2025)
by: Zhang, Jing, et al.
Published: (2025)
GraphGPT: Generative Pre-trained Graph Eulerian Transformer
by: Zhao, Qifang, et al.
Published: (2023)
by: Zhao, Qifang, et al.
Published: (2023)
SEE: Sememe Entanglement Encoding for Transformer-bases Models Compression
by: Zhang, Jing, et al.
Published: (2024)
by: Zhang, Jing, et al.
Published: (2024)
Early stopping by correlating online indicators in neural networks
by: Ferro, Manuel Vilares, et al.
Published: (2024)
by: Ferro, Manuel Vilares, et al.
Published: (2024)
TimelyGPT: Extrapolatable Transformer Pre-training for Long-term Time-Series Forecasting in Healthcare
by: Song, Ziyang, et al.
Published: (2023)
by: Song, Ziyang, et al.
Published: (2023)
Constrained multi-fidelity Bayesian optimization with automatic stop condition
by: Foumani, Zahra Zanjani, et al.
Published: (2025)
by: Foumani, Zahra Zanjani, et al.
Published: (2025)
Transformer in Touch: A Survey
by: Gao, Jing, et al.
Published: (2024)
by: Gao, Jing, et al.
Published: (2024)
xTrimoPGLM: Unified 100B-Scale Pre-trained Transformer for Deciphering the Language of Protein
by: Chen, Bo, et al.
Published: (2024)
by: Chen, Bo, et al.
Published: (2024)
TrajGPT-R: Generating Urban Mobility Trajectory with Reinforcement Learning-Enhanced Generative Pre-trained Transformer
by: Wang, Jiawei, et al.
Published: (2026)
by: Wang, Jiawei, et al.
Published: (2026)
A Pre-trained Reaction Embedding Descriptor Capturing Bond Transformation Patterns
by: Liu, Weiqi, et al.
Published: (2026)
by: Liu, Weiqi, et al.
Published: (2026)
Enhancing robustness of data-driven SHM models: adversarial training with circle loss
by: Yang, Xiangli, et al.
Published: (2024)
by: Yang, Xiangli, et al.
Published: (2024)
Don't stop me now: Rethinking Validation Criteria for Model Parameter Selection
by: Apicella, Andrea, et al.
Published: (2026)
by: Apicella, Andrea, et al.
Published: (2026)
Cross-domain Random Pre-training with Prototypes for Reinforcement Learning
by: Liu, Xin, et al.
Published: (2023)
by: Liu, Xin, et al.
Published: (2023)
LENS: Large Pre-trained Transformer for Exploring Financial Time Series Regularities
by: Xu, Yuanjian, et al.
Published: (2024)
by: Xu, Yuanjian, et al.
Published: (2024)
PrefixGPT: Prefix Adder Optimization by a Generative Pre-trained Transformer
by: Ding, Ruogu, et al.
Published: (2025)
by: Ding, Ruogu, et al.
Published: (2025)
MoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts
by: Lin, Xi Victoria, et al.
Published: (2024)
by: Lin, Xi Victoria, et al.
Published: (2024)
Physics-Guided Tiny-Mamba Transformer for Reliability-Aware Early Fault Warning
by: Li, Changyu, et al.
Published: (2026)
by: Li, Changyu, et al.
Published: (2026)
Priming: Hybrid State Space Models From Pre-trained Transformers
by: Chattopadhyay, Aditya, et al.
Published: (2026)
by: Chattopadhyay, Aditya, et al.
Published: (2026)
Self-Improving Pretraining: using post-trained models to pretrain better models
by: Tan, Ellen Xiaoqing, et al.
Published: (2026)
by: Tan, Ellen Xiaoqing, et al.
Published: (2026)
On the Learn-to-Optimize Capabilities of Transformers in In-Context Sparse Recovery
by: Liu, Renpu, et al.
Published: (2024)
by: Liu, Renpu, et al.
Published: (2024)
Layer Specialization Underlying Compositional Reasoning in Transformers
by: Liu, Jing
Published: (2025)
by: Liu, Jing
Published: (2025)
Graph Tokenization for Bridging Graphs and Transformers
by: Guo, Zeyuan, et al.
Published: (2026)
by: Guo, Zeyuan, et al.
Published: (2026)
Leveraging 2D Information for Long-term Time Series Forecasting with Vanilla Transformers
by: Cheng, Xin, et al.
Published: (2024)
by: Cheng, Xin, et al.
Published: (2024)
Bidirectional Generative Pre-training for Improving Healthcare Time-series Representation Learning
by: Song, Ziyang, et al.
Published: (2024)
by: Song, Ziyang, et al.
Published: (2024)
Are Transformers More Robust? Towards Exact Robustness Verification for Transformers
by: Liao, Brian Hsuan-Cheng, et al.
Published: (2022)
by: Liao, Brian Hsuan-Cheng, et al.
Published: (2022)
Heterogeneous Graph Pre-training Based Model for Secure and Efficient Prediction of Default Risk Propagation among Bond Issuers
by: Li, Xurui, et al.
Published: (2025)
by: Li, Xurui, et al.
Published: (2025)
Zero-shot Meta-learning for Tabular Prediction Tasks with Adversarially Pre-trained Transformer
by: Wu, Yulun, et al.
Published: (2025)
by: Wu, Yulun, et al.
Published: (2025)
Causal prompting model-based offline reinforcement learning
by: Yu, Xuehui, et al.
Published: (2024)
by: Yu, Xuehui, et al.
Published: (2024)
Task-Agnostic Pre-training and Task-Guided Fine-tuning for Versatile Diffusion Planner
by: Fan, Chenyou, et al.
Published: (2024)
by: Fan, Chenyou, et al.
Published: (2024)
One for All: A Non-Linear Transformer can Enable Cross-Domain Generalization for In-Context Reinforcement Learning
by: He, Bowen, et al.
Published: (2026)
by: He, Bowen, et al.
Published: (2026)
Traj-Transformer: Diffusion Models with Transformer for GPS Trajectory Generation
by: Zhang, Zhiyang, et al.
Published: (2025)
by: Zhang, Zhiyang, et al.
Published: (2025)
Pre-trained Language Models Improve the Few-shot Prompt Ability of Decision Transformer
by: Yang, Yu, et al.
Published: (2024)
by: Yang, Yu, et al.
Published: (2024)
Transformers trained on proteins can learn to attend to Euclidean distance
by: Ellmen, Isaac, et al.
Published: (2025)
by: Ellmen, Isaac, et al.
Published: (2025)
Reinforcement Learning With Sparse-Executing Actions via Sparsity Regularization
by: Pang, Jing-Cheng, et al.
Published: (2021)
by: Pang, Jing-Cheng, et al.
Published: (2021)
pUniFind: a unified large pre-trained deep learning model pushing the limit of mass spectra interpretation
by: Zhao, Jiale, et al.
Published: (2025)
by: Zhao, Jiale, et al.
Published: (2025)
Decoupling Weighing and Selecting for Integrating Multiple Graph Pre-training Tasks
by: Fan, Tianyu, et al.
Published: (2024)
by: Fan, Tianyu, et al.
Published: (2024)
TripCast: Pre-training of Masked 2D Transformers for Trip Time Series Forecasting
by: Liao, Yuhua, et al.
Published: (2024)
by: Liao, Yuhua, et al.
Published: (2024)
Graph VQ-Transformer (GVT): Fast and Accurate Molecular Generation via High-Fidelity Discrete Latents
by: Zheng, Haozhuo, et al.
Published: (2025)
by: Zheng, Haozhuo, et al.
Published: (2025)
Similar Items
-
Why pre-training is beneficial for downstream classification tasks?
by: Jiang, Xin, et al.
Published: (2024) -
Graph Generative Pre-trained Transformer
by: Chen, Xiaohui, et al.
Published: (2025) -
The Scaling Law for LoRA Base on Mutual Information Upper Bound
by: Zhang, Jing, et al.
Published: (2025) -
GraphGPT: Generative Pre-trained Graph Eulerian Transformer
by: Zhao, Qifang, et al.
Published: (2023) -
SEE: Sememe Entanglement Encoding for Transformer-bases Models Compression
by: Zhang, Jing, et al.
Published: (2024)