First-Passage Prediction of Grokking Delay: ACalibrated Law under AdamW with Causal Validation
Fuente:
arXiv
Saved in:
| Main Authors: | Khanh, Truong Xuan, Hoa, Truong Quynh, Trung, Luu Duc, Duc, Phan Thanh |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Norm-Separation Delay Law of Grokking: A First-Principles Theory of Delayed Generalization
by: Khanh, Truong Xuan, et al.
Published: (2026)
by: Khanh, Truong Xuan, et al.
Published: (2026)
Spectral Entropy Collapse as a Phase Transition in Delayed Generalisation: An Interventional and Predictive Framework for Grokkin
by: Khanh, Truong Xuan, et al.
Published: (2026)
by: Khanh, Truong Xuan, et al.
Published: (2026)
Entropy Collapse: A Universal Failure Mode of Intelligent Systems
by: Khanh, Truong Xuan, et al.
Published: (2025)
by: Khanh, Truong Xuan, et al.
Published: (2025)
Intelligence as Trajectory-Dominant Pareto Optimization
by: Khanh, Truong Xuan, et al.
Published: (2026)
by: Khanh, Truong Xuan, et al.
Published: (2026)
Dynamic Intelligence Ceilings: Measuring Long-Horizon Limits of Planning and Creativity in Artificial Systems
by: Khanh, Truong Xuan, et al.
Published: (2026)
by: Khanh, Truong Xuan, et al.
Published: (2026)
Norm-Hierarchy Transitions in Representation Learning: When and Why Neural Networks Abandon Shortcuts
by: Khanh, Truong Xuan, et al.
Published: (2026)
by: Khanh, Truong Xuan, et al.
Published: (2026)
Detecting Neurovascular Instability from Multimodal Physiological Signals Using Wearable-Compatible Edge AI: A Responsible Computational Framework
by: Hoa, Truong Quynh, et al.
Published: (2026)
by: Hoa, Truong Quynh, et al.
Published: (2026)
Phase Transitions in Driven Informational Systems: A Two-Field Perspective on Learning Theory and Non-Equilibrium Chemistry
by: Khanh, Truong Xuan
Published: (2026)
by: Khanh, Truong Xuan
Published: (2026)
Beyond First-Order: Training LLMs with Stochastic Conjugate Subgradients and AdamW
by: Zhang, Di, et al.
Published: (2025)
by: Zhang, Di, et al.
Published: (2025)
Solving Cyclic Antibandwidth Problem by SAT
by: Xuan, Hieu Truong, et al.
Published: (2026)
by: Xuan, Hieu Truong, et al.
Published: (2026)
APOLLO: SGD-like Memory, AdamW-level Performance
by: Zhu, Hanqing, et al.
Published: (2024)
by: Zhu, Hanqing, et al.
Published: (2024)
Optimizer-Induced Mode Connectivity: From AdamW to Muon
by: Zhang, Fangzhao, et al.
Published: (2026)
by: Zhang, Fangzhao, et al.
Published: (2026)
How to set AdamW's weight decay as you scale model and dataset size
by: Wang, Xi, et al.
Published: (2024)
by: Wang, Xi, et al.
Published: (2024)
Navigating LLM Valley: From AdamW to Memory-Efficient and Matrix-Based Optimizers
by: Ranganath, Aditya
Published: (2026)
by: Ranganath, Aditya
Published: (2026)
BIMA: Bijective Maximum Likelihood Learning Approach to Hallucination Prediction and Mitigation in Large Vision-Language Models
by: Tran, Huu-Thien, et al.
Published: (2025)
by: Tran, Huu-Thien, et al.
Published: (2025)
Efficient and Concise Explanations for Object Detection with Gaussian-Class Activation Mapping Explainer
by: Nguyen, Quoc Khanh, et al.
Published: (2024)
by: Nguyen, Quoc Khanh, et al.
Published: (2024)
BRAIN: Bias-Mitigation Continual Learning Approach to Vision-Brain Understanding
by: Nguyen, Xuan-Bac, et al.
Published: (2025)
by: Nguyen, Xuan-Bac, et al.
Published: (2025)
SKGE: Spherical Knowledge Graph Embedding with Geometric Regularization
by: Quan, Xuan-Truong, et al.
Published: (2025)
by: Quan, Xuan-Truong, et al.
Published: (2025)
Optical Quantum Mixed-State Reconstruction With Multiple Deep Learning Approaches
by: Luu, Nhan Trong, et al.
Published: (2024)
by: Luu, Nhan Trong, et al.
Published: (2024)
RobustFSM: Submodular Maximization in Federated Setting with Malicious Clients
by: Tran, Duc A., et al.
Published: (2025)
by: Tran, Duc A., et al.
Published: (2025)
A Vietnamese Dataset for Text Segmentation and Multiple Choices Reading Comprehension
by: Hai, Toan Nguyen, et al.
Published: (2025)
by: Hai, Toan Nguyen, et al.
Published: (2025)
Auto-Prompting with Retrieval Guidance for Frame Detection in Logistics
by: Duc, Do Minh, et al.
Published: (2025)
by: Duc, Do Minh, et al.
Published: (2025)
AgriKD: Cross-Architecture Knowledge Distillation for Efficient Leaf Disease Classification
by: Le, Minh-Dung, et al.
Published: (2026)
by: Le, Minh-Dung, et al.
Published: (2026)
Accuracy-Robustness Trade Off via Spiking Neural Network Gradient Sparsity Trail
by: Nhan, Luu Trong, et al.
Published: (2025)
by: Nhan, Luu Trong, et al.
Published: (2025)
A Concept is More Than a Word: Diversified Unlearning in Text-to-Image Diffusion Models
by: Pham, Duc Hao, et al.
Published: (2026)
by: Pham, Duc Hao, et al.
Published: (2026)
Causal-Enhanced AI Agents for Medical Research Screening
by: Ngo, Duc, et al.
Published: (2026)
by: Ngo, Duc, et al.
Published: (2026)
Enhancing the Fairness and Performance of Edge Cameras with Explainable AI
by: Nguyen, Truong Thanh Hung, et al.
Published: (2024)
by: Nguyen, Truong Thanh Hung, et al.
Published: (2024)
MSC: A Marine Wildlife Video Dataset with Grounded Segmentation and Clip-Level Captioning
by: Truong, Quang-Trung, et al.
Published: (2025)
by: Truong, Quang-Trung, et al.
Published: (2025)
Secure and Efficient UAV-Based Face Detection via Homomorphic Encryption and Edge Computing
by: Van Duc, Nguyen, et al.
Published: (2025)
by: Van Duc, Nguyen, et al.
Published: (2025)
Variational Quantum Rainbow Deep Q-Network for Optimizing Resource Allocation Problem
by: Nguyen, Truong Thanh Hung, et al.
Published: (2025)
by: Nguyen, Truong Thanh Hung, et al.
Published: (2025)
Unintentional Consequences: Generative AI Use for Cybercrime
by: Luu, Truong Jack, et al.
Published: (2025)
by: Luu, Truong Jack, et al.
Published: (2025)
Motion2Meaning: A Clinician-Centered Framework for Contestable LLM in Parkinson's Disease Gait Interpretation
by: Nguyen, Loc Phuc Truong, et al.
Published: (2025)
by: Nguyen, Loc Phuc Truong, et al.
Published: (2025)
CryoSAMU: Enhancing 3D Cryo-EM Density Maps of Protein Structures at Intermediate Resolution with Structure-Aware Multimodal U-Nets
by: Zhang, Chenwei, et al.
Published: (2025)
by: Zhang, Chenwei, et al.
Published: (2025)
Crossing Linguistic Horizons: Finetuning and Comprehensive Evaluation of Vietnamese Large Language Models
by: Truong, Sang T., et al.
Published: (2024)
by: Truong, Sang T., et al.
Published: (2024)
SimInterview: Transforming Business Education through Large Language Model-Based Simulated Multilingual Interview Training System
by: Nguyen, Truong Thanh Hung, et al.
Published: (2025)
by: Nguyen, Truong Thanh Hung, et al.
Published: (2025)
BOLIMES: Boruta and LIME optiMized fEature Selection for Gene Expression Classification
by: Phan, Bich-Chung, et al.
Published: (2025)
by: Phan, Bich-Chung, et al.
Published: (2025)
Spectral Flattening Is All Muon Needs: How Orthogonalization Controls Learning Rate and Convergence
by: Nguyen, Tien-Phat, et al.
Published: (2026)
by: Nguyen, Tien-Phat, et al.
Published: (2026)
XEdgeAI: A Human-centered Industrial Inspection Framework with Data-centric Explainable Edge AI Approach
by: Nguyen, Truong Thanh Hung, et al.
Published: (2024)
by: Nguyen, Truong Thanh Hung, et al.
Published: (2024)
HanoiWorld : A Joint Embedding Predictive Architecture BasedWorld Model for Autonomous Vehicle Controller
by: Dat, Tran Tien, et al.
Published: (2026)
by: Dat, Tran Tien, et al.
Published: (2026)
An Empirical Study on Self-correcting Large Language Models for Data Science Code Generation
by: Quoc, Thai Tang, et al.
Published: (2024)
by: Quoc, Thai Tang, et al.
Published: (2024)
Similar Items
-
The Norm-Separation Delay Law of Grokking: A First-Principles Theory of Delayed Generalization
by: Khanh, Truong Xuan, et al.
Published: (2026) -
Spectral Entropy Collapse as a Phase Transition in Delayed Generalisation: An Interventional and Predictive Framework for Grokkin
by: Khanh, Truong Xuan, et al.
Published: (2026) -
Entropy Collapse: A Universal Failure Mode of Intelligent Systems
by: Khanh, Truong Xuan, et al.
Published: (2025) -
Intelligence as Trajectory-Dominant Pareto Optimization
by: Khanh, Truong Xuan, et al.
Published: (2026) -
Dynamic Intelligence Ceilings: Measuring Long-Horizon Limits of Planning and Creativity in Artificial Systems
by: Khanh, Truong Xuan, et al.
Published: (2026)