Disentangling Feature Structure: A Mathematically Provable Two-Stage Training Dynamics in Transformers
Fuente:
arXiv
Guardado en:
| Autores principales: | Gong, Zixuan, Li, Shijia, Liu, Yong, Teng, Jiaye |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
What Makes Looped Transformers Perform Better Than Non-Recursive Ones
por: Gong, Zixuan, et al.
Publicado: (2025)
por: Gong, Zixuan, et al.
Publicado: (2025)
Unveiling Induction Heads: Provable Training Dynamics and Feature Learning in Transformers
por: Chen, Siyu, et al.
Publicado: (2024)
por: Chen, Siyu, et al.
Publicado: (2024)
Two-Stage Regularization-Based Structured Pruning for LLMs
por: Feng, Mingkuan, et al.
Publicado: (2025)
por: Feng, Mingkuan, et al.
Publicado: (2025)
Brain-Inspired Two-Stage Approach: Enhancing Mathematical Reasoning by Imitating Human Thought Processes
por: Chen, Yezeng, et al.
Publicado: (2024)
por: Chen, Yezeng, et al.
Publicado: (2024)
GraphDancer: Training LLMs to Explore and Reason over Graphs via Two-Stage Curriculum Post-Training
por: Bai, Yuyang, et al.
Publicado: (2026)
por: Bai, Yuyang, et al.
Publicado: (2026)
2SSP: A Two-Stage Framework for Structured Pruning of LLMs
por: Sandri, Fabrizio, et al.
Publicado: (2025)
por: Sandri, Fabrizio, et al.
Publicado: (2025)
$Q\sharp$: Provably Optimal Distributional RL for LLM Post-Training
por: Zhou, Jin Peng, et al.
Publicado: (2025)
por: Zhou, Jin Peng, et al.
Publicado: (2025)
An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
por: Hao, Yuren, et al.
Publicado: (2025)
por: Hao, Yuren, et al.
Publicado: (2025)
Latent Concept Disentanglement in Transformer-based Language Models
por: Hong, Guan Zhe, et al.
Publicado: (2025)
por: Hong, Guan Zhe, et al.
Publicado: (2025)
Can A Gamer Train A Mathematical Reasoning Model?
por: Shin, Andrew
Publicado: (2025)
por: Shin, Andrew
Publicado: (2025)
Compositional Generalization from Learned Skills via CoT Training: A Theoretical and Structural Analysis for Reasoning
por: Yao, Xinhao, et al.
Publicado: (2025)
por: Yao, Xinhao, et al.
Publicado: (2025)
LycheeCluster: Efficient Long-Context Inference with Structure-Aware Chunking and Hierarchical KV Indexing
por: Li, Dongfang, et al.
Publicado: (2026)
por: Li, Dongfang, et al.
Publicado: (2026)
Dual-Modality Multi-Stage Adversarial Safety Training: Robustifying Multimodal Web Agents Against Cross-Modal Attacks
por: Liu, Haoyu, et al.
Publicado: (2026)
por: Liu, Haoyu, et al.
Publicado: (2026)
Lil: Less is Less When Applying Post-Training Sparse-Attention Algorithms in Long-Decode Stage
por: Hu, Junhao, et al.
Publicado: (2026)
por: Hu, Junhao, et al.
Publicado: (2026)
Towards a Theoretical Understanding of Synthetic Data in LLM Post-Training: A Reverse-Bottleneck Perspective
por: Gan, Zeyu, et al.
Publicado: (2024)
por: Gan, Zeyu, et al.
Publicado: (2024)
Evolving Demonstration Optimization for Chain-of-Thought Feature Transformation
por: Wang, Xinyuan, et al.
Publicado: (2026)
por: Wang, Xinyuan, et al.
Publicado: (2026)
Improving Word Translation via Two-Stage Contrastive Learning
por: Li, Yaoyiran, et al.
Publicado: (2022)
por: Li, Yaoyiran, et al.
Publicado: (2022)
Emergence of Episodic Memory in Transformers: Characterizing Changes in Temporal Structure of Attention Scores During Training
por: Mistry, Deven Mahesh, et al.
Publicado: (2025)
por: Mistry, Deven Mahesh, et al.
Publicado: (2025)
POET-X: Memory-efficient LLM Training by Scaling Orthogonal Transformation
por: Qiu, Zeju, et al.
Publicado: (2026)
por: Qiu, Zeju, et al.
Publicado: (2026)
SpanNorm: Reconciling Training Stability and Performance in Deep Transformers
por: Wang, Chao, et al.
Publicado: (2026)
por: Wang, Chao, et al.
Publicado: (2026)
Reparameterized LLM Training via Orthogonal Equivalence Transformation
por: Qiu, Zeju, et al.
Publicado: (2025)
por: Qiu, Zeju, et al.
Publicado: (2025)
Provable Scaling Laws for the Test-Time Compute of Large Language Models
por: Chen, Yanxi, et al.
Publicado: (2024)
por: Chen, Yanxi, et al.
Publicado: (2024)
Provable Interactive Learning with Hindsight Instruction Feedback
por: Misra, Dipendra, et al.
Publicado: (2024)
por: Misra, Dipendra, et al.
Publicado: (2024)
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning
por: Zheng, Congmin, et al.
Publicado: (2025)
por: Zheng, Congmin, et al.
Publicado: (2025)
Time-Aware Feature Selection: Adaptive Temporal Masking for Stable Sparse Autoencoder Training
por: Li, T. Ed, et al.
Publicado: (2025)
por: Li, T. Ed, et al.
Publicado: (2025)
On the Mathematical Relationship Between Layer Normalization and Dynamic Activation Functions
por: Stollenwerk, Felix
Publicado: (2025)
por: Stollenwerk, Felix
Publicado: (2025)
Off-the-Shelf LLMs as Process Scorers: Training-Free Alternative to PRMs for Mathematical Reasoning
por: Chegini, Atoosa, et al.
Publicado: (2026)
por: Chegini, Atoosa, et al.
Publicado: (2026)
STAT: Shrinking Transformers After Training
por: Flynn, Megan, et al.
Publicado: (2024)
por: Flynn, Megan, et al.
Publicado: (2024)
TSO: Self-Training with Scaled Preference Optimization
por: Chen, Kaihui, et al.
Publicado: (2024)
por: Chen, Kaihui, et al.
Publicado: (2024)
Tuning without Peeking: Provable Generalization Bounds and Robust LLM Post-Training
por: Labiad, Ismail, et al.
Publicado: (2025)
por: Labiad, Ismail, et al.
Publicado: (2025)
Efficient Mathematical Reasoning Models via Dynamic Pruning and Knowledge Distillation
por: Yu, Fengming, et al.
Publicado: (2025)
por: Yu, Fengming, et al.
Publicado: (2025)
DISPO: Enhancing Training Efficiency and Stability in Reinforcement Learning for Large Language Model Mathematical Reasoning
por: Karaman, Batuhan K., et al.
Publicado: (2026)
por: Karaman, Batuhan K., et al.
Publicado: (2026)
Transformers as Decision Makers: Provable In-Context Reinforcement Learning via Supervised Pretraining
por: Lin, Licong, et al.
Publicado: (2023)
por: Lin, Licong, et al.
Publicado: (2023)
Provable Length Generalization in Sequence Prediction via Spectral Filtering
por: Marsden, Annie, et al.
Publicado: (2024)
por: Marsden, Annie, et al.
Publicado: (2024)
Interpretable Steering of Large Language Models with Feature Guided Activation Additions
por: Soo, Samuel, et al.
Publicado: (2025)
por: Soo, Samuel, et al.
Publicado: (2025)
Diagnosing Transformers: Illuminating Feature Spaces for Clinical Decision-Making
por: Hsu, Aliyah R., et al.
Publicado: (2023)
por: Hsu, Aliyah R., et al.
Publicado: (2023)
Teaching Transformers Causal Reasoning through Axiomatic Training
por: Vashishtha, Aniket, et al.
Publicado: (2024)
por: Vashishtha, Aniket, et al.
Publicado: (2024)
Can Post-Training Transform LLMs into Causal Reasoners?
por: Chen, Junqi, et al.
Publicado: (2026)
por: Chen, Junqi, et al.
Publicado: (2026)
Closing the Loop: Learning to Generate Writing Feedback via Language Model Simulated Student Revisions
por: Nair, Inderjeet, et al.
Publicado: (2024)
por: Nair, Inderjeet, et al.
Publicado: (2024)
Learning Dynamics in Continual Pre-Training for Large Language Models
por: Wang, Xingjin, et al.
Publicado: (2025)
por: Wang, Xingjin, et al.
Publicado: (2025)
Ejemplares similares
-
What Makes Looped Transformers Perform Better Than Non-Recursive Ones
por: Gong, Zixuan, et al.
Publicado: (2025) -
Unveiling Induction Heads: Provable Training Dynamics and Feature Learning in Transformers
por: Chen, Siyu, et al.
Publicado: (2024) -
Two-Stage Regularization-Based Structured Pruning for LLMs
por: Feng, Mingkuan, et al.
Publicado: (2025) -
Brain-Inspired Two-Stage Approach: Enhancing Mathematical Reasoning by Imitating Human Thought Processes
por: Chen, Yezeng, et al.
Publicado: (2024) -
GraphDancer: Training LLMs to Explore and Reason over Graphs via Two-Stage Curriculum Post-Training
por: Bai, Yuyang, et al.
Publicado: (2026)