Layer Importance for Mathematical Reasoning is Forged in Pre-Training and Invariant after Post-Training
Fuente:
arXiv
Saved in:
| Main Authors: | Nepal, Aadim, Shrestha, Safal, Shrestha, Anubhav, Kim, Minwu, Naghiyev, Jalal, Shwartz-Ziv, Ravid, Ross, Keith |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Warm Up Before You Train: Unlocking General Reasoning in Resource-Constrained Settings
by: Shrestha, Safal, et al.
Published: (2025)
by: Shrestha, Safal, et al.
Published: (2025)
On the Limits of Layer Pruning for Generative Reasoning in Large Language Models
by: Shrestha, Safal, et al.
Published: (2026)
by: Shrestha, Safal, et al.
Published: (2026)
Reinforcement Learning vs. Distillation: Understanding Accuracy and Capability in LLM Reasoning
by: Kim, Minwu, et al.
Published: (2025)
by: Kim, Minwu, et al.
Published: (2025)
Training Reasoning Models on Saturated Problems via Failure-Prefix Conditioning
by: Kim, Minwu, et al.
Published: (2026)
by: Kim, Minwu, et al.
Published: (2026)
Mathematical Reasoning in Large Language Models: Assessing Logical and Arithmetic Errors across Wide Numerical Ranges
by: Shrestha, Safal, et al.
Published: (2025)
by: Shrestha, Safal, et al.
Published: (2025)
Layer by Layer: Uncovering Hidden Representations in Language Models
by: Skean, Oscar, et al.
Published: (2025)
by: Skean, Oscar, et al.
Published: (2025)
When Attention Collapses: How Degenerate Layers in LLMs Enable Smaller, Stronger Models
by: Sanyal, Sunny, et al.
Published: (2024)
by: Sanyal, Sunny, et al.
Published: (2024)
Video Representation Learning with Joint-Embedding Predictive Architectures
by: Drozdov, Katrina, et al.
Published: (2024)
by: Drozdov, Katrina, et al.
Published: (2024)
AI Must Embrace Specialization via Superhuman Adaptable Intelligence
by: Goldfeder, Judah, et al.
Published: (2026)
by: Goldfeder, Judah, et al.
Published: (2026)
Learning to Compress: Local Rank and Information Compression in Deep Neural Networks
by: Patel, Niket, et al.
Published: (2024)
by: Patel, Niket, et al.
Published: (2024)
On Training in Imagination
by: Timor, Nadav, et al.
Published: (2026)
by: Timor, Nadav, et al.
Published: (2026)
From Tokens to Thoughts: How LLMs and Humans Trade Compression for Meaning
by: Shani, Chen, et al.
Published: (2025)
by: Shani, Chen, et al.
Published: (2025)
Variance-Covariance Regularization Improves Representation Learning
by: Zhu, Jiachen, et al.
Published: (2023)
by: Zhu, Jiachen, et al.
Published: (2023)
The Illusion of Progress: Re-evaluating Hallucination Detection in LLMs
by: Janiak, Denis, et al.
Published: (2025)
by: Janiak, Denis, et al.
Published: (2025)
Nwāchā Munā: A Devanagari Speech Corpus and Proximal Transfer Benchmark for Nepal Bhasha ASR
by: Sharma, Rishikesh Kumar, et al.
Published: (2026)
by: Sharma, Rishikesh Kumar, et al.
Published: (2026)
An Information-Theoretic Perspective on Variance-Invariance-Covariance Regularization
by: Shwartz-Ziv, Ravid, et al.
Published: (2023)
by: Shwartz-Ziv, Ravid, et al.
Published: (2023)
Does Representation Matter? Exploring Intermediate Layers in Large Language Models
by: Skean, Oscar, et al.
Published: (2024)
by: Skean, Oscar, et al.
Published: (2024)
UAT-LITE: Inference-Time Uncertainty-Aware Attention for Pretrained Transformers
by: Hossain, Elias, et al.
Published: (2026)
by: Hossain, Elias, et al.
Published: (2026)
Golden Layers and Where to Find Them: Improved Knowledge Editing for Large Language Models Via Layer Gradient Analysis
by: Datta, Shrestha, et al.
Published: (2026)
by: Datta, Shrestha, et al.
Published: (2026)
Efficient Multi-Hop Question Answering over Knowledge Graphs via LLM Planning and Embedding-Guided Search
by: Shrestha, Manil, et al.
Published: (2025)
by: Shrestha, Manil, et al.
Published: (2025)
Conformal Prediction for Risk-Controlled Medical Entity Extraction Across Clinical Domains
by: Shrestha, Manil, et al.
Published: (2026)
by: Shrestha, Manil, et al.
Published: (2026)
Beyond the Loss Curve: Scaling Laws, Active Learning, and the Limits of Learning from Exact Posteriors
by: Khorasani, Arian, et al.
Published: (2026)
by: Khorasani, Arian, et al.
Published: (2026)
Scaling Behaviors of LLM Reinforcement Learning Post-Training: An Empirical Study in Mathematical Reasoning
by: Tan, Zelin, et al.
Published: (2025)
by: Tan, Zelin, et al.
Published: (2025)
Antislop: A Comprehensive Framework for Identifying and Eliminating Repetitive Patterns in Language Models
by: Paech, Samuel, et al.
Published: (2025)
by: Paech, Samuel, et al.
Published: (2025)
Theoretical Perspectives on Data Quality and Synergistic Effects in Pre- and Post-Training Reasoning Models
by: Javanmard, Adel, et al.
Published: (2026)
by: Javanmard, Adel, et al.
Published: (2026)
The Coverage Principle: How Pre-Training Enables Post-Training
by: Chen, Fan, et al.
Published: (2025)
by: Chen, Fan, et al.
Published: (2025)
Rethinking Leveraging Pre-Trained Multi-Layer Representations for Speaker Verification
by: Kim, Jin Sob, et al.
Published: (2025)
by: Kim, Jin Sob, et al.
Published: (2025)
JEPA as a Neural Tokenizer: Learning Robust Speech Representations with Density Adaptive Attention
by: Ioannides, Georgios, et al.
Published: (2025)
by: Ioannides, Georgios, et al.
Published: (2025)
Distribution Matching via Generalized Consistency Models
by: Shrestha, Sagar, et al.
Published: (2025)
by: Shrestha, Sagar, et al.
Published: (2025)
Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same Coin
by: Queipo-de-Llano, Enrique, et al.
Published: (2025)
by: Queipo-de-Llano, Enrique, et al.
Published: (2025)
A superpersuasive autonomous policy debating system
by: Roush, Allen, et al.
Published: (2025)
by: Roush, Allen, et al.
Published: (2025)
Understanding the Emergence of Seemingly Useless Features in Next-Token Predictors
by: Rofin, Mark, et al.
Published: (2026)
by: Rofin, Mark, et al.
Published: (2026)
Self-Supervised Pre-Training for Precipitation Post-Processor
by: An, Sojung, et al.
Published: (2023)
by: An, Sojung, et al.
Published: (2023)
Visual Grounding Methods for VQA are Working for the Wrong Reasons!
by: Shrestha, Robik, et al.
Published: (2020)
by: Shrestha, Robik, et al.
Published: (2020)
The Entropy Enigma: Success and Failure of Entropy Minimization
by: Press, Ori, et al.
Published: (2024)
by: Press, Ori, et al.
Published: (2024)
Trust-Based Incentive Mechanisms in Semi-Decentralized Federated Learning Systems
by: Shrestha, Ajay Kumar
Published: (2026)
by: Shrestha, Ajay Kumar
Published: (2026)
Enhanced Deep Q-Learning for 2D Self-Driving Cars: Implementation and Evaluation on a Custom Track Environment
by: Pathak, Sagar, et al.
Published: (2024)
by: Pathak, Sagar, et al.
Published: (2024)
Secure Multiparty Generative AI
by: Shrestha, Manil, et al.
Published: (2024)
by: Shrestha, Manil, et al.
Published: (2024)
In-Context Learning for Label-Efficient Cancer Image Classification in Oncology
by: Shrestha, Mobina, et al.
Published: (2025)
by: Shrestha, Mobina, et al.
Published: (2025)
Soft Clustering Anchors for Self-Supervised Speech Representation Learning in Joint Embedding Prediction Architectures
by: Ioannides, Georgios, et al.
Published: (2026)
by: Ioannides, Georgios, et al.
Published: (2026)
Similar Items
-
Warm Up Before You Train: Unlocking General Reasoning in Resource-Constrained Settings
by: Shrestha, Safal, et al.
Published: (2025) -
On the Limits of Layer Pruning for Generative Reasoning in Large Language Models
by: Shrestha, Safal, et al.
Published: (2026) -
Reinforcement Learning vs. Distillation: Understanding Accuracy and Capability in LLM Reasoning
by: Kim, Minwu, et al.
Published: (2025) -
Training Reasoning Models on Saturated Problems via Failure-Prefix Conditioning
by: Kim, Minwu, et al.
Published: (2026) -
Mathematical Reasoning in Large Language Models: Assessing Logical and Arithmetic Errors across Wide Numerical Ranges
by: Shrestha, Safal, et al.
Published: (2025)