Dr. Post-Training: A Data Regularization Perspective on LLM Post-Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hu, Pingbang, Liu, Xueshen, Mao, Z. Morley, Ma, Jiaqi W. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Reliable Cryptographic Framework for Empirical Machine Unlearning Evaluation
von: Tu, Yiwen, et al.
Veröffentlicht: (2024)
von: Tu, Yiwen, et al.
Veröffentlicht: (2024)
A Unified Theory of Random Projection for Influence Functions
von: Hu, Pingbang, et al.
Veröffentlicht: (2026)
von: Hu, Pingbang, et al.
Veröffentlicht: (2026)
GraSS: Scalable Data Attribution with Gradient Sparsification and Sparse Projection
von: Hu, Pingbang, et al.
Veröffentlicht: (2025)
von: Hu, Pingbang, et al.
Veröffentlicht: (2025)
PostTrainBench: Can LLM Agents Automate LLM Post-Training?
von: Rank, Ben, et al.
Veröffentlicht: (2026)
von: Rank, Ben, et al.
Veröffentlicht: (2026)
Astro: Activation-guided Structured Regularization for Outlier-Robust LLM Post-Training Quantization
von: Chen, Xi, et al.
Veröffentlicht: (2026)
von: Chen, Xi, et al.
Veröffentlicht: (2026)
Towards a Theoretical Understanding of Synthetic Data in LLM Post-Training: A Reverse-Bottleneck Perspective
von: Gan, Zeyu, et al.
Veröffentlicht: (2024)
von: Gan, Zeyu, et al.
Veröffentlicht: (2024)
Automatic Configuration of LLM Post-Training Pipelines
von: Chwa, Channe, et al.
Veröffentlicht: (2026)
von: Chwa, Channe, et al.
Veröffentlicht: (2026)
Theoretical Perspectives on Data Quality and Synergistic Effects in Pre- and Post-Training Reasoning Models
von: Javanmard, Adel, et al.
Veröffentlicht: (2026)
von: Javanmard, Adel, et al.
Veröffentlicht: (2026)
RaanA: A Fast, Flexible, and Data-Efficient Post-Training Quantization Algorithm
von: Yang, Yongyi, et al.
Veröffentlicht: (2025)
von: Yang, Yongyi, et al.
Veröffentlicht: (2025)
DAQ: Delta-Aware Quantization for Post-Training LLM Weight Compression
von: Yu, Xiaoming, et al.
Veröffentlicht: (2026)
von: Yu, Xiaoming, et al.
Veröffentlicht: (2026)
Probe-Based Data Attribution: Discovering and Mitigating Undesirable Behaviors in LLM Post-Training
von: Xiao, Frank, et al.
Veröffentlicht: (2026)
von: Xiao, Frank, et al.
Veröffentlicht: (2026)
Zipping the Thought: When and How Compressed Reasoning Data Works in LLM Post-Training
von: Matsutani, Kohsei, et al.
Veröffentlicht: (2026)
von: Matsutani, Kohsei, et al.
Veröffentlicht: (2026)
BiLLM: Pushing the Limit of Post-Training Quantization for LLMs
von: Huang, Wei, et al.
Veröffentlicht: (2024)
von: Huang, Wei, et al.
Veröffentlicht: (2024)
PT$^2$-LLM: Post-Training Ternarization for Large Language Models
von: Yan, Xianglong, et al.
Veröffentlicht: (2025)
von: Yan, Xianglong, et al.
Veröffentlicht: (2025)
DVPO: Distributional Value Modeling-based Policy Optimization for LLM Post-Training
von: Zhu, Dingwei, et al.
Veröffentlicht: (2025)
von: Zhu, Dingwei, et al.
Veröffentlicht: (2025)
AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training
von: Han, Zhenyu, et al.
Veröffentlicht: (2025)
von: Han, Zhenyu, et al.
Veröffentlicht: (2025)
Benchmarking Post-Training Quantization in LLMs: Comprehensive Taxonomy, Unified Evaluation, and Comparative Analysis
von: Zhao, Jiaqi, et al.
Veröffentlicht: (2025)
von: Zhao, Jiaqi, et al.
Veröffentlicht: (2025)
Learn To be Efficient: Build Structured Sparsity in Large Language Models
von: Zheng, Haizhong, et al.
Veröffentlicht: (2024)
von: Zheng, Haizhong, et al.
Veröffentlicht: (2024)
Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training
von: Liu, Yixin, et al.
Veröffentlicht: (2026)
von: Liu, Yixin, et al.
Veröffentlicht: (2026)
GAC: Noise-Aware Adaptive Mixing for Hybrid SFT-RL Post-Training
von: Hu, Yuelin, et al.
Veröffentlicht: (2026)
von: Hu, Yuelin, et al.
Veröffentlicht: (2026)
Front-Loading Reasoning: The Synergy between Pretraining and Post-Training Data
von: Akter, Syeda Nahida, et al.
Veröffentlicht: (2025)
von: Akter, Syeda Nahida, et al.
Veröffentlicht: (2025)
Post-training an LLM for RAG? Train on Self-Generated Demonstrations
von: Finlayson, Matthew, et al.
Veröffentlicht: (2025)
von: Finlayson, Matthew, et al.
Veröffentlicht: (2025)
On Distinguishing Capability Elicitation from Capability Creation in Post-Training: A Free-Energy Perspective
von: Li, Yuhao, et al.
Veröffentlicht: (2026)
von: Li, Yuhao, et al.
Veröffentlicht: (2026)
Scaling Behaviors of LLM Reinforcement Learning Post-Training: An Empirical Study in Mathematical Reasoning
von: Tan, Zelin, et al.
Veröffentlicht: (2025)
von: Tan, Zelin, et al.
Veröffentlicht: (2025)
RedOne 2.0: Rethinking Domain-specific LLM Post-Training in Social Networking Services
von: Zhao, Fei, et al.
Veröffentlicht: (2025)
von: Zhao, Fei, et al.
Veröffentlicht: (2025)
Learning-Zone Energy: Online Data Selection for Efficient RL Post-Training
von: Cui, Peng, et al.
Veröffentlicht: (2026)
von: Cui, Peng, et al.
Veröffentlicht: (2026)
CoScale-RL: Efficient Post-Training by Co-Scaling Data and Computation
von: Chen, Yutong, et al.
Veröffentlicht: (2026)
von: Chen, Yutong, et al.
Veröffentlicht: (2026)
Pushing the Limits of Block Rotations in Post-Training Quantization
von: Sanjeet, Sai, et al.
Veröffentlicht: (2026)
von: Sanjeet, Sai, et al.
Veröffentlicht: (2026)
Breaking the Capability Ceiling of LLM Post-Training by Reintroducing Markov States
von: Yuan, Yurun, et al.
Veröffentlicht: (2026)
von: Yuan, Yurun, et al.
Veröffentlicht: (2026)
Self-Supervised Pre-Training for Precipitation Post-Processor
von: An, Sojung, et al.
Veröffentlicht: (2023)
von: An, Sojung, et al.
Veröffentlicht: (2023)
Intrinsically Interpretable Attention via Sparse Post-Training
von: Draye, Florent, et al.
Veröffentlicht: (2025)
von: Draye, Florent, et al.
Veröffentlicht: (2025)
Post-Training Statistical Calibration for Higher Activation Sparsity
von: Chua, Vui Seng, et al.
Veröffentlicht: (2024)
von: Chua, Vui Seng, et al.
Veröffentlicht: (2024)
Beacon: Post-Training Quantization with Integrated Grid Selection
von: Zhang, Shihao, et al.
Veröffentlicht: (2025)
von: Zhang, Shihao, et al.
Veröffentlicht: (2025)
PITA: Preference-Guided Inference-Time Alignment for LLM Post-Training
von: Bobbili, Sarat Chandra, et al.
Veröffentlicht: (2025)
von: Bobbili, Sarat Chandra, et al.
Veröffentlicht: (2025)
$Q\sharp$: Provably Optimal Distributional RL for LLM Post-Training
von: Zhou, Jin Peng, et al.
Veröffentlicht: (2025)
von: Zhou, Jin Peng, et al.
Veröffentlicht: (2025)
RiskPO: Risk-based Policy Optimization via Verifiable Reward for LLM Post-Training
von: Ren, Tao, et al.
Veröffentlicht: (2025)
von: Ren, Tao, et al.
Veröffentlicht: (2025)
Layer Importance for Mathematical Reasoning is Forged in Pre-Training and Invariant after Post-Training
von: Nepal, Aadim, et al.
Veröffentlicht: (2025)
von: Nepal, Aadim, et al.
Veröffentlicht: (2025)
Efficient Ensembles Improve Training Data Attribution
von: Deng, Junwei, et al.
Veröffentlicht: (2024)
von: Deng, Junwei, et al.
Veröffentlicht: (2024)
How Reasoning Evolves from Post-Training Data: An Empirical Study Using Chess
von: Dionisopoulos, Lucas, et al.
Veröffentlicht: (2026)
von: Dionisopoulos, Lucas, et al.
Veröffentlicht: (2026)
A Quantitative Characterization of Forgetting in Post-Training
von: Balasubramanian, Krishnakumar, et al.
Veröffentlicht: (2026)
von: Balasubramanian, Krishnakumar, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
A Reliable Cryptographic Framework for Empirical Machine Unlearning Evaluation
von: Tu, Yiwen, et al.
Veröffentlicht: (2024) -
A Unified Theory of Random Projection for Influence Functions
von: Hu, Pingbang, et al.
Veröffentlicht: (2026) -
GraSS: Scalable Data Attribution with Gradient Sparsification and Sparse Projection
von: Hu, Pingbang, et al.
Veröffentlicht: (2025) -
PostTrainBench: Can LLM Agents Automate LLM Post-Training?
von: Rank, Ben, et al.
Veröffentlicht: (2026) -
Astro: Activation-guided Structured Regularization for Outlier-Robust LLM Post-Training Quantization
von: Chen, Xi, et al.
Veröffentlicht: (2026)