Towards a Theoretical Understanding of Synthetic Data in LLM Post-Training: A Reverse-Bottleneck Perspective
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gan, Zeyu, Liu, Yong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BiLLM: Pushing the Limit of Post-Training Quantization for LLMs
von: Huang, Wei, et al.
Veröffentlicht: (2024)
von: Huang, Wei, et al.
Veröffentlicht: (2024)
Understanding R1-Zero-Like Training: A Critical Perspective
von: Liu, Zichen, et al.
Veröffentlicht: (2025)
von: Liu, Zichen, et al.
Veröffentlicht: (2025)
Towards a Unified View of Large Language Model Post-Training
von: Lv, Xingtai, et al.
Veröffentlicht: (2025)
von: Lv, Xingtai, et al.
Veröffentlicht: (2025)
Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training
von: Liu, Yixin, et al.
Veröffentlicht: (2026)
von: Liu, Yixin, et al.
Veröffentlicht: (2026)
LoopTool: Closing the Data-Training Loop for Robust LLM Tool Calls
von: Zhang, Kangning, et al.
Veröffentlicht: (2025)
von: Zhang, Kangning, et al.
Veröffentlicht: (2025)
Compositional Generalization from Learned Skills via CoT Training: A Theoretical and Structural Analysis for Reasoning
von: Yao, Xinhao, et al.
Veröffentlicht: (2025)
von: Yao, Xinhao, et al.
Veröffentlicht: (2025)
Post-training an LLM for RAG? Train on Self-Generated Demonstrations
von: Finlayson, Matthew, et al.
Veröffentlicht: (2025)
von: Finlayson, Matthew, et al.
Veröffentlicht: (2025)
LASA: Language-Agnostic Semantic Alignment at the Semantic Bottleneck for LLM Safety
von: Yang, Junxiao, et al.
Veröffentlicht: (2026)
von: Yang, Junxiao, et al.
Veröffentlicht: (2026)
Does Training on Synthetic Data Make Models Less Robust?
von: Zhang, Lingze, et al.
Veröffentlicht: (2025)
von: Zhang, Lingze, et al.
Veröffentlicht: (2025)
Breaking the Capability Ceiling of LLM Post-Training by Reintroducing Markov States
von: Yuan, Yurun, et al.
Veröffentlicht: (2026)
von: Yuan, Yurun, et al.
Veröffentlicht: (2026)
PITA: Preference-Guided Inference-Time Alignment for LLM Post-Training
von: Bobbili, Sarat Chandra, et al.
Veröffentlicht: (2025)
von: Bobbili, Sarat Chandra, et al.
Veröffentlicht: (2025)
$Q\sharp$: Provably Optimal Distributional RL for LLM Post-Training
von: Zhou, Jin Peng, et al.
Veröffentlicht: (2025)
von: Zhou, Jin Peng, et al.
Veröffentlicht: (2025)
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training
von: Pan, Chengjun, et al.
Veröffentlicht: (2026)
von: Pan, Chengjun, et al.
Veröffentlicht: (2026)
Towards Understanding Safety Alignment: A Mechanistic Perspective from Safety Neurons
von: Chen, Jianhui, et al.
Veröffentlicht: (2024)
von: Chen, Jianhui, et al.
Veröffentlicht: (2024)
How Bad is Training on Synthetic Data? A Statistical Analysis of Language Model Collapse
von: Seddik, Mohamed El Amine, et al.
Veröffentlicht: (2024)
von: Seddik, Mohamed El Amine, et al.
Veröffentlicht: (2024)
ABBEL: LLM Agents Acting through Belief Bottlenecks Expressed in Language
von: Lidayan, Aly, et al.
Veröffentlicht: (2025)
von: Lidayan, Aly, et al.
Veröffentlicht: (2025)
A Theoretical Perspective for Speculative Decoding Algorithm
von: Yin, Ming, et al.
Veröffentlicht: (2024)
von: Yin, Ming, et al.
Veröffentlicht: (2024)
Lightweight Safety Guardrails via Synthetic Data and RL-guided Adversarial Training
von: Ilin, Aleksei, et al.
Veröffentlicht: (2025)
von: Ilin, Aleksei, et al.
Veröffentlicht: (2025)
Toward Understanding BERT-Like Pre-Training for DNA Foundation Models
von: Liang, Chaoqi, et al.
Veröffentlicht: (2023)
von: Liang, Chaoqi, et al.
Veröffentlicht: (2023)
Synthetic vs. Gold: The Role of LLM Generated Labels and Data in Cyberbullying Detection
von: Kazemi, Arefeh, et al.
Veröffentlicht: (2025)
von: Kazemi, Arefeh, et al.
Veröffentlicht: (2025)
CoT-Space: A Theoretical Framework for Internal Slow-Thinking via Reinforcement Learning
von: Gan, Zeyu, et al.
Veröffentlicht: (2025)
von: Gan, Zeyu, et al.
Veröffentlicht: (2025)
XL-Suite: Cross-Lingual Synthetic Training and Evaluation Data for Open-Ended Generation
von: Iyer, Vivek, et al.
Veröffentlicht: (2025)
von: Iyer, Vivek, et al.
Veröffentlicht: (2025)
ShiftAddLLM: Accelerating Pretrained LLMs via Post-Training Multiplication-Less Reparameterization
von: You, Haoran, et al.
Veröffentlicht: (2024)
von: You, Haoran, et al.
Veröffentlicht: (2024)
Dr. Post-Training: A Data Regularization Perspective on LLM Post-Training
von: Hu, Pingbang, et al.
Veröffentlicht: (2026)
von: Hu, Pingbang, et al.
Veröffentlicht: (2026)
Towards Efficient Resume Understanding: A Multi-Granularity Multi-Modal Pre-Training Approach
von: Jiang, Feihu, et al.
Veröffentlicht: (2024)
von: Jiang, Feihu, et al.
Veröffentlicht: (2024)
Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning
von: Zhao, Shiwan, et al.
Veröffentlicht: (2026)
von: Zhao, Shiwan, et al.
Veröffentlicht: (2026)
Heterogeneity in Formal Linguistic Competence of Language Models: Is Data the Real Bottleneck?
von: Renduchintala, H S V N S Kowndinya, et al.
Veröffentlicht: (2026)
von: Renduchintala, H S V N S Kowndinya, et al.
Veröffentlicht: (2026)
Learning to Reason Efficiently with A* Post-Training
von: Opedal, Andreas, et al.
Veröffentlicht: (2026)
von: Opedal, Andreas, et al.
Veröffentlicht: (2026)
Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls
von: Kang, Feiyang, et al.
Veröffentlicht: (2025)
von: Kang, Feiyang, et al.
Veröffentlicht: (2025)
Marco-o1 v2: Towards Widening The Distillation Bottleneck for Reasoning Models
von: Yin, Huifeng, et al.
Veröffentlicht: (2025)
von: Yin, Huifeng, et al.
Veröffentlicht: (2025)
KALAVAI: Predicting When Independent Specialist Fusion Works -- A Quantitative Model for Post-Hoc Cooperative LLM Training
von: Kumaresan, Ramchand
Veröffentlicht: (2026)
von: Kumaresan, Ramchand
Veröffentlicht: (2026)
Policy Learning with a Language Bottleneck
von: Srivastava, Megha, et al.
Veröffentlicht: (2024)
von: Srivastava, Megha, et al.
Veröffentlicht: (2024)
Understanding the planning of LLM agents: A survey
von: Huang, Xu, et al.
Veröffentlicht: (2024)
von: Huang, Xu, et al.
Veröffentlicht: (2024)
Towards a Theoretical Understanding of the 'Reversal Curse' via Training Dynamics
von: Zhu, Hanlin, et al.
Veröffentlicht: (2024)
von: Zhu, Hanlin, et al.
Veröffentlicht: (2024)
How Instruction and Reasoning Data shape Post-Training: Data Quality through the Lens of Layer-wise Gradients
von: Li, Ming, et al.
Veröffentlicht: (2025)
von: Li, Ming, et al.
Veröffentlicht: (2025)
Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn Search Agents
von: Wang, Guoqing, et al.
Veröffentlicht: (2025)
von: Wang, Guoqing, et al.
Veröffentlicht: (2025)
You Can Generate It Again: Data-to-Text Generation with Verification and Correction Prompting
von: Ren, Xuan, et al.
Veröffentlicht: (2023)
von: Ren, Xuan, et al.
Veröffentlicht: (2023)
Guiding LLM Post-training Data Engineering with Model Internals from Sparse Autoencoders
von: Jing, Yi, et al.
Veröffentlicht: (2026)
von: Jing, Yi, et al.
Veröffentlicht: (2026)
Towards Best Practices for Open Datasets for LLM Training
von: Baack, Stefan, et al.
Veröffentlicht: (2025)
von: Baack, Stefan, et al.
Veröffentlicht: (2025)
Understanding Post-hoc Explainers: The Case of Anchors
von: Lopardo, Gianluigi, et al.
Veröffentlicht: (2023)
von: Lopardo, Gianluigi, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
BiLLM: Pushing the Limit of Post-Training Quantization for LLMs
von: Huang, Wei, et al.
Veröffentlicht: (2024) -
Understanding R1-Zero-Like Training: A Critical Perspective
von: Liu, Zichen, et al.
Veröffentlicht: (2025) -
Towards a Unified View of Large Language Model Post-Training
von: Lv, Xingtai, et al.
Veröffentlicht: (2025) -
Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training
von: Liu, Yixin, et al.
Veröffentlicht: (2026) -
LoopTool: Closing the Data-Training Loop for Robust LLM Tool Calls
von: Zhang, Kangning, et al.
Veröffentlicht: (2025)