Post-training an LLM for RAG? Train on Self-Generated Demonstrations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Finlayson, Matthew, Kulikov, Ilia, Bikel, Daniel M., Oguz, Barlas, Chen, Xilun, Pappu, Aasish |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Self-Improving Pretraining: using post-trained models to pretrain better models
von: Tan, Ellen Xiaoqing, et al.
Veröffentlicht: (2026)
von: Tan, Ellen Xiaoqing, et al.
Veröffentlicht: (2026)
J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning
von: Whitehouse, Chenxi, et al.
Veröffentlicht: (2025)
von: Whitehouse, Chenxi, et al.
Veröffentlicht: (2025)
Backtracking Improves Generation Safety
von: Zhang, Yiming, et al.
Veröffentlicht: (2024)
von: Zhang, Yiming, et al.
Veröffentlicht: (2024)
Learning Facts at Scale with Active Reading
von: Lin, Jessy, et al.
Veröffentlicht: (2025)
von: Lin, Jessy, et al.
Veröffentlicht: (2025)
Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training
von: Liu, Yixin, et al.
Veröffentlicht: (2026)
von: Liu, Yixin, et al.
Veröffentlicht: (2026)
BiLLM: Pushing the Limit of Post-Training Quantization for LLMs
von: Huang, Wei, et al.
Veröffentlicht: (2024)
von: Huang, Wei, et al.
Veröffentlicht: (2024)
Guiding LLM Post-training Data Engineering with Model Internals from Sparse Autoencoders
von: Jing, Yi, et al.
Veröffentlicht: (2026)
von: Jing, Yi, et al.
Veröffentlicht: (2026)
Breaking the Capability Ceiling of LLM Post-Training by Reintroducing Markov States
von: Yuan, Yurun, et al.
Veröffentlicht: (2026)
von: Yuan, Yurun, et al.
Veröffentlicht: (2026)
PITA: Preference-Guided Inference-Time Alignment for LLM Post-Training
von: Bobbili, Sarat Chandra, et al.
Veröffentlicht: (2025)
von: Bobbili, Sarat Chandra, et al.
Veröffentlicht: (2025)
$Q\sharp$: Provably Optimal Distributional RL for LLM Post-Training
von: Zhou, Jin Peng, et al.
Veröffentlicht: (2025)
von: Zhou, Jin Peng, et al.
Veröffentlicht: (2025)
FLAME: Factuality-Aware Alignment for Large Language Models
von: Lin, Sheng-Chieh, et al.
Veröffentlicht: (2024)
von: Lin, Sheng-Chieh, et al.
Veröffentlicht: (2024)
Prioritized Replay for RL Post-training
von: Fatemi, Mehdi
Veröffentlicht: (2026)
von: Fatemi, Mehdi
Veröffentlicht: (2026)
Self-Trained Verification for Training- and Test-Time Self-Improvement
von: Wu, Chen Henry, et al.
Veröffentlicht: (2026)
von: Wu, Chen Henry, et al.
Veröffentlicht: (2026)
Tuning without Peeking: Provable Generalization Bounds and Robust LLM Post-Training
von: Labiad, Ismail, et al.
Veröffentlicht: (2025)
von: Labiad, Ismail, et al.
Veröffentlicht: (2025)
Can Post-Training Transform LLMs into Causal Reasoners?
von: Chen, Junqi, et al.
Veröffentlicht: (2026)
von: Chen, Junqi, et al.
Veröffentlicht: (2026)
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training
von: Pan, Chengjun, et al.
Veröffentlicht: (2026)
von: Pan, Chengjun, et al.
Veröffentlicht: (2026)
Post-Trained MoE Can Skip Half Experts via Self-Distillation
von: Lv, Xingtai, et al.
Veröffentlicht: (2026)
von: Lv, Xingtai, et al.
Veröffentlicht: (2026)
Automatic Pair Construction for Contrastive Post-training
von: Xu, Canwen, et al.
Veröffentlicht: (2023)
von: Xu, Canwen, et al.
Veröffentlicht: (2023)
From Decoding to Meta-Generation: Inference-time Algorithms for Large Language Models
von: Welleck, Sean, et al.
Veröffentlicht: (2024)
von: Welleck, Sean, et al.
Veröffentlicht: (2024)
ShiftAddLLM: Accelerating Pretrained LLMs via Post-Training Multiplication-Less Reparameterization
von: You, Haoran, et al.
Veröffentlicht: (2024)
von: You, Haoran, et al.
Veröffentlicht: (2024)
Self-Rewarding PPO: Aligning Large Language Models with Demonstrations Only
von: Zhang, Qingru, et al.
Veröffentlicht: (2025)
von: Zhang, Qingru, et al.
Veröffentlicht: (2025)
Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators
von: Roytburg, Dani, et al.
Veröffentlicht: (2025)
von: Roytburg, Dani, et al.
Veröffentlicht: (2025)
Comparative Analysis of Demonstration Selection Algorithms for LLM In-Context Learning
von: Shu, Dong, et al.
Veröffentlicht: (2024)
von: Shu, Dong, et al.
Veröffentlicht: (2024)
Post-training for Efficient Communication via Convention Formation
von: Hua, Yilun, et al.
Veröffentlicht: (2025)
von: Hua, Yilun, et al.
Veröffentlicht: (2025)
Efficient Post-training Quantization with FP8 Formats
von: Shen, Haihao, et al.
Veröffentlicht: (2023)
von: Shen, Haihao, et al.
Veröffentlicht: (2023)
Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations
von: Roytburg, Dani, et al.
Veröffentlicht: (2026)
von: Roytburg, Dani, et al.
Veröffentlicht: (2026)
StateX: Enhancing RNN Recall via Post-training State Expansion
von: Shen, Xingyu, et al.
Veröffentlicht: (2025)
von: Shen, Xingyu, et al.
Veröffentlicht: (2025)
Towards a Theoretical Understanding of Synthetic Data in LLM Post-Training: A Reverse-Bottleneck Perspective
von: Gan, Zeyu, et al.
Veröffentlicht: (2024)
von: Gan, Zeyu, et al.
Veröffentlicht: (2024)
Interactions Across Blocks in Post-Training Quantization of Large Language Models
von: Shabanovi, Khasmamad, et al.
Veröffentlicht: (2024)
von: Shabanovi, Khasmamad, et al.
Veröffentlicht: (2024)
PoTPTQ: A Two-step Power-of-Two Post-training for LLMs
von: Wang, Xinyu, et al.
Veröffentlicht: (2025)
von: Wang, Xinyu, et al.
Veröffentlicht: (2025)
The Price of Agreement: Measuring LLM Sycophancy in Agentic Financial Applications
von: Zhao, Zhenyu, et al.
Veröffentlicht: (2026)
von: Zhao, Zhenyu, et al.
Veröffentlicht: (2026)
Arena Learning: Build Data Flywheel for LLMs Post-training via Simulated Chatbot Arena
von: Luo, Haipeng, et al.
Veröffentlicht: (2024)
von: Luo, Haipeng, et al.
Veröffentlicht: (2024)
ProofOptimizer: Training Language Models to Simplify Proofs without Human Demonstrations
von: Gu, Alex, et al.
Veröffentlicht: (2025)
von: Gu, Alex, et al.
Veröffentlicht: (2025)
Search, Verify and Feedback: Towards Next Generation Post-training Paradigm of Foundation Models via Verifier Engineering
von: Guan, Xinyan, et al.
Veröffentlicht: (2024)
von: Guan, Xinyan, et al.
Veröffentlicht: (2024)
RadioRAG: Online Retrieval-augmented Generation for Radiology Question Answering
von: Arasteh, Soroosh Tayebi, et al.
Veröffentlicht: (2024)
von: Arasteh, Soroosh Tayebi, et al.
Veröffentlicht: (2024)
P$^2$ Law: Scaling Law for Post-Training After Model Pruning
von: Chen, Xiaodong, et al.
Veröffentlicht: (2024)
von: Chen, Xiaodong, et al.
Veröffentlicht: (2024)
KALAVAI: Predicting When Independent Specialist Fusion Works -- A Quantitative Model for Post-Hoc Cooperative LLM Training
von: Kumaresan, Ramchand
Veröffentlicht: (2026)
von: Kumaresan, Ramchand
Veröffentlicht: (2026)
Post-training makes large language models less human-like
von: Binz, Marcel, et al.
Veröffentlicht: (2026)
von: Binz, Marcel, et al.
Veröffentlicht: (2026)
Tracing the Representation Geometry of Language Models from Pretraining to Post-training
von: Li, Melody Zixuan, et al.
Veröffentlicht: (2025)
von: Li, Melody Zixuan, et al.
Veröffentlicht: (2025)
Post-Training Sparse Attention with Double Sparsity
von: Yang, Shuo, et al.
Veröffentlicht: (2024)
von: Yang, Shuo, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Self-Improving Pretraining: using post-trained models to pretrain better models
von: Tan, Ellen Xiaoqing, et al.
Veröffentlicht: (2026) -
J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning
von: Whitehouse, Chenxi, et al.
Veröffentlicht: (2025) -
Backtracking Improves Generation Safety
von: Zhang, Yiming, et al.
Veröffentlicht: (2024) -
Learning Facts at Scale with Active Reading
von: Lin, Jessy, et al.
Veröffentlicht: (2025) -
Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training
von: Liu, Yixin, et al.
Veröffentlicht: (2026)