Breaking the Capability Ceiling of LLM Post-Training by Reintroducing Markov States
Fuente:
arXiv
Salvato in:
| Autori principali: | Yuan, Yurun, Xie, Tengyang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Reinforce LLM Reasoning through Multi-Agent Reflection
di: Yuan, Yurun, et al.
Pubblicazione: (2025)
di: Yuan, Yurun, et al.
Pubblicazione: (2025)
Trajectory Bellman Residual Minimization: A Simple Value-Based Method for LLM Reasoning
di: Yuan, Yurun, et al.
Pubblicazione: (2025)
di: Yuan, Yurun, et al.
Pubblicazione: (2025)
MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate
di: Wang, Jianze, et al.
Pubblicazione: (2026)
di: Wang, Jianze, et al.
Pubblicazione: (2026)
Self-Play with Adversarial Critic: Provable and Scalable Offline Alignment for Language Models
di: Ji, Xiang, et al.
Pubblicazione: (2024)
di: Ji, Xiang, et al.
Pubblicazione: (2024)
BiLLM: Pushing the Limit of Post-Training Quantization for LLMs
di: Huang, Wei, et al.
Pubblicazione: (2024)
di: Huang, Wei, et al.
Pubblicazione: (2024)
Post-training an LLM for RAG? Train on Self-Generated Demonstrations
di: Finlayson, Matthew, et al.
Pubblicazione: (2025)
di: Finlayson, Matthew, et al.
Pubblicazione: (2025)
Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
di: Rosset, Corby, et al.
Pubblicazione: (2024)
di: Rosset, Corby, et al.
Pubblicazione: (2024)
SafePred: A Predictive Guardrail for Computer-Using Agents via World Models
di: Chen, Yurun, et al.
Pubblicazione: (2026)
di: Chen, Yurun, et al.
Pubblicazione: (2026)
Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training
di: Liu, Yixin, et al.
Pubblicazione: (2026)
di: Liu, Yixin, et al.
Pubblicazione: (2026)
PITA: Preference-Guided Inference-Time Alignment for LLM Post-Training
di: Bobbili, Sarat Chandra, et al.
Pubblicazione: (2025)
di: Bobbili, Sarat Chandra, et al.
Pubblicazione: (2025)
$Q\sharp$: Provably Optimal Distributional RL for LLM Post-Training
di: Zhou, Jin Peng, et al.
Pubblicazione: (2025)
di: Zhou, Jin Peng, et al.
Pubblicazione: (2025)
The Depth Ceiling: On the Limits of Large Language Models in Discovering Latent Planning
di: Xu, Yi, et al.
Pubblicazione: (2026)
di: Xu, Yi, et al.
Pubblicazione: (2026)
Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF
di: Xie, Tengyang, et al.
Pubblicazione: (2024)
di: Xie, Tengyang, et al.
Pubblicazione: (2024)
R$^2$PO: Decoupling Training Trajectories from Inference Responses for LLM Reasoning
di: Wang, Jingchu, et al.
Pubblicazione: (2026)
di: Wang, Jingchu, et al.
Pubblicazione: (2026)
ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
di: Lu, Jiarui, et al.
Pubblicazione: (2024)
di: Lu, Jiarui, et al.
Pubblicazione: (2024)
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training
di: Pan, Chengjun, et al.
Pubblicazione: (2026)
di: Pan, Chengjun, et al.
Pubblicazione: (2026)
CeRA: Overcoming the Linear Ceiling of Low-Rank Adaptation via Capacity Expansion
di: Chen, Hung-Hsuan
Pubblicazione: (2026)
di: Chen, Hung-Hsuan
Pubblicazione: (2026)
GraphDancer: Training LLMs to Explore and Reason over Graphs via Two-Stage Curriculum Post-Training
di: Bai, Yuyang, et al.
Pubblicazione: (2026)
di: Bai, Yuyang, et al.
Pubblicazione: (2026)
AMAQ: Adaptive Mixed-bit Activation Quantization for Collaborative Parameter Efficient Fine-tuning
di: Song, Yurun, et al.
Pubblicazione: (2025)
di: Song, Yurun, et al.
Pubblicazione: (2025)
ShiftAddLLM: Accelerating Pretrained LLMs via Post-Training Multiplication-Less Reparameterization
di: You, Haoran, et al.
Pubblicazione: (2024)
di: You, Haoran, et al.
Pubblicazione: (2024)
CounterCurate: Enhancing Physical and Semantic Visio-Linguistic Compositional Reasoning via Counterfactual Examples
di: Zhang, Jianrui, et al.
Pubblicazione: (2024)
di: Zhang, Jianrui, et al.
Pubblicazione: (2024)
Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization
di: Huang, Audrey, et al.
Pubblicazione: (2024)
di: Huang, Audrey, et al.
Pubblicazione: (2024)
Offline Reinforcement Learning in Large State Spaces: Algorithms and Guarantees
di: Jiang, Nan, et al.
Pubblicazione: (2025)
di: Jiang, Nan, et al.
Pubblicazione: (2025)
Towards a Theoretical Understanding of Synthetic Data in LLM Post-Training: A Reverse-Bottleneck Perspective
di: Gan, Zeyu, et al.
Pubblicazione: (2024)
di: Gan, Zeyu, et al.
Pubblicazione: (2024)
Atom of Thoughts for Markov LLM Test-Time Scaling
di: Teng, Fengwei, et al.
Pubblicazione: (2025)
di: Teng, Fengwei, et al.
Pubblicazione: (2025)
ModelGPT: Unleashing LLM's Capabilities for Tailored Model Generation
di: Tang, Zihao, et al.
Pubblicazione: (2024)
di: Tang, Zihao, et al.
Pubblicazione: (2024)
KALAVAI: Predicting When Independent Specialist Fusion Works -- A Quantitative Model for Post-Hoc Cooperative LLM Training
di: Kumaresan, Ramchand
Pubblicazione: (2026)
di: Kumaresan, Ramchand
Pubblicazione: (2026)
Learning to Reason Efficiently with A* Post-Training
di: Opedal, Andreas, et al.
Pubblicazione: (2026)
di: Opedal, Andreas, et al.
Pubblicazione: (2026)
Post-Training Sparse Attention with Double Sparsity
di: Yang, Shuo, et al.
Pubblicazione: (2024)
di: Yang, Shuo, et al.
Pubblicazione: (2024)
When Greedy Wins: Emergent Exploitation Bias in Meta-Bandit LLM Training
di: Chen, Sanxing, et al.
Pubblicazione: (2025)
di: Chen, Sanxing, et al.
Pubblicazione: (2025)
Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators
di: Roytburg, Dani, et al.
Pubblicazione: (2025)
di: Roytburg, Dani, et al.
Pubblicazione: (2025)
Capability Instruction Tuning: A New Paradigm for Dynamic LLM Routing
di: Zhang, Yi-Kai, et al.
Pubblicazione: (2025)
di: Zhang, Yi-Kai, et al.
Pubblicazione: (2025)
Can Post-Training Transform LLMs into Causal Reasoners?
di: Chen, Junqi, et al.
Pubblicazione: (2026)
di: Chen, Junqi, et al.
Pubblicazione: (2026)
Mapping Post-Training Forgetting in Language Models at Scale
di: Harmon, Jackson, et al.
Pubblicazione: (2025)
di: Harmon, Jackson, et al.
Pubblicazione: (2025)
Breaking the Ceiling of the LLM Community by Treating Token Generation as a Classification for Ensembling
di: Yu, Yao-Ching, et al.
Pubblicazione: (2024)
di: Yu, Yao-Ching, et al.
Pubblicazione: (2024)
Muon is Scalable for LLM Training
di: Liu, Jingyuan, et al.
Pubblicazione: (2025)
di: Liu, Jingyuan, et al.
Pubblicazione: (2025)
LoopTool: Closing the Data-Training Loop for Robust LLM Tool Calls
di: Zhang, Kangning, et al.
Pubblicazione: (2025)
di: Zhang, Kangning, et al.
Pubblicazione: (2025)
Lil: Less is Less When Applying Post-Training Sparse-Attention Algorithms in Long-Decode Stage
di: Hu, Junhao, et al.
Pubblicazione: (2026)
di: Hu, Junhao, et al.
Pubblicazione: (2026)
Tuning without Peeking: Provable Generalization Bounds and Robust LLM Post-Training
di: Labiad, Ismail, et al.
Pubblicazione: (2025)
di: Labiad, Ismail, et al.
Pubblicazione: (2025)
Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability
di: Lin, Zicheng, et al.
Pubblicazione: (2024)
di: Lin, Zicheng, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Reinforce LLM Reasoning through Multi-Agent Reflection
di: Yuan, Yurun, et al.
Pubblicazione: (2025) -
Trajectory Bellman Residual Minimization: A Simple Value-Based Method for LLM Reasoning
di: Yuan, Yurun, et al.
Pubblicazione: (2025) -
MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate
di: Wang, Jianze, et al.
Pubblicazione: (2026) -
Self-Play with Adversarial Critic: Provable and Scalable Offline Alignment for Language Models
di: Ji, Xiang, et al.
Pubblicazione: (2024) -
BiLLM: Pushing the Limit of Post-Training Quantization for LLMs
di: Huang, Wei, et al.
Pubblicazione: (2024)