Skip-Connected Policy Optimization for Implicit Advantage
Fuente:
arXiv
Salvato in:
| Autori principali: | Teng, Fengwei, Bai, Jinyi, Yao, Xinhao, Wang, Demi Ruohan, Zhao, Jiahao, Guo, Zhijiang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Atom of Thoughts for Markov LLM Test-Time Scaling
di: Teng, Fengwei, et al.
Pubblicazione: (2025)
di: Teng, Fengwei, et al.
Pubblicazione: (2025)
GAGPO: Generalized Advantage Grouped Policy Optimization
di: Zhu, Siyuan, et al.
Pubblicazione: (2026)
di: Zhu, Siyuan, et al.
Pubblicazione: (2026)
REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
di: Hu, Jian, et al.
Pubblicazione: (2025)
di: Hu, Jian, et al.
Pubblicazione: (2025)
Probe and Skip: Self-Predictive Token Skipping for Efficient Long-Context LLM Inference
di: Wu, Zimeng, et al.
Pubblicazione: (2026)
di: Wu, Zimeng, et al.
Pubblicazione: (2026)
DASH: Input-Aware Dynamic Layer Skipping for Efficient LLM Inference with Markov Decision Policies
di: Yang, Ning, et al.
Pubblicazione: (2025)
di: Yang, Ning, et al.
Pubblicazione: (2025)
FFN-SkipLLM: A Hidden Gem for Autoregressive Decoding with Adaptive Feed Forward Skipping
di: Jaiswal, Ajay, et al.
Pubblicazione: (2024)
di: Jaiswal, Ajay, et al.
Pubblicazione: (2024)
Efficient RLVR Training via Weighted Mutual Information Data Selection
di: Zhou, Xinyu, et al.
Pubblicazione: (2026)
di: Zhou, Xinyu, et al.
Pubblicazione: (2026)
Learning to Skip the Middle Layers of Transformers
di: Lawson, Tim, et al.
Pubblicazione: (2025)
di: Lawson, Tim, et al.
Pubblicazione: (2025)
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections
di: Wang, Bo, et al.
Pubblicazione: (2025)
di: Wang, Bo, et al.
Pubblicazione: (2025)
When Inverse Data Outperforms: Exploring the Pitfalls of Mixed Data in Multi-Stage Fine-Tuning
di: Deng, Mengyi, et al.
Pubblicazione: (2025)
di: Deng, Mengyi, et al.
Pubblicazione: (2025)
Modeling Bilingual Sentence Processing: Evaluating RNN and Transformer Architectures for Cross-Language Structural Priming
di: Zhang, Demi, et al.
Pubblicazione: (2024)
di: Zhang, Demi, et al.
Pubblicazione: (2024)
IIET: Efficient Numerical Transformer via Implicit Iterative Euler Method
di: Liu, Xinyu, et al.
Pubblicazione: (2025)
di: Liu, Xinyu, et al.
Pubblicazione: (2025)
Stabilizing Policy Optimization via Logits Convexity
di: Chen, Hongzhan, et al.
Pubblicazione: (2026)
di: Chen, Hongzhan, et al.
Pubblicazione: (2026)
Skipformer: A Skip-and-Recover Strategy for Efficient Speech Recognition
di: Zhu, Wenjing, et al.
Pubblicazione: (2024)
di: Zhu, Wenjing, et al.
Pubblicazione: (2024)
DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning
di: Jiang, Guochao, et al.
Pubblicazione: (2026)
di: Jiang, Guochao, et al.
Pubblicazione: (2026)
AutoPSV: Automated Process-Supervised Verifier
di: Lu, Jianqiao, et al.
Pubblicazione: (2024)
di: Lu, Jianqiao, et al.
Pubblicazione: (2024)
On SkipGram Word Embedding Models with Negative Sampling: Unified Framework and Impact of Noise Distributions
di: Liu, Dezhi, et al.
Pubblicazione: (2020)
di: Liu, Dezhi, et al.
Pubblicazione: (2020)
Multi-Objective and Mixed-Reward Reinforcement Learning via Reward-Decorrelated Policy Optimization
di: Bai, Yang, et al.
Pubblicazione: (2026)
di: Bai, Yang, et al.
Pubblicazione: (2026)
IAPO: Information-Aware Policy Optimization for Token-Efficient Reasoning
di: He, Yinhan, et al.
Pubblicazione: (2026)
di: He, Yinhan, et al.
Pubblicazione: (2026)
PoSE: Efficient Context Window Extension of LLMs via Positional Skip-wise Training
di: Zhu, Dawei, et al.
Pubblicazione: (2023)
di: Zhu, Dawei, et al.
Pubblicazione: (2023)
Self-Supervised Prompt Optimization
di: Xiang, Jinyu, et al.
Pubblicazione: (2025)
di: Xiang, Jinyu, et al.
Pubblicazione: (2025)
ConfLayers: Adaptive Confidence-based Layer Skipping for Self-Speculative Decoding
di: Amer, Walaa, et al.
Pubblicazione: (2026)
di: Amer, Walaa, et al.
Pubblicazione: (2026)
Implicit Optimization Bias of Next-Token Prediction in Linear Models
di: Thrampoulidis, Christos
Pubblicazione: (2024)
di: Thrampoulidis, Christos
Pubblicazione: (2024)
On the Role of Preference Variance in Preference Optimization
di: Guo, Jiacheng, et al.
Pubblicazione: (2025)
di: Guo, Jiacheng, et al.
Pubblicazione: (2025)
On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization
di: Lin, Yong, et al.
Pubblicazione: (2024)
di: Lin, Yong, et al.
Pubblicazione: (2024)
InfiFPO: Implicit Model Fusion via Preference Optimization in Large Language Models
di: Gu, Yanggan, et al.
Pubblicazione: (2025)
di: Gu, Yanggan, et al.
Pubblicazione: (2025)
Causally-Enhanced Reinforcement Policy Optimization
di: Wang, Xiangqi, et al.
Pubblicazione: (2025)
di: Wang, Xiangqi, et al.
Pubblicazione: (2025)
Stabilizing Efficient Reasoning with Step-Level Advantage Selection
di: Wang, Han, et al.
Pubblicazione: (2026)
di: Wang, Han, et al.
Pubblicazione: (2026)
Soft Adaptive Policy Optimization
di: Gao, Chang, et al.
Pubblicazione: (2025)
di: Gao, Chang, et al.
Pubblicazione: (2025)
EPO: Entropy-regularized Policy Optimization for LLM Agents Reinforcement Learning
di: Xu, Wujiang, et al.
Pubblicazione: (2025)
di: Xu, Wujiang, et al.
Pubblicazione: (2025)
Merging Beyond: Streaming LLM Updates via Activation-Guided Rotations
di: Yao, Yuxuan, et al.
Pubblicazione: (2026)
di: Yao, Yuxuan, et al.
Pubblicazione: (2026)
Outcome-Grounded Advantage Reshaping for Fine-Grained Credit Assignment in Mathematical Reasoning
di: Li, Ziheng, et al.
Pubblicazione: (2026)
di: Li, Ziheng, et al.
Pubblicazione: (2026)
Diffusion LMs Can Approximate Optimal Infilling Lengths Implicitly
di: Liu, Hengchang, et al.
Pubblicazione: (2026)
di: Liu, Hengchang, et al.
Pubblicazione: (2026)
AAPO: Enhancing the Reasoning Capabilities of LLMs with Advantage Margin
di: Xiong, Jian, et al.
Pubblicazione: (2025)
di: Xiong, Jian, et al.
Pubblicazione: (2025)
Skip-It? Theoretical Conditions for Layer Skipping in Vision-Language Models
di: Hartman, Max, et al.
Pubblicazione: (2025)
di: Hartman, Max, et al.
Pubblicazione: (2025)
DCPO: Dynamic Clipping Policy Optimization
di: Yang, Shihui, et al.
Pubblicazione: (2025)
di: Yang, Shihui, et al.
Pubblicazione: (2025)
MOOSComp: Improving Lightweight Long-Context Compressor via Mitigating Over-Smoothing and Incorporating Outlier Scores
di: Zhou, Fengwei, et al.
Pubblicazione: (2025)
di: Zhou, Fengwei, et al.
Pubblicazione: (2025)
Skip \n: A Simple Method to Reduce Hallucination in Large Vision-Language Models
di: Han, Zongbo, et al.
Pubblicazione: (2024)
di: Han, Zongbo, et al.
Pubblicazione: (2024)
Position-Agnostic Pre-Projection for Transformer Attention: Nonlinear Feature Construction and Content Skip Before Q/K/V
di: Shinde, Chirag
Pubblicazione: (2026)
di: Shinde, Chirag
Pubblicazione: (2026)
Understanding Reference Policies in Direct Preference Optimization
di: Liu, Yixin, et al.
Pubblicazione: (2024)
di: Liu, Yixin, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Atom of Thoughts for Markov LLM Test-Time Scaling
di: Teng, Fengwei, et al.
Pubblicazione: (2025) -
GAGPO: Generalized Advantage Grouped Policy Optimization
di: Zhu, Siyuan, et al.
Pubblicazione: (2026) -
REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
di: Hu, Jian, et al.
Pubblicazione: (2025) -
Probe and Skip: Self-Predictive Token Skipping for Efficient Long-Context LLM Inference
di: Wu, Zimeng, et al.
Pubblicazione: (2026) -
DASH: Input-Aware Dynamic Layer Skipping for Efficient LLM Inference with Markov Decision Policies
di: Yang, Ning, et al.
Pubblicazione: (2025)