GAPO: Robust Advantage Estimation for Real-World Code LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Jianqing, Hao, Zhezheng, Xia, Wei, Dong, Hande, Wang, Hong, Wei, Chenxing, Zhou, Yuyan, Qi, Yubin, Lin, Qiang, Cao, Jian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LEPO: Latent Reasoning Policy Optimization for Large Language Models
by: Zhou, Yuyan, et al.
Published: (2026)
by: Zhou, Yuyan, et al.
Published: (2026)
AP2O-Coder: Adaptively Progressive Preference Optimization for Reducing Compilation and Runtime Errors in LLM-Generated Code
by: Zhang, Jianqing, et al.
Published: (2025)
by: Zhang, Jianqing, et al.
Published: (2025)
ReCreate: Reasoning and Creating Domain Agents Driven by Experience
by: Hao, Zhezheng, et al.
Published: (2026)
by: Hao, Zhezheng, et al.
Published: (2026)
Scheduling Your LLM Reinforcement Learning with Reasoning Trees
by: Wang, Hong, et al.
Published: (2025)
by: Wang, Hong, et al.
Published: (2025)
Scaling Human-AI Coding Collaboration Requires a Governable Consensus Layer
by: Wang, Tianfu, et al.
Published: (2026)
by: Wang, Tianfu, et al.
Published: (2026)
Unleashing the True Potential of LLMs: A Feedback-Triggered Self-Correction with Long-Term Multipath Decoding
by: Li, Jipeng, et al.
Published: (2025)
by: Li, Jipeng, et al.
Published: (2025)
Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective
by: Hao, Zhezheng, et al.
Published: (2025)
by: Hao, Zhezheng, et al.
Published: (2025)
Evolve as a Team: Collaborative Self-Evolution for LLM-based Multi-Agent Systems
by: Hao, Zhezheng, et al.
Published: (2026)
by: Hao, Zhezheng, et al.
Published: (2026)
UniSVG: A Unified Dataset for Vector Graphic Understanding and Generation with Multimodal Large Language Models
by: Li, Jinke, et al.
Published: (2025)
by: Li, Jinke, et al.
Published: (2025)
Why Classical Dentures Are a Success in GAPO Patients?
by: Mohamed A. Abdel‐Kader, et al.
Published: (2025)
by: Mohamed A. Abdel‐Kader, et al.
Published: (2025)
Accelerating IC Thermal Simulation Data Generation via Block Krylov and Operator Action
by: Wang, Hong, et al.
Published: (2025)
by: Wang, Hong, et al.
Published: (2025)
AAPO: Enhancing the Reasoning Capabilities of LLMs with Advantage Margin
by: Xiong, Jian, et al.
Published: (2025)
by: Xiong, Jian, et al.
Published: (2025)
ReDit: Reward Dithering for Improved LLM Policy Optimization
by: Wei, Chenxing, et al.
Published: (2025)
by: Wei, Chenxing, et al.
Published: (2025)
GAPO: Learning Preferential Prompt through Generative Adversarial Policy Optimization
by: Gu, Zhouhong, et al.
Published: (2025)
by: Gu, Zhouhong, et al.
Published: (2025)
DSO: Dual-Scale Neural Operators for Stable Long-term Fluid Dynamics Forecasting
by: Dong, Huanshuo, et al.
Published: (2026)
by: Dong, Huanshuo, et al.
Published: (2026)
LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs
by: Xia, Yunhui, et al.
Published: (2025)
by: Xia, Yunhui, et al.
Published: (2025)
Scaling Laws Behind Code Understanding Model
by: Lin, Jiayi, et al.
Published: (2024)
by: Lin, Jiayi, et al.
Published: (2024)
Test-Time Policy Adaptation for Enhanced Multi-Turn Interactions with LLMs
by: Wei, Chenxing, et al.
Published: (2025)
by: Wei, Chenxing, et al.
Published: (2025)
Multi-class Support Vector Machine with Maximizing Minimum Margin
by: Hao, Zhezheng, et al.
Published: (2023)
by: Hao, Zhezheng, et al.
Published: (2023)
Ensemble-Based Uncertainty Estimation for Code Correctness Estimation
by: Wei, Yunxiang, et al.
Published: (2026)
by: Wei, Yunxiang, et al.
Published: (2026)
Real‐Time Decoding For Surface Code
by: Jia‐Ning Li, et al.
Published: (2026)
by: Jia‐Ning Li, et al.
Published: (2026)
Cutoff for the Swendsen-Wang dynamics on the complete graph
by: Blanca, Antonio, et al.
Published: (2025)
by: Blanca, Antonio, et al.
Published: (2025)
Real-to-Sim Grasp: Rethinking the Gap between Simulation and Real World in Grasp Detection
by: Cai, Jia-Feng, et al.
Published: (2024)
by: Cai, Jia-Feng, et al.
Published: (2024)
Learned or Memorized ? Quantifying Memorization Advantage in Code LLMs
by: Euraste, Djiré Albérick, et al.
Published: (2026)
by: Euraste, Djiré Albérick, et al.
Published: (2026)
Hierarchical Context Pruning: Optimizing Real-World Code Completion with Repository-Level Pretrained Code LLMs
by: Zhang, Lei, et al.
Published: (2024)
by: Zhang, Lei, et al.
Published: (2024)
Code2Worlds: Empowering Coding LLMs for 4D World Generation
by: Zhang, Yi, et al.
Published: (2026)
by: Zhang, Yi, et al.
Published: (2026)
Adaptive Guidance for Local Training in Heterogeneous Federated Learning
by: Zhang, Jianqing, et al.
Published: (2024)
by: Zhang, Jianqing, et al.
Published: (2024)
Operational Robustness of LLMs on Code Generation
by: Paul, Debalina Ghosh, et al.
Published: (2026)
by: Paul, Debalina Ghosh, et al.
Published: (2026)
Evaluating LLMs Code Reasoning Under Real-World Context
by: Liu, Changshu
Published: (2026)
by: Liu, Changshu
Published: (2026)
Robust Bandwidth Estimation for Real-Time Communication with Offline Reinforcement Learning
by: Kai, Jian, et al.
Published: (2025)
by: Kai, Jian, et al.
Published: (2025)
Low-Complexity Channel Estimation for RIS-Assisted ISAC System
by: Zhen, Chen, et al.
Published: (2025)
by: Zhen, Chen, et al.
Published: (2025)
Exponential Advantage from One More Replica in Estimating Nonlinear Properties of Quantum States
by: Ye, Qi, et al.
Published: (2025)
by: Ye, Qi, et al.
Published: (2025)
Exploring Robustness of Multilingual LLMs on Real-World Noisy Data
by: Aliakbarzadeh, Amirhossein, et al.
Published: (2025)
by: Aliakbarzadeh, Amirhossein, et al.
Published: (2025)
Why Tree-Style Branching Matters for Thought Advantage Estimation in GRPO
by: Wang, Hongcheng, et al.
Published: (2025)
by: Wang, Hongcheng, et al.
Published: (2025)
A Real‐World Pharmacovigilance Study of Fruquintinib Based on the FDA Adverse Event Reporting System ( FAERS ) Database
by: Yajing Xu, et al.
Published: (2025)
by: Yajing Xu, et al.
Published: (2025)
SDA: Steering-Driven Distribution Alignment for Open LLMs without Fine-Tuning
by: Xia, Wei, et al.
Published: (2025)
by: Xia, Wei, et al.
Published: (2025)
Improving Code Search with Hard Negative Sampling Based on Fine-tuning
by: Dong, Hande, et al.
Published: (2023)
by: Dong, Hande, et al.
Published: (2023)
Cardiovascular Safety Landscape of ADT in Prostate Cancer Treatment Based on Real‐World Analysis
by: Wei Wang, et al.
Published: (2025)
by: Wei Wang, et al.
Published: (2025)
KTAE: A Model-Free Algorithm to Key-Tokens Advantage Estimation in Mathematical Reasoning
by: Sun, Wei, et al.
Published: (2025)
by: Sun, Wei, et al.
Published: (2025)
CodeSense: a Real-World Benchmark and Dataset for Code Semantic Reasoning
by: Roy, Monoshi Kumar, et al.
Published: (2025)
by: Roy, Monoshi Kumar, et al.
Published: (2025)
Similar Items
-
LEPO: Latent Reasoning Policy Optimization for Large Language Models
by: Zhou, Yuyan, et al.
Published: (2026) -
AP2O-Coder: Adaptively Progressive Preference Optimization for Reducing Compilation and Runtime Errors in LLM-Generated Code
by: Zhang, Jianqing, et al.
Published: (2025) -
ReCreate: Reasoning and Creating Domain Agents Driven by Experience
by: Hao, Zhezheng, et al.
Published: (2026) -
Scheduling Your LLM Reinforcement Learning with Reasoning Trees
by: Wang, Hong, et al.
Published: (2025) -
Scaling Human-AI Coding Collaboration Requires a Governable Consensus Layer
by: Wang, Tianfu, et al.
Published: (2026)