Token-level Proximal Policy Optimization for Query Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Ouyang, Yichen, Wang, Lu, Yang, Fangkai, Zhao, Pu, Huang, Chenghua, Liu, Jianfeng, Pang, Bochen, Yang, Yaming, Zhan, Yuefeng, Sun, Hao, Lin, Qingwei, Rajmohan, Saravan, Deng, Weiwei, Zhang, Dongmei, Sun, Feng, Zhang, Qi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Pretrain Value, Not Reward: Decoupled Value Policy Optimization
by: Huang, Chenghua, et al.
Published: (2025)
by: Huang, Chenghua, et al.
Published: (2025)
Self-Evolved Reward Learning for LLMs
by: Huang, Chenghua, et al.
Published: (2024)
by: Huang, Chenghua, et al.
Published: (2024)
LettinGo: Explore User Profile Generation for Recommendation System
by: Wang, Lu, et al.
Published: (2025)
by: Wang, Lu, et al.
Published: (2025)
DoVer: Intervention-Driven Auto Debugging for LLM Multi-Agent Systems
by: Ma, Ming, et al.
Published: (2025)
by: Ma, Ming, et al.
Published: (2025)
WarriorMath: Enhancing the Mathematical Ability of Large Language Models with a Defect-aware Framework
by: Chen, Yue, et al.
Published: (2025)
by: Chen, Yue, et al.
Published: (2025)
From Reasoning to Answer: Empirical, Attention-Based and Mechanistic Insights into Distilled DeepSeek R1 Models
by: Zhang, Jue, et al.
Published: (2025)
by: Zhang, Jue, et al.
Published: (2025)
WarriorCoder: Learning from Expert Battles to Augment Code Large Language Models
by: Feng, Huawen, et al.
Published: (2024)
by: Feng, Huawen, et al.
Published: (2024)
The Vision of Autonomic Computing: Can LLMs Make It a Reality?
by: Zhang, Zhiyang, et al.
Published: (2024)
by: Zhang, Zhiyang, et al.
Published: (2024)
ExeCoder: Empowering Large Language Models with Executability Representation for Code Translation
by: He, Minghua, et al.
Published: (2025)
by: He, Minghua, et al.
Published: (2025)
Learning to Refine: Self-Refinement of Parallel Reasoning in LLMs
by: Wang, Qibin, et al.
Published: (2025)
by: Wang, Qibin, et al.
Published: (2025)
AXIS: Efficient Human-Agent-Computer Interaction with API-First LLM-Based Agents
by: Lu, Junting, et al.
Published: (2024)
by: Lu, Junting, et al.
Published: (2024)
AutoRAG-HP: Automatic Online Hyper-Parameter Tuning for Retrieval-Augmented Generation
by: Fu, Jia, et al.
Published: (2024)
by: Fu, Jia, et al.
Published: (2024)
VEM: Environment-Free Exploration for Training GUI Agent with Value Environment Model
by: Zheng, Jiani, et al.
Published: (2025)
by: Zheng, Jiani, et al.
Published: (2025)
EfficientRAG: Efficient Retriever for Multi-Hop Question Answering
by: Zhuang, Ziyuan, et al.
Published: (2024)
by: Zhuang, Ziyuan, et al.
Published: (2024)
DUET: Joint Exploration of User Item Profiles in Recommendation System
by: Chen, Yue, et al.
Published: (2026)
by: Chen, Yue, et al.
Published: (2026)
RepoGenesis: Benchmarking End-to-End Microservice Generation from Readme to Repository
by: Peng, Zhiyuan, et al.
Published: (2026)
by: Peng, Zhiyuan, et al.
Published: (2026)
LLM Reasoning as Trajectories: Step-Specific Representation Geometry and Correctness Signals
by: Sun, Lihao, et al.
Published: (2026)
by: Sun, Lihao, et al.
Published: (2026)
COIN: Chance-Constrained Imitation Learning for Uncertainty-aware Adaptive Resource Oversubscription Policy
by: Wang, Lu, et al.
Published: (2024)
by: Wang, Lu, et al.
Published: (2024)
AdNanny: One Reasoning LLM for All Offline Ads Recommendation Tasks
by: Hu, Nan, et al.
Published: (2026)
by: Hu, Nan, et al.
Published: (2026)
RePrompt: Reasoning-Augmented Reprompting for Text-to-Image Generation via Reinforcement Learning
by: Wu, Mingrui, et al.
Published: (2025)
by: Wu, Mingrui, et al.
Published: (2025)
Enabling Autonomic Microservice Management through Self-Learning Agents
by: Yu, Fenglin, et al.
Published: (2025)
by: Yu, Fenglin, et al.
Published: (2025)
AdaptFlow: Adaptive Workflow Optimization via Meta-Learning
by: Zhu, Runchuan, et al.
Published: (2025)
by: Zhu, Runchuan, et al.
Published: (2025)
AI Delegates with a Dual Focus: Ensuring Privacy and Strategic Self-Disclosure
by: Zhang, Zhiyang, et al.
Published: (2024)
by: Zhang, Zhiyang, et al.
Published: (2024)
A Tale of Two Graphs: Separating Knowledge Exploration from Outline Structure for Open-Ended Deep Research
by: Shi, Zhuofan, et al.
Published: (2026)
by: Shi, Zhuofan, et al.
Published: (2026)
Contrastive Attribution in the Wild: An Interpretability Analysis of LLM Failures on Realistic Benchmarks
by: Tan, Rongyuan, et al.
Published: (2026)
by: Tan, Rongyuan, et al.
Published: (2026)
Beyond State Consistency: Behavior Consistency in Text-Based World Models
by: Huang, Youling, et al.
Published: (2026)
by: Huang, Youling, et al.
Published: (2026)
Call Me When Necessary: LLMs can Efficiently and Faithfully Reason over Structured Environments
by: Cheng, Sitao, et al.
Published: (2024)
by: Cheng, Sitao, et al.
Published: (2024)
Nissist: An Incident Mitigation Copilot based on Troubleshooting Guides
by: An, Kaikai, et al.
Published: (2024)
by: An, Kaikai, et al.
Published: (2024)
Skeleton-Guided-Translation: A Benchmarking Framework for Code Repository Translation with Fine-Grained Quality Evaluation
by: Zhang, Xing, et al.
Published: (2025)
by: Zhang, Xing, et al.
Published: (2025)
Navigating the Unknown: A Chat-Based Collaborative Interface for Personalized Exploratory Tasks
by: Peng, Yingzhe, et al.
Published: (2024)
by: Peng, Yingzhe, et al.
Published: (2024)
Thread: A Logic-Based Data Organization Paradigm for How-To Question Answering with Retrieval Augmented Generation
by: An, Kaikai, et al.
Published: (2024)
by: An, Kaikai, et al.
Published: (2024)
MEETING DELEGATE: Benchmarking LLMs on Attending Meetings on Our Behalf
by: Hu, Lingxiang, et al.
Published: (2025)
by: Hu, Lingxiang, et al.
Published: (2025)
Distill Not Only Data but Also Rewards: Can Smaller Language Models Surpass Larger Ones?
by: Zhang, Yudi, et al.
Published: (2025)
by: Zhang, Yudi, et al.
Published: (2025)
StreamAdapter: Efficient Test Time Adaptation from Contextual Streams
by: Muhtar, Dilxat, et al.
Published: (2024)
by: Muhtar, Dilxat, et al.
Published: (2024)
An Advanced Reinforcement Learning Framework for Online Scheduling of Deferrable Workloads in Cloud Computing
by: Dong, Hang, et al.
Published: (2024)
by: Dong, Hang, et al.
Published: (2024)
Contrastive Learning with Negative Sampling Correction
by: Wang, Lu, et al.
Published: (2024)
by: Wang, Lu, et al.
Published: (2024)
Privacy in Action: Towards Realistic Privacy Mitigation and Evaluation for LLM-Powered Agents
by: Wang, Shouju, et al.
Published: (2025)
by: Wang, Shouju, et al.
Published: (2025)
Text2Grad: Reinforcement Learning from Natural Language Feedback
by: Wang, Hanyang, et al.
Published: (2025)
by: Wang, Hanyang, et al.
Published: (2025)
CI-Work: Benchmarking Contextual Integrity in Enterprise LLM Agents
by: Fu, Wenjie, et al.
Published: (2026)
by: Fu, Wenjie, et al.
Published: (2026)
API Agents vs. GUI Agents: Divergence and Convergence
by: Zhang, Chaoyun, et al.
Published: (2025)
by: Zhang, Chaoyun, et al.
Published: (2025)
Similar Items
-
Pretrain Value, Not Reward: Decoupled Value Policy Optimization
by: Huang, Chenghua, et al.
Published: (2025) -
Self-Evolved Reward Learning for LLMs
by: Huang, Chenghua, et al.
Published: (2024) -
LettinGo: Explore User Profile Generation for Recommendation System
by: Wang, Lu, et al.
Published: (2025) -
DoVer: Intervention-Driven Auto Debugging for LLM Multi-Agent Systems
by: Ma, Ming, et al.
Published: (2025) -
WarriorMath: Enhancing the Mathematical Ability of Large Language Models with a Defect-aware Framework
by: Chen, Yue, et al.
Published: (2025)