Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Xiaozhe, Fang, Xinyu, Ding, Shengyuan, Li, Yang, Li, Linyang, Duan, Haodong, Liu, Qingwen, Chen, Kai |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
NP-Engine: Empowering Optimization Reasoning in Large Language Models with Verifiable Synthetic NP Problems
by: Li, Xiaozhe, et al.
Published: (2025)
by: Li, Xiaozhe, et al.
Published: (2025)
OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces
by: Li, Xiaozhe, et al.
Published: (2026)
by: Li, Xiaozhe, et al.
Published: (2026)
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems
by: Li, Xiaozhe, et al.
Published: (2025)
by: Li, Xiaozhe, et al.
Published: (2025)
Beyond Mode Collapse: Distribution Matching for Diverse Reasoning
by: Li, Xiaozhe, et al.
Published: (2026)
by: Li, Xiaozhe, et al.
Published: (2026)
Timely Machine: Awareness of Time Makes Test-Time Scaling Agentic
by: Ma, Yichuan, et al.
Published: (2026)
by: Ma, Yichuan, et al.
Published: (2026)
Escaping the Context Bottleneck: Active Context Curation for LLM Agents via Reinforcement Learning
by: Li, Xiaozhe, et al.
Published: (2026)
by: Li, Xiaozhe, et al.
Published: (2026)
What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents
by: Li, Xiaozhe, et al.
Published: (2026)
by: Li, Xiaozhe, et al.
Published: (2026)
AutoForge: Automated Environment Synthesis for Agentic Reinforcement Learning
by: Cai, Shihao, et al.
Published: (2025)
by: Cai, Shihao, et al.
Published: (2025)
Bridging Pattern-Aware Complexity with NP-Hard Optimization: A Unifying Framework and Empirical Study
by: Saidi, Olivier
Published: (2025)
by: Saidi, Olivier
Published: (2025)
Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message
by: Duan, Wei, et al.
Published: (2025)
by: Duan, Wei, et al.
Published: (2025)
NPG-Muse: Scaling Long Chain-of-Thought Reasoning with NP-Hard Graph Problems
by: Wang, Yuyao, et al.
Published: (2025)
by: Wang, Yuyao, et al.
Published: (2025)
ChronoForge-RL: Chronological Forging through Reinforcement Learning for Enhanced Video Understanding
by: Chen, Kehua
Published: (2025)
by: Chen, Kehua
Published: (2025)
Redundancy Principles for MLLMs Benchmarks
by: Zhang, Zicheng, et al.
Published: (2025)
by: Zhang, Zicheng, et al.
Published: (2025)
Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks
by: Wang, Chonghua, et al.
Published: (2024)
by: Wang, Chonghua, et al.
Published: (2024)
HardSecBench: Benchmarking the Security Awareness of LLMs for Hardware Code Generation
by: Chen, Qirui, et al.
Published: (2026)
by: Chen, Qirui, et al.
Published: (2026)
M-$LLM^3$REC: A Motivation-Aware User-Item Interaction Framework for Enhancing Recommendation Accuracy with LLMs
by: Chen, Lining, et al.
Published: (2025)
by: Chen, Lining, et al.
Published: (2025)
Strategy-Aware Optimization Modeling with Reasoning LLMs
by: Zhao, Ruiqing, et al.
Published: (2026)
by: Zhao, Ruiqing, et al.
Published: (2026)
Judge Before Answer: Can MLLM Discern the False Premise in Question?
by: Li, Jidong, et al.
Published: (2025)
by: Li, Jidong, et al.
Published: (2025)
HardSATGEN: Understanding the Difficulty of Hard SAT Formula Generation and A Strong Structure-Hardness-Aware Baseline
by: Li, Yang, et al.
Published: (2023)
by: Li, Yang, et al.
Published: (2023)
Klear-AgentForge: Forging Agentic Intelligence through Posttraining Scaling
by: Wang, Qi, et al.
Published: (2025)
by: Wang, Qi, et al.
Published: (2025)
Towards Open-Ended Emotional Support Conversations in LLMs via Reinforcement Learning with Future-Oriented Rewards
by: Yang, Ting, et al.
Published: (2025)
by: Yang, Ting, et al.
Published: (2025)
COINBench: Moving Beyond Individual Perspectives to Collective Intent Understanding
by: Li, Xiaozhe, et al.
Published: (2026)
by: Li, Xiaozhe, et al.
Published: (2026)
Sorting by Strip Swaps is NP-Hard
by: Roy, Swapnoneel, et al.
Published: (2025)
by: Roy, Swapnoneel, et al.
Published: (2025)
Universal NP-Hardness of Clustering under General Utilities
by: Majumdar, Angshul
Published: (2026)
by: Majumdar, Angshul
Published: (2026)
ProActor: Timing-Aware Reinforcement Learning for Proactive Task Scheduling Agents
by: Ding, Lei, et al.
Published: (2026)
by: Ding, Lei, et al.
Published: (2026)
Conformal Symplectic Optimization for Stable Reinforcement Learning
by: Lyu, Yao, et al.
Published: (2024)
by: Lyu, Yao, et al.
Published: (2024)
ConsintBench: Evaluating Language Models on Real-World Consumer Intent Understanding
by: Li, Xiaozhe, et al.
Published: (2025)
by: Li, Xiaozhe, et al.
Published: (2025)
STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens
by: Liu, Shiqi, et al.
Published: (2026)
by: Liu, Shiqi, et al.
Published: (2026)
Research and Design on Intelligent Recognition of Unordered Targets for Robots Based on Reinforcement Learning
by: Mao, Yiting, et al.
Published: (2025)
by: Mao, Yiting, et al.
Published: (2025)
NP-Hard Lower Bound Complexity for Semantic Self-Verification
by: Young, Robin
Published: (2025)
by: Young, Robin
Published: (2025)
Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning
by: Li, Ang, et al.
Published: (2025)
by: Li, Ang, et al.
Published: (2025)
ProMed: Shapley Information Gain Guided Reinforcement Learning for Proactive Medical LLMs
by: Ding, Hongxin, et al.
Published: (2025)
by: Ding, Hongxin, et al.
Published: (2025)
Deep Reinforcement Learning-based Obstacle Avoidance for Robot Movement in Warehouse Environments
by: Li, Keqin, et al.
Published: (2024)
by: Li, Keqin, et al.
Published: (2024)
Yukthi Opus: A Multi-Chain Hybrid Metaheuristic for Large-Scale NP-Hard Optimization
by: Vikraman, SB Danush, et al.
Published: (2026)
by: Vikraman, SB Danush, et al.
Published: (2026)
Jailbreak-R1: Exploring the Jailbreak Capabilities of LLMs via Reinforcement Learning
by: Guo, Weiyang, et al.
Published: (2025)
by: Guo, Weiyang, et al.
Published: (2025)
SATURN: SAT-based Reinforcement Learning to Unleash LLMs Reasoning
by: Liu, Huanyu, et al.
Published: (2025)
by: Liu, Huanyu, et al.
Published: (2025)
CudaForge: An Agent Framework with Hardware Feedback for CUDA Kernel Optimization
by: Zhang, Zijian, et al.
Published: (2025)
by: Zhang, Zijian, et al.
Published: (2025)
On Token's Dilemma: Dynamic MoE with Drift-Aware Token Assignment for Continual Learning of Large Vision Language Models
by: Zhao, Chongyang, et al.
Published: (2026)
by: Zhao, Chongyang, et al.
Published: (2026)
Knowledge-Guided Prompt Learning for Request Quality Assurance in Public Code Review
by: Li, Lin, et al.
Published: (2024)
by: Li, Lin, et al.
Published: (2024)
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach
by: Zhang, Xinnan, et al.
Published: (2025)
by: Zhang, Xinnan, et al.
Published: (2025)
Similar Items
-
NP-Engine: Empowering Optimization Reasoning in Large Language Models with Verifiable Synthetic NP Problems
by: Li, Xiaozhe, et al.
Published: (2025) -
OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces
by: Li, Xiaozhe, et al.
Published: (2026) -
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems
by: Li, Xiaozhe, et al.
Published: (2025) -
Beyond Mode Collapse: Distribution Matching for Diverse Reasoning
by: Li, Xiaozhe, et al.
Published: (2026) -
Timely Machine: Awareness of Time Makes Test-Time Scaling Agentic
by: Ma, Yichuan, et al.
Published: (2026)