Precision over Diversity: High-Precision Reward Generalizes to Robust Instruction Following
Fuente:
arXiv
Saved in:
| Main Authors: | Zeng, Yirong, Liu, Yufei, Ding, Xiao, Hou, Yutai, Wang, Yuxian, Song, Haonan, Ning, Wu, Tu, Dandan, Zhang, Qixun, Cai, Bibo, He, Yuxiang, Liu, Ting |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Tool-Overuse Illusion: Why Does LLM Prefer External Tools over Internal Knowledge?
by: Zeng, Yirong, et al.
Published: (2026)
by: Zeng, Yirong, et al.
Published: (2026)
AutoTool: Automatic Scaling of Tool-Use Capabilities in RL via Decoupled Entropy Constraints
by: Zeng, Yirong, et al.
Published: (2026)
by: Zeng, Yirong, et al.
Published: (2026)
iTool: Reinforced Fine-Tuning with Dynamic Deficiency Calibration for Advanced Tool Use
by: Zeng, Yirong, et al.
Published: (2025)
by: Zeng, Yirong, et al.
Published: (2025)
Tool Zero: Training Tool-Augmented LLMs via Pure RL from Scratch
by: Zeng, Yirong, et al.
Published: (2025)
by: Zeng, Yirong, et al.
Published: (2025)
Concise and Precise Context Compression for Tool-Using Language Models
by: Xu, Yang, et al.
Published: (2024)
by: Xu, Yang, et al.
Published: (2024)
Towards Generalizable and Faithful Logic Reasoning over Natural Language via Resolution Refutation
by: Sun, Zhouhao, et al.
Published: (2024)
by: Sun, Zhouhao, et al.
Published: (2024)
Complex Instruction Following with Diverse Style Policies in Football Games
by: Sun, Chenglu, et al.
Published: (2025)
by: Sun, Chenglu, et al.
Published: (2025)
Self-Route: Automatic Mode Switching via Capability Estimation for Efficient Reasoning
by: He, Yang, et al.
Published: (2025)
by: He, Yang, et al.
Published: (2025)
DeepTool: Scaling Interleaved Deliberation in Tool-Integrated Reasoning via Process-Supervised Reinforcement Learning
by: He, Yang, et al.
Published: (2026)
by: He, Yang, et al.
Published: (2026)
ExpeTrans: LLMs Are Experiential Transfer Learners
by: Gao, Jinglong, et al.
Published: (2025)
by: Gao, Jinglong, et al.
Published: (2025)
MAESTRO: Meta-learning Adaptive Estimation of Scalarization Trade-offs for Reward Optimization
by: Zhao, Yang, et al.
Published: (2026)
by: Zhao, Yang, et al.
Published: (2026)
MangaNinja: Line Art Colorization with Precise Reference Following
by: Liu, Zhiheng, et al.
Published: (2025)
by: Liu, Zhiheng, et al.
Published: (2025)
PACIFIC: a framework for generating benchmarks to check Precise Automatically Checked Instruction Following In Code
by: Dreyfuss, Itay, et al.
Published: (2025)
by: Dreyfuss, Itay, et al.
Published: (2025)
Precision Imaging for Intraindividual Investigation of the Reward Response
by: Matthew Mattoni, et al.
Published: (2026)
by: Matthew Mattoni, et al.
Published: (2026)
Straggler-Resilient Federated Learning over A Hybrid Conventional and Pinching Antenna Network
by: Wu, Bibo, et al.
Published: (2025)
by: Wu, Bibo, et al.
Published: (2025)
Enhancing Causal Reasoning in Large Language Models: A Causal Attribution Model for Precision Fine-Tuning
by: Cai, Hengrui, et al.
Published: (2023)
by: Cai, Hengrui, et al.
Published: (2023)
FedRecon: Missing Modality Reconstruction in Heterogeneous Distributed Environments
by: Liu, Junming, et al.
Published: (2025)
by: Liu, Junming, et al.
Published: (2025)
Reinforcement Learning with Robust Rubric Rewards
by: Yu, Ya-Qi, et al.
Published: (2026)
by: Yu, Ya-Qi, et al.
Published: (2026)
Precision Bioprinted Living Facial Mask for Deep Burn Wound Repair
by: Xuanqi Liu, et al.
Published: (2025)
by: Xuanqi Liu, et al.
Published: (2025)
GR-Ben: A General Reasoning Benchmark for Evaluating Process Reward Models
by: Sun, Zhouhao, et al.
Published: (2026)
by: Sun, Zhouhao, et al.
Published: (2026)
Sequential Regiodivergent Polyol Sulfonylation and Functionalization Enable Precise Engineering of Carbohydrates.
by: Zhou, Siai, et al.
Published: (2026)
by: Zhou, Siai, et al.
Published: (2026)
Safety Geometry Collapse in Multimodal LLMs and Adaptive Drift Correction
by: Guo, Jiahe, et al.
Published: (2026)
by: Guo, Jiahe, et al.
Published: (2026)
Robust High-Precision Time Transfer over 91-km Hollow-Core Fiber: Immunity to Dispersion and Nonlinearity
by: Liu, Bo, et al.
Published: (2026)
by: Liu, Bo, et al.
Published: (2026)
FollowTable: A Benchmark for Instruction-Following Table Retrieval
by: Jin, Rihui, et al.
Published: (2026)
by: Jin, Rihui, et al.
Published: (2026)
UltraIF: Advancing Instruction Following from the Wild
by: An, Kaikai, et al.
Published: (2025)
by: An, Kaikai, et al.
Published: (2025)
Pairing Real-Time Piano Transcription with Symbol-level Tracking for Precise and Robust Score Following
by: Peter, Silvan, et al.
Published: (2025)
by: Peter, Silvan, et al.
Published: (2025)
Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
by: Su, Yi, et al.
Published: (2025)
by: Su, Yi, et al.
Published: (2025)
Advanced Synthesis, Structural Characterization, and Functional Applications of Precision Polymers
by: Jie Cen, et al.
Published: (2024)
by: Jie Cen, et al.
Published: (2024)
Advanced Microsphere and Hydrogel Platforms for Precision Interventional Therapy of Hepatocellular Carcinoma
by: Yanhui Wang, et al.
Published: (2025)
by: Yanhui Wang, et al.
Published: (2025)
Self‐Propelled Nanomotor for Cancer Precision Combination Therapy
by: Yijie Lu, et al.
Published: (2024)
by: Yijie Lu, et al.
Published: (2024)
Water Quality Assessment of Harvested Rainwater Across China's Loess Plateau: Implications for Precision Irrigation
by: Shaoxiong Ning, et al.
Published: (2026)
by: Shaoxiong Ning, et al.
Published: (2026)
Robust Training of Neural Networks at Arbitrary Precision and Sparsity
by: Ye, Chengxi, et al.
Published: (2024)
by: Ye, Chengxi, et al.
Published: (2024)
IFDECORATOR: Wrapping Instruction Following Reinforcement Learning with Verifiable Rewards
by: Guo, Xu, et al.
Published: (2025)
by: Guo, Xu, et al.
Published: (2025)
IF-RewardBench: Benchmarking Judge Models for Instruction-Following Evaluation
by: Wen, Bosi, et al.
Published: (2026)
by: Wen, Bosi, et al.
Published: (2026)
Large Language Models Are Still Misled by Simple Bias Ensembles
by: Sun, Zhouhao, et al.
Published: (2025)
by: Sun, Zhouhao, et al.
Published: (2025)
Few-Shot Precise Event Spotting via Unified Multi-Entity Graph and Distillation
by: Liu, Zhaoyu, et al.
Published: (2025)
by: Liu, Zhaoyu, et al.
Published: (2025)
QiMeng-PRepair: Precise Code Repair via Edit-Aware Reward Optimization
by: Ke, Changxin, et al.
Published: (2026)
by: Ke, Changxin, et al.
Published: (2026)
Precision Studies of the $η_c$ decay at BESIII
by: Zeng, Yijia
Published: (2026)
by: Zeng, Yijia
Published: (2026)
Engineering Precise and Robust Effective Hamiltonians
by: Chen, Jiahui, et al.
Published: (2025)
by: Chen, Jiahui, et al.
Published: (2025)
Process Reward Models Meet Planning: Generating Precise and Scalable Datasets for Step-Level Rewards
by: Pisano, Raffaele, et al.
Published: (2026)
by: Pisano, Raffaele, et al.
Published: (2026)
Similar Items
-
The Tool-Overuse Illusion: Why Does LLM Prefer External Tools over Internal Knowledge?
by: Zeng, Yirong, et al.
Published: (2026) -
AutoTool: Automatic Scaling of Tool-Use Capabilities in RL via Decoupled Entropy Constraints
by: Zeng, Yirong, et al.
Published: (2026) -
iTool: Reinforced Fine-Tuning with Dynamic Deficiency Calibration for Advanced Tool Use
by: Zeng, Yirong, et al.
Published: (2025) -
Tool Zero: Training Tool-Augmented LLMs via Pure RL from Scratch
by: Zeng, Yirong, et al.
Published: (2025) -
Concise and Precise Context Compression for Tool-Using Language Models
by: Xu, Yang, et al.
Published: (2024)