Saved in:
| Main Authors: | Yu, Zhaoning, Su, Will, Tao, Leitian, Wang, Haozhu, Singh, Aashu, Yu, Hanchao, Wang, Jianyu, Gao, Hongyang, Yuan, Weizhe, Weston, Jason, Yu, Ping, Xu, Jing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2510.02172 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Hybrid Reinforcement: When Reward Is Sparse, It's Better to Be Dense
by: Tao, Leitian, et al.
Published: (2025)
by: Tao, Leitian, et al.
Published: (2025)
MAGE: Model-Level Graph Neural Networks Explanations via Motif-based Graph Generation
by: Yu, Zhaoning, et al.
Published: (2024)
by: Yu, Zhaoning, et al.
Published: (2024)
CoT-Self-Instruct: Building high-quality synthetic prompts for reasoning and non-reasoning tasks
by: Yu, Ping, et al.
Published: (2025)
by: Yu, Ping, et al.
Published: (2025)
The Era of Real-World Human Interaction: RL from User Conversations
by: Jin, Chuanyang, et al.
Published: (2025)
by: Jin, Chuanyang, et al.
Published: (2025)
Self-Taught Evaluators
by: Wang, Tianlu, et al.
Published: (2024)
by: Wang, Tianlu, et al.
Published: (2024)
G2T-LLM: Graph-to-Tree Text Encoding for Molecule Generation with Fine-Tuned Large Language Models
by: Yu, Zhaoning, et al.
Published: (2024)
by: Yu, Zhaoning, et al.
Published: (2024)
Self-Consistency Preference Optimization
by: Prasad, Archiki, et al.
Published: (2024)
by: Prasad, Archiki, et al.
Published: (2024)
Self-Rewarding Language Models
by: Yuan, Weizhe, et al.
Published: (2024)
by: Yuan, Weizhe, et al.
Published: (2024)
Following Length Constraints in Instructions
by: Yuan, Weizhe, et al.
Published: (2024)
by: Yuan, Weizhe, et al.
Published: (2024)
R.I.P.: Better Models by Survival of the Fittest Prompts
by: Yu, Ping, et al.
Published: (2025)
by: Yu, Ping, et al.
Published: (2025)
HOW TO RESTRAIN SADDAM
Published: (1995)
Published: (1995)
Distilling System 2 into System 1
by: Yu, Ping, et al.
Published: (2024)
by: Yu, Ping, et al.
Published: (2024)
System-Level Natural Language Feedback
by: Yuan, Weizhe, et al.
Published: (2023)
by: Yuan, Weizhe, et al.
Published: (2023)
Optimizing Recall or Relevance? A Multi-Task Multi-Head Approach for Item-to-Item Retrieval in Recommendation
by: Zhang, Jiang, et al.
Published: (2025)
by: Zhang, Jiang, et al.
Published: (2025)
Your Weak LLM is Secretly a Strong Teacher for Alignment
by: Tao, Leitian, et al.
Published: (2024)
by: Tao, Leitian, et al.
Published: (2024)
Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
by: Wu, Tianhao, et al.
Published: (2024)
by: Wu, Tianhao, et al.
Published: (2024)
Self-Alignment with Instruction Backtranslation
by: Li, Xian, et al.
Published: (2023)
by: Li, Xian, et al.
Published: (2023)
Verifiable Reasoning for LLM-based Generative Recommendation
by: Lin, Xinyu, et al.
Published: (2026)
by: Lin, Xinyu, et al.
Published: (2026)
Limited Preference Data? Learning Better Reward Model with Latent Space Synthesis
by: Tao, Leitian, et al.
Published: (2025)
by: Tao, Leitian, et al.
Published: (2025)
SPICE: Self-Play In Corpus Environments Improves Reasoning
by: Liu, Bo, et al.
Published: (2025)
by: Liu, Bo, et al.
Published: (2025)
Mitigating Spurious Correlations for Self-supervised Recommendation
by: Lin, Xinyu, et al.
Published: (2022)
by: Lin, Xinyu, et al.
Published: (2022)
RESTRAIN: Reinforcement Learning-Based Secure Framework for Trigger-Action IoT Environment
by: Alam, Md Morshed, et al.
Published: (2025)
by: Alam, Md Morshed, et al.
Published: (2025)
CompCap: Improving Multimodal Large Language Models with Composite Captions
by: Chen, Xiaohui, et al.
Published: (2024)
by: Chen, Xiaohui, et al.
Published: (2024)
TOOLVERIFIER: Generalization to New Tools via Self-Verification
by: Mekala, Dheeraj, et al.
Published: (2024)
by: Mekala, Dheeraj, et al.
Published: (2024)
AnyTool: Self-Reflective, Hierarchical Agents for Large-Scale API Calls
by: Du, Yu, et al.
Published: (2024)
by: Du, Yu, et al.
Published: (2024)
Self-Improving Pretraining: using post-trained models to pretrain better models
by: Tan, Ellen Xiaoqing, et al.
Published: (2026)
by: Tan, Ellen Xiaoqing, et al.
Published: (2026)
Bridging Offline and Online Reinforcement Learning for LLMs
by: Lanchantin, Jack, et al.
Published: (2025)
by: Lanchantin, Jack, et al.
Published: (2025)
Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning
by: Yu, Yongcan, et al.
Published: (2026)
by: Yu, Yongcan, et al.
Published: (2026)
Self-Challenging Language Model Agents
by: Zhou, Yifei, et al.
Published: (2025)
by: Zhou, Yifei, et al.
Published: (2025)
Hierarchical Prompt Decision Transformer: Improving Few-Shot Policy Generalization with Global and Adaptive Guidance
by: Wang, Zhe, et al.
Published: (2024)
by: Wang, Zhe, et al.
Published: (2024)
MotionHint: Self-Supervised Monocular Visual Odometry with Motion Constraints
by: Wang, Cong, et al.
Published: (2021)
by: Wang, Cong, et al.
Published: (2021)
Learning Critically: Selective Self Distillation in Federated Learning on Non-IID Data
by: He, Yuting, et al.
Published: (2025)
by: He, Yuting, et al.
Published: (2025)
MAD-Spear: A Conformity-Driven Prompt Injection Attack on Multi-Agent Debate Systems
by: Cui, Yu, et al.
Published: (2025)
by: Cui, Yu, et al.
Published: (2025)
CodeLutra: Boosting LLM Code Generation via Preference-Guided Refinement
by: Tao, Leitian, et al.
Published: (2024)
by: Tao, Leitian, et al.
Published: (2024)
Save the Good Prefix: Precise Error Penalization via Process-Supervised RL to Enhance LLM Reasoning
by: Liu, Haolin, et al.
Published: (2026)
by: Liu, Haolin, et al.
Published: (2026)
LLM-Based SQL Generation: Prompting, Self-Refinement, and Adaptive Weighted Majority Voting
by: Yang, Yu-Jie, et al.
Published: (2026)
by: Yang, Yu-Jie, et al.
Published: (2026)
Quantitative estimates of the singular values of random i.i.d. matrices
by: Dai, Guozheng, et al.
Published: (2024)
by: Dai, Guozheng, et al.
Published: (2024)
Quantitative estimates of the spectral norm of random matrices with independent columns
by: Dai, Guozheng, et al.
Published: (2023)
by: Dai, Guozheng, et al.
Published: (2023)
J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning
by: Whitehouse, Chenxi, et al.
Published: (2025)
by: Whitehouse, Chenxi, et al.
Published: (2025)
Comments on “On the significance of peak dose in normal tissue toxicity in spatially fractionated radiotherapy: The case of proton minibeam radiation therapy”
by: Zhaoning Wang, et al.
Published: (2025)
by: Zhaoning Wang, et al.
Published: (2025)
Similar Items
-
Hybrid Reinforcement: When Reward Is Sparse, It's Better to Be Dense
by: Tao, Leitian, et al.
Published: (2025) -
MAGE: Model-Level Graph Neural Networks Explanations via Motif-based Graph Generation
by: Yu, Zhaoning, et al.
Published: (2024) -
CoT-Self-Instruct: Building high-quality synthetic prompts for reasoning and non-reasoning tasks
by: Yu, Ping, et al.
Published: (2025) -
The Era of Real-World Human Interaction: RL from User Conversations
by: Jin, Chuanyang, et al.
Published: (2025) -
Self-Taught Evaluators
by: Wang, Tianlu, et al.
Published: (2024)