Steering LLMs via Scalable Interactive Oversight
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Enyu, Xi, Zhiheng, Ma, Long, Zhang, Zhihao, Dou, Shihan, Lei, Zhikai, Wang, Guoteng, Zheng, Rui, Yan, Hang, Gui, Tao, Zhang, Qi, Huang, Xuanjing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BAPO: Stabilizing Off-Policy Reinforcement Learning for LLMs via Balanced Policy Optimization with Adaptive Clipping
von: Xi, Zhiheng, et al.
Veröffentlicht: (2025)
von: Xi, Zhiheng, et al.
Veröffentlicht: (2025)
Subspace Defense: Discarding Adversarial Perturbations by Learning a Subspace for Clean Signals
von: Zheng, Rui, et al.
Veröffentlicht: (2024)
von: Zheng, Rui, et al.
Veröffentlicht: (2024)
Improving RL Exploration for LLM Reasoning through Retrospective Replay
von: Dou, Shihan, et al.
Veröffentlicht: (2025)
von: Dou, Shihan, et al.
Veröffentlicht: (2025)
RMB: Comprehensively Benchmarking Reward Models in LLM Alignment
von: Zhou, Enyu, et al.
Veröffentlicht: (2024)
von: Zhou, Enyu, et al.
Veröffentlicht: (2024)
JFTA-Bench: Evaluate LLM's Ability of Tracking and Analyzing Malfunctions Using Fault Trees
von: Wang, Yuhui, et al.
Veröffentlicht: (2026)
von: Wang, Yuhui, et al.
Veröffentlicht: (2026)
MM-Doc-R1: Training Agents for Long Document Visual Question Answering through Multi-turn Reinforcement Learning
von: Lin, Jiahang, et al.
Veröffentlicht: (2026)
von: Lin, Jiahang, et al.
Veröffentlicht: (2026)
MetaRM: Shifted Distributions Alignment via Meta-Learning
von: Dou, Shihan, et al.
Veröffentlicht: (2024)
von: Dou, Shihan, et al.
Veröffentlicht: (2024)
Enhancing LLM-based Search Agents via Contribution Weighted Group Relative Policy Optimization
von: Wang, Junzhe, et al.
Veröffentlicht: (2026)
von: Wang, Junzhe, et al.
Veröffentlicht: (2026)
LoRAMoE: Alleviate World Knowledge Forgetting in Large Language Models via MoE-Style Plugin
von: Dou, Shihan, et al.
Veröffentlicht: (2023)
von: Dou, Shihan, et al.
Veröffentlicht: (2023)
StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback
von: Dou, Shihan, et al.
Veröffentlicht: (2024)
von: Dou, Shihan, et al.
Veröffentlicht: (2024)
Self-Polish: Enhance Reasoning in Large Language Models via Problem Refinement
von: Xi, Zhiheng, et al.
Veröffentlicht: (2023)
von: Xi, Zhiheng, et al.
Veröffentlicht: (2023)
Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses
von: Lin, Jiahang, et al.
Veröffentlicht: (2026)
von: Lin, Jiahang, et al.
Veröffentlicht: (2026)
Advancing Translation Preference Modeling with RLHF: A Step Towards Cost-Effective Solution
von: Xu, Nuo, et al.
Veröffentlicht: (2024)
von: Xu, Nuo, et al.
Veröffentlicht: (2024)
Unveiling Linguistic Regions in Large Language Models
von: Zhang, Zhihao, et al.
Veröffentlicht: (2024)
von: Zhang, Zhihao, et al.
Veröffentlicht: (2024)
RoCoIns: Enhancing Robustness of Large Language Models through Code-Style Instructions
von: Zhang, Yuansen, et al.
Veröffentlicht: (2024)
von: Zhang, Yuansen, et al.
Veröffentlicht: (2024)
Distill Visual Chart Reasoning Ability from LLMs to MLLMs
von: He, Wei, et al.
Veröffentlicht: (2024)
von: He, Wei, et al.
Veröffentlicht: (2024)
SafeAligner: Safety Alignment against Jailbreak Attacks via Response Disparity Guidance
von: Huang, Caishuang, et al.
Veröffentlicht: (2024)
von: Huang, Caishuang, et al.
Veröffentlicht: (2024)
LLMEval-Med: A Real-world Clinical Benchmark for Medical LLMs with Physician Validation
von: Zhang, Ming, et al.
Veröffentlicht: (2025)
von: Zhang, Ming, et al.
Veröffentlicht: (2025)
From Scores to Preferences: Redefining MOS Benchmarking for Speech Quality Reward Modeling
von: Cao, Yifei, et al.
Veröffentlicht: (2025)
von: Cao, Yifei, et al.
Veröffentlicht: (2025)
LLMEval-Logic: A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardening
von: Zhang, Ming, et al.
Veröffentlicht: (2026)
von: Zhang, Ming, et al.
Veröffentlicht: (2026)
Aligning Large Language Models from Self-Reference AI Feedback with one General Principle
von: Bao, Rong, et al.
Veröffentlicht: (2024)
von: Bao, Rong, et al.
Veröffentlicht: (2024)
CodeChameleon: Personalized Encryption Framework for Jailbreaking Large Language Models
von: Lv, Huijie, et al.
Veröffentlicht: (2024)
von: Lv, Huijie, et al.
Veröffentlicht: (2024)
Why Reinforcement Fine-Tuning Enables MLLMs Preserve Prior Knowledge Better: A Data Perspective
von: Zhang, Zhihao, et al.
Veröffentlicht: (2025)
von: Zhang, Zhihao, et al.
Veröffentlicht: (2025)
What's Wrong with Your Code Generated by Large Language Models? An Extensive Study
von: Dou, Shihan, et al.
Veröffentlicht: (2024)
von: Dou, Shihan, et al.
Veröffentlicht: (2024)
VRPO: Rethinking Value Modeling for Robust RL Training under Noisy Supervision
von: Zhu, Dingwei, et al.
Veröffentlicht: (2025)
von: Zhu, Dingwei, et al.
Veröffentlicht: (2025)
LLaMA Beyond English: An Empirical Study on Language Capability Transfer
von: Zhao, Jun, et al.
Veröffentlicht: (2024)
von: Zhao, Jun, et al.
Veröffentlicht: (2024)
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control
von: Zhang, Jiazheng, et al.
Veröffentlicht: (2026)
von: Zhang, Jiazheng, et al.
Veröffentlicht: (2026)
PFDial: A Structured Dialogue Instruction Fine-tuning Method Based on UML Flowcharts
von: Zhang, Ming, et al.
Veröffentlicht: (2025)
von: Zhang, Ming, et al.
Veröffentlicht: (2025)
Compression Hacking: A Supplementary Perspective on Informatics Properties of Language Models from Geometric Distortion
von: Zang, Jianxiang, et al.
Veröffentlicht: (2025)
von: Zang, Jianxiang, et al.
Veröffentlicht: (2025)
Toward Optimal LLM Alignments Using Two-Player Games
von: Zheng, Rui, et al.
Veröffentlicht: (2024)
von: Zheng, Rui, et al.
Veröffentlicht: (2024)
FRoM-W1: Towards General Humanoid Whole-Body Control with Language Instructions
von: Li, Peng, et al.
Veröffentlicht: (2026)
von: Li, Peng, et al.
Veröffentlicht: (2026)
Beyond Scaling: Measuring and Predicting the Upper Bound of Knowledge Retention in Language Model Pre-Training
von: Jiang, Changhao, et al.
Veröffentlicht: (2025)
von: Jiang, Changhao, et al.
Veröffentlicht: (2025)
EliteKV: Scalable KV Cache Compression via RoPE Frequency Selection and Joint Low-Rank Projection
von: Zhou, Yuhao, et al.
Veröffentlicht: (2025)
von: Zhou, Yuhao, et al.
Veröffentlicht: (2025)
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment
von: Zhang, Jiazheng, et al.
Veröffentlicht: (2025)
von: Zhang, Jiazheng, et al.
Veröffentlicht: (2025)
Multi-Programming Language Sandbox for LLMs
von: Dou, Shihan, et al.
Veröffentlicht: (2024)
von: Dou, Shihan, et al.
Veröffentlicht: (2024)
Better Process Supervision with Bi-directional Rewarding Signals
von: Chen, Wenxiang, et al.
Veröffentlicht: (2025)
von: Chen, Wenxiang, et al.
Veröffentlicht: (2025)
Counteracting Matthew Effect in Self-Improvement of LVLMs through Head-Tail Re-balancing
von: Guo, Xin, et al.
Veröffentlicht: (2025)
von: Guo, Xin, et al.
Veröffentlicht: (2025)
Self-Demos: Eliciting Out-of-Demonstration Generalizability in Large Language Models
von: He, Wei, et al.
Veröffentlicht: (2024)
von: He, Wei, et al.
Veröffentlicht: (2024)
Unlocking the Essence of Beauty: Advanced Aesthetic Reasoning with Relative-Absolute Policy Optimization
von: Liu, Boyang, et al.
Veröffentlicht: (2025)
von: Liu, Boyang, et al.
Veröffentlicht: (2025)
EvaLearn: Quantifying the Learning Capability and Efficiency of LLMs via Sequential Problem Solving
von: Dou, Shihan, et al.
Veröffentlicht: (2025)
von: Dou, Shihan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
BAPO: Stabilizing Off-Policy Reinforcement Learning for LLMs via Balanced Policy Optimization with Adaptive Clipping
von: Xi, Zhiheng, et al.
Veröffentlicht: (2025) -
Subspace Defense: Discarding Adversarial Perturbations by Learning a Subspace for Clean Signals
von: Zheng, Rui, et al.
Veröffentlicht: (2024) -
Improving RL Exploration for LLM Reasoning through Retrospective Replay
von: Dou, Shihan, et al.
Veröffentlicht: (2025) -
RMB: Comprehensively Benchmarking Reward Models in LLM Alignment
von: Zhou, Enyu, et al.
Veröffentlicht: (2024) -
JFTA-Bench: Evaluate LLM's Ability of Tracking and Analyzing Malfunctions Using Fault Trees
von: Wang, Yuhui, et al.
Veröffentlicht: (2026)