Towards Scalable Oversight via Partitioned Human Supervision
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yin, Ren, Ishida, Takashi, Sugiyama, Masashi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Off-Policy Corrected Reward Modeling for Reinforcement Learning from Human Feedback
von: Ackermann, Johannes, et al.
Veröffentlicht: (2025)
von: Ackermann, Johannes, et al.
Veröffentlicht: (2025)
Gradient Regularization Prevents Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards
von: Ackermann, Johannes, et al.
Veröffentlicht: (2026)
von: Ackermann, Johannes, et al.
Veröffentlicht: (2026)
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
von: Ono, Shinnosuke, et al.
Veröffentlicht: (2026)
von: Ono, Shinnosuke, et al.
Veröffentlicht: (2026)
Scalable Oversight for Superhuman AI via Recursive Self-Critiquing
von: Wen, Xueru, et al.
Veröffentlicht: (2025)
von: Wen, Xueru, et al.
Veröffentlicht: (2025)
Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision
von: Sun, Zhiqing, et al.
Veröffentlicht: (2024)
von: Sun, Zhiqing, et al.
Veröffentlicht: (2024)
Generating Chain-of-Thoughts with a Pairwise-Comparison Approach to Searching for the Most Promising Intermediate Thought
von: Zhang, Zhen-Yu, et al.
Veröffentlicht: (2024)
von: Zhang, Zhen-Yu, et al.
Veröffentlicht: (2024)
Great Models Think Alike and this Undermines AI Oversight
von: Goel, Shashwat, et al.
Veröffentlicht: (2025)
von: Goel, Shashwat, et al.
Veröffentlicht: (2025)
Impact of Noisy Supervision in Foundation Model Learning
von: Chen, Hao, et al.
Veröffentlicht: (2024)
von: Chen, Hao, et al.
Veröffentlicht: (2024)
Scaling Laws For Scalable Oversight
von: Engels, Joshua, et al.
Veröffentlicht: (2025)
von: Engels, Joshua, et al.
Veröffentlicht: (2025)
Towards Scalable Automated Alignment of LLMs: A Survey
von: Cao, Boxi, et al.
Veröffentlicht: (2024)
von: Cao, Boxi, et al.
Veröffentlicht: (2024)
Building a Precise Video Language with Human-AI Oversight
von: Lin, Zhiqiu, et al.
Veröffentlicht: (2026)
von: Lin, Zhiqiu, et al.
Veröffentlicht: (2026)
Action-Agnostic Point-Level Supervision for Temporal Action Detection
von: Yoshida, Shuhei M., et al.
Veröffentlicht: (2024)
von: Yoshida, Shuhei M., et al.
Veröffentlicht: (2024)
Towards Scalable Meta-Learning of near-optimal Interpretable Models via Synthetic Model Generations
von: Myint, Kyaw Hpone, et al.
Veröffentlicht: (2025)
von: Myint, Kyaw Hpone, et al.
Veröffentlicht: (2025)
Aligning Human and Machine Attention for Enhanced Supervised Learning
von: Chriqui, Avihay, et al.
Veröffentlicht: (2025)
von: Chriqui, Avihay, et al.
Veröffentlicht: (2025)
Guided Self-Evolving LLMs with Minimal Human Supervision
von: Yu, Wenhao, et al.
Veröffentlicht: (2025)
von: Yu, Wenhao, et al.
Veröffentlicht: (2025)
Auto-ICL: In-Context Learning without Human Supervision
von: Yang, Jinghan, et al.
Veröffentlicht: (2023)
von: Yang, Jinghan, et al.
Veröffentlicht: (2023)
Vision-Language Model Fine-Tuning via Simple Parameter-Efficient Modification
von: Li, Ming, et al.
Veröffentlicht: (2024)
von: Li, Ming, et al.
Veröffentlicht: (2024)
Steering LLMs via Scalable Interactive Oversight
von: Zhou, Enyu, et al.
Veröffentlicht: (2026)
von: Zhou, Enyu, et al.
Veröffentlicht: (2026)
Modeling Human Beliefs about AI Behavior for Scalable Oversight
von: Lang, Leon, et al.
Veröffentlicht: (2025)
von: Lang, Leon, et al.
Veröffentlicht: (2025)
Extracting effective solutions hidden in large language models via generated comprehensive specialists: case studies in developing electronic devices
von: Tomita, Hikari, et al.
Veröffentlicht: (2024)
von: Tomita, Hikari, et al.
Veröffentlicht: (2024)
Scalable Chain of Thoughts via Elastic Reasoning
von: Xu, Yuhui, et al.
Veröffentlicht: (2025)
von: Xu, Yuhui, et al.
Veröffentlicht: (2025)
Towards Few-Shot Adaptation of Foundation Models via Multitask Finetuning
von: Xu, Zhuoyan, et al.
Veröffentlicht: (2024)
von: Xu, Zhuoyan, et al.
Veröffentlicht: (2024)
Self-Supervised Prompt Optimization
von: Xiang, Jinyu, et al.
Veröffentlicht: (2025)
von: Xiang, Jinyu, et al.
Veröffentlicht: (2025)
Rotate, Clip, and Partition: Towards W2A4KV4 Quantization by Integrating Rotation and Learnable Non-uniform Quantizer
von: Choi, Euntae, et al.
Veröffentlicht: (2025)
von: Choi, Euntae, et al.
Veröffentlicht: (2025)
Imitating Language via Scalable Inverse Reinforcement Learning
von: Wulfmeier, Markus, et al.
Veröffentlicht: (2024)
von: Wulfmeier, Markus, et al.
Veröffentlicht: (2024)
CoPE: Clipped RoPE as A Scalable Free Lunch for Long Context LLMs
von: Li, Haoran, et al.
Veröffentlicht: (2026)
von: Li, Haoran, et al.
Veröffentlicht: (2026)
Plan Optimization to Bilingual Dictionary Induction for Low-Resource Language Families
von: Nasution, Arbi Haza, et al.
Veröffentlicht: (2020)
von: Nasution, Arbi Haza, et al.
Veröffentlicht: (2020)
Towards Unified Alignment Between Agents, Humans, and Environment
von: Yang, Zonghan, et al.
Veröffentlicht: (2024)
von: Yang, Zonghan, et al.
Veröffentlicht: (2024)
Aligning Large Language Models via Fine-grained Supervision
von: Xu, Dehong, et al.
Veröffentlicht: (2024)
von: Xu, Dehong, et al.
Veröffentlicht: (2024)
Differentially Private Zeroth-Order Methods for Scalable Large Language Model Finetuning
von: Liu, Z, et al.
Veröffentlicht: (2024)
von: Liu, Z, et al.
Veröffentlicht: (2024)
Scalable Prompt Routing via Fine-Grained Latent Task Discovery
von: Zhang, Yunyi, et al.
Veröffentlicht: (2026)
von: Zhang, Yunyi, et al.
Veröffentlicht: (2026)
Muon is Scalable for LLM Training
von: Liu, Jingyuan, et al.
Veröffentlicht: (2025)
von: Liu, Jingyuan, et al.
Veröffentlicht: (2025)
Stabilizing LLM Supervised Fine-Tuning via Explicit Distributional Control
von: Wang, Xinyu, et al.
Veröffentlicht: (2026)
von: Wang, Xinyu, et al.
Veröffentlicht: (2026)
Divide-Then-Aggregate: An Efficient Tool Learning Method via Parallel Tool Invocation
von: Zhu, Dongsheng, et al.
Veröffentlicht: (2025)
von: Zhu, Dongsheng, et al.
Veröffentlicht: (2025)
Scalable Best-of-N Selection for Large Language Models via Self-Certainty
von: Kang, Zhewei, et al.
Veröffentlicht: (2025)
von: Kang, Zhewei, et al.
Veröffentlicht: (2025)
VI-CuRL: Stabilizing Verifier-Independent RL Reasoning via Confidence-Guided Variance Reduction
von: Cai, Xin-Qiang, et al.
Veröffentlicht: (2026)
von: Cai, Xin-Qiang, et al.
Veröffentlicht: (2026)
Q-SFT: Q-Learning for Language Models via Supervised Fine-Tuning
von: Hong, Joey, et al.
Veröffentlicht: (2024)
von: Hong, Joey, et al.
Veröffentlicht: (2024)
LLMs as Scalable, General-Purpose Simulators For Evolving Digital Agent Training
von: Wang, Yiming, et al.
Veröffentlicht: (2025)
von: Wang, Yiming, et al.
Veröffentlicht: (2025)
SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
von: Fu, Yuqian, et al.
Veröffentlicht: (2025)
von: Fu, Yuqian, et al.
Veröffentlicht: (2025)
Aligning LLMs with Human Uncertainty: A Beta-Bernoulli Calibrator for LLM Forecasting
von: Dai, Hui, et al.
Veröffentlicht: (2026)
von: Dai, Hui, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Off-Policy Corrected Reward Modeling for Reinforcement Learning from Human Feedback
von: Ackermann, Johannes, et al.
Veröffentlicht: (2025) -
Gradient Regularization Prevents Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards
von: Ackermann, Johannes, et al.
Veröffentlicht: (2026) -
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
von: Ono, Shinnosuke, et al.
Veröffentlicht: (2026) -
Scalable Oversight for Superhuman AI via Recursive Self-Critiquing
von: Wen, Xueru, et al.
Veröffentlicht: (2025) -
Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision
von: Sun, Zhiqing, et al.
Veröffentlicht: (2024)