Why Does RLAIF Work At All?
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Young, Robin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Offline RLAIF: Piloting VLM Feedback for RL via SFO
von: Beck, Jacob
Veröffentlicht: (2025)
von: Beck, Jacob
Veröffentlicht: (2025)
RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
von: Lee, Harrison, et al.
Veröffentlicht: (2023)
von: Lee, Harrison, et al.
Veröffentlicht: (2023)
Information-theoretic Distinctions Between Deception and Confusion
von: Young, Robin
Veröffentlicht: (2025)
von: Young, Robin
Veröffentlicht: (2025)
Does Deep Active Learning Work in the Wild?
von: Ren, Simiao, et al.
Veröffentlicht: (2023)
von: Ren, Simiao, et al.
Veröffentlicht: (2023)
Why Linear Recurrent Memory Works in Partially Observable Reinforcement Learning
von: Zhao, Yike, et al.
Veröffentlicht: (2026)
von: Zhao, Yike, et al.
Veröffentlicht: (2026)
Why Goal-Conditioned Reinforcement Learning Works: Relation to Dual Control
von: Lawrence, Nathan P., et al.
Veröffentlicht: (2025)
von: Lawrence, Nathan P., et al.
Veröffentlicht: (2025)
Why Do Unlearnable Examples Work: A Novel Perspective of Mutual Information
von: Zhu, Yifan, et al.
Veröffentlicht: (2026)
von: Zhu, Yifan, et al.
Veröffentlicht: (2026)
Why Do Some Inputs Break Low-Bit LLM Quantization?
von: Chang, Ting-Yun, et al.
Veröffentlicht: (2025)
von: Chang, Ting-Yun, et al.
Veröffentlicht: (2025)
Does Your Wildfire Prediction Model Actually Work, or Just Score Well?
von: Xu, Yangshuang, et al.
Veröffentlicht: (2026)
von: Xu, Yangshuang, et al.
Veröffentlicht: (2026)
Why Representation Engineering Works: A Theoretical and Empirical Study in Vision-Language Models
von: Tian, Bowei, et al.
Veröffentlicht: (2025)
von: Tian, Bowei, et al.
Veröffentlicht: (2025)
What Is the Alignment Tax?
von: Young, Robin
Veröffentlicht: (2026)
von: Young, Robin
Veröffentlicht: (2026)
Infinite Width Models That Work: Why Feature Learning Doesn't Matter as Much as You Think
von: Sernau, Luke
Veröffentlicht: (2024)
von: Sernau, Luke
Veröffentlicht: (2024)
Feature-Enhanced Machine Learning for All-Cause Mortality Prediction in Healthcare Data
von: Lee, HyeYoung, et al.
Veröffentlicht: (2025)
von: Lee, HyeYoung, et al.
Veröffentlicht: (2025)
Why Adam Works Better with $β_1 = β_2$: The Missing Gradient Scale Invariance Principle
von: Fernández-Hernández, Alberto, et al.
Veröffentlicht: (2026)
von: Fernández-Hernández, Alberto, et al.
Veröffentlicht: (2026)
Adaptation Odyssey in LLMs: Why Does Additional Pretraining Sometimes Fail to Improve?
von: Öncel, Fırat, et al.
Veröffentlicht: (2024)
von: Öncel, Fırat, et al.
Veröffentlicht: (2024)
One Size Does Not Fit All: Token-Wise Adaptive Compression for KV Cache
von: Lu, Liming, et al.
Veröffentlicht: (2026)
von: Lu, Liming, et al.
Veröffentlicht: (2026)
Why Does ChatGPT "Delve" So Much? Exploring the Sources of Lexical Overrepresentation in Large Language Models
von: Juzek, Tom S., et al.
Veröffentlicht: (2024)
von: Juzek, Tom S., et al.
Veröffentlicht: (2024)
Why Does Stochastic Gradient Descent Slow Down in Low-Precision Training?
von: Yun, Vincent-Daniel
Veröffentlicht: (2025)
von: Yun, Vincent-Daniel
Veröffentlicht: (2025)
One Size Does Not Fit All: A Distribution-Aware Sparsification for More Precise Model Merging
von: Luo, Yingfeng, et al.
Veröffentlicht: (2025)
von: Luo, Yingfeng, et al.
Veröffentlicht: (2025)
Does Graph Prompt Work? A Data Operation Perspective with Theoretical Analysis
von: Wang, Qunzhong, et al.
Veröffentlicht: (2024)
von: Wang, Qunzhong, et al.
Veröffentlicht: (2024)
Does This Gradient Spark Joy?
von: Osband, Ian
Veröffentlicht: (2026)
von: Osband, Ian
Veröffentlicht: (2026)
AllMatch: Exploiting All Unlabeled Data for Semi-Supervised Learning
von: Wu, Zhiyu, et al.
Veröffentlicht: (2024)
von: Wu, Zhiyu, et al.
Veröffentlicht: (2024)
Why Do Language Model Agents Whistleblow?
von: Agrawal, Kushal, et al.
Veröffentlicht: (2025)
von: Agrawal, Kushal, et al.
Veröffentlicht: (2025)
Magic Words or Methodical Work? Challenging Conventional Wisdom in LLM-Based Political Text Annotation
von: McLaren, Lorcan, et al.
Veröffentlicht: (2026)
von: McLaren, Lorcan, et al.
Veröffentlicht: (2026)
The Depth Delusion: Why Transformers Should Be Wider, Not Deeper
von: Fahim, Md Muhtasim Munif, et al.
Veröffentlicht: (2026)
von: Fahim, Md Muhtasim Munif, et al.
Veröffentlicht: (2026)
Why pre-training is beneficial for downstream classification tasks?
von: Jiang, Xin, et al.
Veröffentlicht: (2024)
von: Jiang, Xin, et al.
Veröffentlicht: (2024)
Why and How Auxiliary Tasks Improve JEPA Representations
von: Yu, Jiacan, et al.
Veröffentlicht: (2025)
von: Yu, Jiacan, et al.
Veröffentlicht: (2025)
Why Gradients Rapidly Increase Near the End of Training
von: Defazio, Aaron
Veröffentlicht: (2025)
von: Defazio, Aaron
Veröffentlicht: (2025)
Why Transformers Need Adam: A Hessian Perspective
von: Zhang, Yushun, et al.
Veröffentlicht: (2024)
von: Zhang, Yushun, et al.
Veröffentlicht: (2024)
DoWhy-GCM: An extension of DoWhy for causal inference in graphical causal models
von: Blöbaum, Patrick, et al.
Veröffentlicht: (2022)
von: Blöbaum, Patrick, et al.
Veröffentlicht: (2022)
Context is All You Need
von: Delanois, Jean Erik, et al.
Veröffentlicht: (2026)
von: Delanois, Jean Erik, et al.
Veröffentlicht: (2026)
WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
von: Drouin, Alexandre, et al.
Veröffentlicht: (2024)
von: Drouin, Alexandre, et al.
Veröffentlicht: (2024)
Why the Maximum Second Derivative of Activations Matters for Adversarial Robustness
von: Yu, Yunrui, et al.
Veröffentlicht: (2026)
von: Yu, Yunrui, et al.
Veröffentlicht: (2026)
Why Self-Inconsistency Arises in GNN Explanations and How to Exploit It
von: Tai, Wenxin, et al.
Veröffentlicht: (2026)
von: Tai, Wenxin, et al.
Veröffentlicht: (2026)
Why Inference in Large Models Becomes Decomposable After Training
von: Jin, Jidong
Veröffentlicht: (2026)
von: Jin, Jidong
Veröffentlicht: (2026)
Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why
von: Armandpour, Mohammadreza, et al.
Veröffentlicht: (2026)
von: Armandpour, Mohammadreza, et al.
Veröffentlicht: (2026)
Why Uncertainty Calibration Matters for Reliable Perturbation-based Explanations
von: Decker, Thomas, et al.
Veröffentlicht: (2025)
von: Decker, Thomas, et al.
Veröffentlicht: (2025)
Why Do Transformers Fail to Forecast Time Series In-Context?
von: Zhou, Yufa, et al.
Veröffentlicht: (2025)
von: Zhou, Yufa, et al.
Veröffentlicht: (2025)
Not All Adapters Matter: Selective Adapter Freezing for Memory-Efficient Fine-Tuning of Language Models
von: Son, Hyegang, et al.
Veröffentlicht: (2024)
von: Son, Hyegang, et al.
Veröffentlicht: (2024)
Learning Through Noise: Why Subliminal Learning Works and When It Fails
von: Brockers, Vincent C., et al.
Veröffentlicht: (2026)
von: Brockers, Vincent C., et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Offline RLAIF: Piloting VLM Feedback for RL via SFO
von: Beck, Jacob
Veröffentlicht: (2025) -
RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
von: Lee, Harrison, et al.
Veröffentlicht: (2023) -
Information-theoretic Distinctions Between Deception and Confusion
von: Young, Robin
Veröffentlicht: (2025) -
Does Deep Active Learning Work in the Wild?
von: Ren, Simiao, et al.
Veröffentlicht: (2023) -
Why Linear Recurrent Memory Works in Partially Observable Reinforcement Learning
von: Zhao, Yike, et al.
Veröffentlicht: (2026)