Zhang, Q., Wei, H., & Ying, L. (2024). Reinforcement Learning from Human Feedback without Reward Inference: Model-Free Algorithm and Instance-Dependent Analysis.
Chicago Style (17th ed.) CitationZhang, Qining, Honghao Wei, and Lei Ying. Reinforcement Learning from Human Feedback Without Reward Inference: Model-Free Algorithm and Instance-Dependent Analysis. 2024.
MLA (9th ed.) CitationZhang, Qining, et al. Reinforcement Learning from Human Feedback Without Reward Inference: Model-Free Algorithm and Instance-Dependent Analysis. 2024.
Warning: These citations may not always be 100% accurate.