Tan, Y., Jiang, Y., Li, Y., Liu, J., Bu, X., Su, W., . . . Zheng, B. (2025). Equilibrate RLHF: Towards Balancing Helpfulness-Safety Trade-off in Large Language Models.
Chicago Style (17th ed.) CitationTan, Yingshui, Yilei Jiang, Yanshi Li, Jiaheng Liu, Xingyuan Bu, Wenbo Su, Xiangyu Yue, Xiaoyong Zhu, and Bo Zheng. Equilibrate RLHF: Towards Balancing Helpfulness-Safety Trade-off in Large Language Models. 2025.
MLA (9th ed.) CitationTan, Yingshui, et al. Equilibrate RLHF: Towards Balancing Helpfulness-Safety Trade-off in Large Language Models. 2025.
Warning: These citations may not always be 100% accurate.