Sinha, A., Elango, S., & Liu, D. (2026). Expected Return Causes Outcome-Level Mode Collapse in Reinforcement Learning and How to Fix It with Inverse Probability Scaling.
Chicago-Zitierstil (17. Ausg.)Sinha, Abhijeet, Sundari Elango, und Dianbo Liu. Expected Return Causes Outcome-Level Mode Collapse in Reinforcement Learning and How to Fix It with Inverse Probability Scaling. 2026.
MLA-Zitierstil (9. Ausg.)Sinha, Abhijeet, et al. Expected Return Causes Outcome-Level Mode Collapse in Reinforcement Learning and How to Fix It with Inverse Probability Scaling. 2026.
Achtung: Diese Zitate sind unter Umständen nicht zu 100% korrekt.