Sahoo, S., & Junkin, J. (2025). The Horcrux: Mechanistically Interpretable Task Decomposition for Detecting and Mitigating Reward Hacking in Embodied AI Systems.
Chicago Style (17th ed.) CitationSahoo, Subramanyam, and Jared Junkin. The Horcrux: Mechanistically Interpretable Task Decomposition for Detecting and Mitigating Reward Hacking in Embodied AI Systems. 2025.
MLA (9th ed.) CitationSahoo, Subramanyam, and Jared Junkin. The Horcrux: Mechanistically Interpretable Task Decomposition for Detecting and Mitigating Reward Hacking in Embodied AI Systems. 2025.
Warning: These citations may not always be 100% accurate.