SuS: Strategy-aware Surprise for Intrinsic Exploration

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kashirskiy, Mark, Makarov, Ilya
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908768820789248
author Kashirskiy, Mark
Makarov, Ilya
author_facet Kashirskiy, Mark
Makarov, Ilya
contents We propose Strategy-aware Surprise (SuS), a novel intrinsic motivation framework that uses pre-post prediction mismatch as a novelty signal for exploration in reinforcement learning. Unlike traditional curiosity-driven methods that rely solely on state prediction error, SuS introduces two complementary components: Strategy Stability (SS) and Strategy Surprise (SuS). SS measures consistency in behavioral strategy across temporal steps, while SuS captures unexpected outcomes relative to the agent's current strategy representation. Our combined reward formulation leverages both signals through learned weighting coefficients. We evaluate SuS on mathematical reasoning tasks using large language models, demonstrating significant improvements in both accuracy and solution diversity. Ablation studies confirm that removing either component results in at least 10% performance degradation, validating the synergistic nature of our approach. SuS achieves 17.4% improvement in Pass@1 and 26.4% improvement in Pass@5 compared to baseline methods, while maintaining higher strategy diversity throughout training.
format Preprint
id arxiv_https___arxiv_org_abs_2601_10349
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SuS: Strategy-aware Surprise for Intrinsic Exploration
Kashirskiy, Mark
Makarov, Ilya
Machine Learning
Artificial Intelligence
Computation and Language
Computer Science and Game Theory
68T05, 68T07
I.2.6; I.2.8
We propose Strategy-aware Surprise (SuS), a novel intrinsic motivation framework that uses pre-post prediction mismatch as a novelty signal for exploration in reinforcement learning. Unlike traditional curiosity-driven methods that rely solely on state prediction error, SuS introduces two complementary components: Strategy Stability (SS) and Strategy Surprise (SuS). SS measures consistency in behavioral strategy across temporal steps, while SuS captures unexpected outcomes relative to the agent's current strategy representation. Our combined reward formulation leverages both signals through learned weighting coefficients. We evaluate SuS on mathematical reasoning tasks using large language models, demonstrating significant improvements in both accuracy and solution diversity. Ablation studies confirm that removing either component results in at least 10% performance degradation, validating the synergistic nature of our approach. SuS achieves 17.4% improvement in Pass@1 and 26.4% improvement in Pass@5 compared to baseline methods, while maintaining higher strategy diversity throughout training.
title SuS: Strategy-aware Surprise for Intrinsic Exploration
topic Machine Learning
Artificial Intelligence
Computation and Language
Computer Science and Game Theory
68T05, 68T07
I.2.6; I.2.8
url https://arxiv.org/abs/2601.10349