Counting Reward Automata: Sample Efficient Reinforcement Learning Through the Exploitation of Reward Function Structure
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bester, Tristan, Rosman, Benjamin, James, Steven, Tasse, Geraud Nangue |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Beyond Sliding Windows: Learning to Manage Memory in Non-Markovian Environments
von: Tasse, Geraud Nangue, et al.
Veröffentlicht: (2025)
von: Tasse, Geraud Nangue, et al.
Veröffentlicht: (2025)
Skill Machines: Temporal Logic Skill Composition in Reinforcement Learning
von: Tasse, Geraud Nangue, et al.
Veröffentlicht: (2022)
von: Tasse, Geraud Nangue, et al.
Veröffentlicht: (2022)
Reward Machines for Deep RL in Noisy and Uncertain Environments
von: Li, Andrew C., et al.
Veröffentlicht: (2024)
von: Li, Andrew C., et al.
Veröffentlicht: (2024)
Argumentation and Machine Learning
von: Rago, Antonio, et al.
Veröffentlicht: (2024)
von: Rago, Antonio, et al.
Veröffentlicht: (2024)
Unsupervised Automata Learning via Discrete Optimization
von: Lutz, Simon, et al.
Veröffentlicht: (2023)
von: Lutz, Simon, et al.
Veröffentlicht: (2023)
Solving Combinatorial Counting Problems with Weighted First-Order Model Counting
von: Wang, Yuanhong, et al.
Veröffentlicht: (2026)
von: Wang, Yuanhong, et al.
Veröffentlicht: (2026)
IntSat: Integer Linear Programming by Conflict-Driven Constraint-Learning
von: Nieuwenhuis, Robert, et al.
Veröffentlicht: (2024)
von: Nieuwenhuis, Robert, et al.
Veröffentlicht: (2024)
MoralityGym: A Benchmark for Evaluating Hierarchical Moral Alignment in Sequential Decision-Making Agents
von: Rosen, Simon, et al.
Veröffentlicht: (2026)
von: Rosen, Simon, et al.
Veröffentlicht: (2026)
Unsupervised Hierarchical Skill Discovery
von: Harvey, Damion, et al.
Veröffentlicht: (2026)
von: Harvey, Damion, et al.
Veröffentlicht: (2026)
Reasoning and Planning with Dynamically Changing Norms
von: Olson, Taylor, et al.
Veröffentlicht: (2026)
von: Olson, Taylor, et al.
Veröffentlicht: (2026)
Inverse-Free Wilson Loops for Transformers: A Practical Diagnostic for Invariance and Order Sensitivity
von: Chang, Edward Y., et al.
Veröffentlicht: (2025)
von: Chang, Edward Y., et al.
Veröffentlicht: (2025)
Logic interpretations of ANN partition cells
von: Schmitt, Ingo
Veröffentlicht: (2024)
von: Schmitt, Ingo
Veröffentlicht: (2024)
$γ(3,4)$ `Attention' in Cognitive Agents: Ontology-Free Knowledge Representations With Promise Theoretic Semantics
von: Burgess, Mark
Veröffentlicht: (2025)
von: Burgess, Mark
Veröffentlicht: (2025)
Rewarding Creativity: A Human-Aligned Generative Reward Model for Reinforcement Learning in Storytelling
von: Li, Zhaoyan, et al.
Veröffentlicht: (2026)
von: Li, Zhaoyan, et al.
Veröffentlicht: (2026)
On the Conditions for Domain Stability for Machine Learning: a Mathematical Approach
von: Pedroza, Gabriel
Veröffentlicht: (2024)
von: Pedroza, Gabriel
Veröffentlicht: (2024)
Compositional Instruction Following with Language Models and Reinforcement Learning
von: Cohen, Vanya, et al.
Veröffentlicht: (2025)
von: Cohen, Vanya, et al.
Veröffentlicht: (2025)
FrankenBot: Brain-Morphic Modular Orchestration for Robotic Manipulation with Vision-Language Models
von: Wang, Shiyi, et al.
Veröffentlicht: (2025)
von: Wang, Shiyi, et al.
Veröffentlicht: (2025)
Learning Rules Explaining Interactive Theorem Proving Tactic Prediction
von: Zhang, Liao, et al.
Veröffentlicht: (2024)
von: Zhang, Liao, et al.
Veröffentlicht: (2024)
Structure Transfer: an Inference-Based Calculus for the Transformation of Representations
von: Raggi, Daniel, et al.
Veröffentlicht: (2025)
von: Raggi, Daniel, et al.
Veröffentlicht: (2025)
Goal-Driven Query Answering over First- and Second-Order Dependencies with Equality
von: Tsamoura, Efthymia, et al.
Veröffentlicht: (2024)
von: Tsamoura, Efthymia, et al.
Veröffentlicht: (2024)
The Opaque Law of Artificial Intelligence
von: Calderonio, Vincenzo
Veröffentlicht: (2023)
von: Calderonio, Vincenzo
Veröffentlicht: (2023)
Unique Characterisability and Learnability of Temporal Queries Mediated by an Ontology
von: Jung, Jean Christoph, et al.
Veröffentlicht: (2023)
von: Jung, Jean Christoph, et al.
Veröffentlicht: (2023)
Understanding Syllogistic Reasoning in LLMs from Formal and Natural Language Perspectives
von: Poddar, Aheli, et al.
Veröffentlicht: (2025)
von: Poddar, Aheli, et al.
Veröffentlicht: (2025)
Automated Theorem Provers Help Improve Large Language Model Reasoning
von: McGinness, Lachlan, et al.
Veröffentlicht: (2024)
von: McGinness, Lachlan, et al.
Veröffentlicht: (2024)
TPTP World Infrastructure for Non-classical Logics
von: Steen, Alexander, et al.
Veröffentlicht: (2025)
von: Steen, Alexander, et al.
Veröffentlicht: (2025)
On The Role of Intentionality in Knowledge Representation: Analyzing Scene Context for Cognitive Agents with a Tiny Language Model
von: Burgess, Mark
Veröffentlicht: (2025)
von: Burgess, Mark
Veröffentlicht: (2025)
Globally Interpretable Classifiers via Boolean Formulas with Dynamic Propositions
von: Jaakkola, Reijo, et al.
Veröffentlicht: (2024)
von: Jaakkola, Reijo, et al.
Veröffentlicht: (2024)
ProofBridge: Auto-Formalization of Natural Language Proofs in Lean via Joint Embeddings
von: Jana, Prithwish, et al.
Veröffentlicht: (2025)
von: Jana, Prithwish, et al.
Veröffentlicht: (2025)
Anticipating Oblivious Opponents in Stochastic Games
von: Kalat, Shadi Tasdighi, et al.
Veröffentlicht: (2024)
von: Kalat, Shadi Tasdighi, et al.
Veröffentlicht: (2024)
Considerations on Approaches and Metrics in Automated Theorem Generation/Finding in Geometry
von: Quaresma, Pedro, et al.
Veröffentlicht: (2024)
von: Quaresma, Pedro, et al.
Veröffentlicht: (2024)
Sutra: Tensor-Op RNNs as a Compilation Target for Vector Symbolic Architectures
von: Leonhart, Emma
Veröffentlicht: (2026)
von: Leonhart, Emma
Veröffentlicht: (2026)
AI Art is Theft: Labour, Extraction, and Exploitation, Or, On the Dangers of Stochastic Pollocks
von: Goetze, Trystan S.
Veröffentlicht: (2024)
von: Goetze, Trystan S.
Veröffentlicht: (2024)
Active Automata Learning with Advice
von: Fica, Michał, et al.
Veröffentlicht: (2025)
von: Fica, Michał, et al.
Veröffentlicht: (2025)
Uncertainty-Aware Reward Discounting for Mitigating Reward Hacking
von: Singha, Disha
Veröffentlicht: (2026)
von: Singha, Disha
Veröffentlicht: (2026)
Learning How to Cube
von: Erata, Ferhat, et al.
Veröffentlicht: (2026)
von: Erata, Ferhat, et al.
Veröffentlicht: (2026)
CountPath: Automating Fragment Counting in Digital Pathology
von: Vieira, Ana Beatriz, et al.
Veröffentlicht: (2025)
von: Vieira, Ana Beatriz, et al.
Veröffentlicht: (2025)
Locality, Consistency, and the Tractability Frontier
von: Simas, Tristan
Veröffentlicht: (2026)
von: Simas, Tristan
Veröffentlicht: (2026)
Reward Model Interpretability via Optimal and Pessimal Tokens
von: Christian, Brian, et al.
Veröffentlicht: (2025)
von: Christian, Brian, et al.
Veröffentlicht: (2025)
Complex Event Recognition with Symbolic Register Transducers: Extended Technical Report
von: Alevizos, Elias, et al.
Veröffentlicht: (2024)
von: Alevizos, Elias, et al.
Veröffentlicht: (2024)
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy
von: Lan, Guangchen, et al.
Veröffentlicht: (2026)
von: Lan, Guangchen, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Beyond Sliding Windows: Learning to Manage Memory in Non-Markovian Environments
von: Tasse, Geraud Nangue, et al.
Veröffentlicht: (2025) -
Skill Machines: Temporal Logic Skill Composition in Reinforcement Learning
von: Tasse, Geraud Nangue, et al.
Veröffentlicht: (2022) -
Reward Machines for Deep RL in Noisy and Uncertain Environments
von: Li, Andrew C., et al.
Veröffentlicht: (2024) -
Argumentation and Machine Learning
von: Rago, Antonio, et al.
Veröffentlicht: (2024) -
Unsupervised Automata Learning via Discrete Optimization
von: Lutz, Simon, et al.
Veröffentlicht: (2023)