Counting Reward Automata: Sample Efficient Reinforcement Learning Through the Exploitation of Reward Function Structure
Fuente:
arXiv
Guardado en:
| Autores principales: | Bester, Tristan, Rosman, Benjamin, James, Steven, Tasse, Geraud Nangue |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Beyond Sliding Windows: Learning to Manage Memory in Non-Markovian Environments
por: Tasse, Geraud Nangue, et al.
Publicado: (2025)
por: Tasse, Geraud Nangue, et al.
Publicado: (2025)
Skill Machines: Temporal Logic Skill Composition in Reinforcement Learning
por: Tasse, Geraud Nangue, et al.
Publicado: (2022)
por: Tasse, Geraud Nangue, et al.
Publicado: (2022)
Reward Machines for Deep RL in Noisy and Uncertain Environments
por: Li, Andrew C., et al.
Publicado: (2024)
por: Li, Andrew C., et al.
Publicado: (2024)
Argumentation and Machine Learning
por: Rago, Antonio, et al.
Publicado: (2024)
por: Rago, Antonio, et al.
Publicado: (2024)
Unsupervised Automata Learning via Discrete Optimization
por: Lutz, Simon, et al.
Publicado: (2023)
por: Lutz, Simon, et al.
Publicado: (2023)
Solving Combinatorial Counting Problems with Weighted First-Order Model Counting
por: Wang, Yuanhong, et al.
Publicado: (2026)
por: Wang, Yuanhong, et al.
Publicado: (2026)
IntSat: Integer Linear Programming by Conflict-Driven Constraint-Learning
por: Nieuwenhuis, Robert, et al.
Publicado: (2024)
por: Nieuwenhuis, Robert, et al.
Publicado: (2024)
MoralityGym: A Benchmark for Evaluating Hierarchical Moral Alignment in Sequential Decision-Making Agents
por: Rosen, Simon, et al.
Publicado: (2026)
por: Rosen, Simon, et al.
Publicado: (2026)
Unsupervised Hierarchical Skill Discovery
por: Harvey, Damion, et al.
Publicado: (2026)
por: Harvey, Damion, et al.
Publicado: (2026)
Reasoning and Planning with Dynamically Changing Norms
por: Olson, Taylor, et al.
Publicado: (2026)
por: Olson, Taylor, et al.
Publicado: (2026)
Inverse-Free Wilson Loops for Transformers: A Practical Diagnostic for Invariance and Order Sensitivity
por: Chang, Edward Y., et al.
Publicado: (2025)
por: Chang, Edward Y., et al.
Publicado: (2025)
Logic interpretations of ANN partition cells
por: Schmitt, Ingo
Publicado: (2024)
por: Schmitt, Ingo
Publicado: (2024)
$γ(3,4)$ `Attention' in Cognitive Agents: Ontology-Free Knowledge Representations With Promise Theoretic Semantics
por: Burgess, Mark
Publicado: (2025)
por: Burgess, Mark
Publicado: (2025)
Rewarding Creativity: A Human-Aligned Generative Reward Model for Reinforcement Learning in Storytelling
por: Li, Zhaoyan, et al.
Publicado: (2026)
por: Li, Zhaoyan, et al.
Publicado: (2026)
On the Conditions for Domain Stability for Machine Learning: a Mathematical Approach
por: Pedroza, Gabriel
Publicado: (2024)
por: Pedroza, Gabriel
Publicado: (2024)
Compositional Instruction Following with Language Models and Reinforcement Learning
por: Cohen, Vanya, et al.
Publicado: (2025)
por: Cohen, Vanya, et al.
Publicado: (2025)
FrankenBot: Brain-Morphic Modular Orchestration for Robotic Manipulation with Vision-Language Models
por: Wang, Shiyi, et al.
Publicado: (2025)
por: Wang, Shiyi, et al.
Publicado: (2025)
Learning Rules Explaining Interactive Theorem Proving Tactic Prediction
por: Zhang, Liao, et al.
Publicado: (2024)
por: Zhang, Liao, et al.
Publicado: (2024)
Structure Transfer: an Inference-Based Calculus for the Transformation of Representations
por: Raggi, Daniel, et al.
Publicado: (2025)
por: Raggi, Daniel, et al.
Publicado: (2025)
Goal-Driven Query Answering over First- and Second-Order Dependencies with Equality
por: Tsamoura, Efthymia, et al.
Publicado: (2024)
por: Tsamoura, Efthymia, et al.
Publicado: (2024)
The Opaque Law of Artificial Intelligence
por: Calderonio, Vincenzo
Publicado: (2023)
por: Calderonio, Vincenzo
Publicado: (2023)
Unique Characterisability and Learnability of Temporal Queries Mediated by an Ontology
por: Jung, Jean Christoph, et al.
Publicado: (2023)
por: Jung, Jean Christoph, et al.
Publicado: (2023)
Understanding Syllogistic Reasoning in LLMs from Formal and Natural Language Perspectives
por: Poddar, Aheli, et al.
Publicado: (2025)
por: Poddar, Aheli, et al.
Publicado: (2025)
Automated Theorem Provers Help Improve Large Language Model Reasoning
por: McGinness, Lachlan, et al.
Publicado: (2024)
por: McGinness, Lachlan, et al.
Publicado: (2024)
TPTP World Infrastructure for Non-classical Logics
por: Steen, Alexander, et al.
Publicado: (2025)
por: Steen, Alexander, et al.
Publicado: (2025)
On The Role of Intentionality in Knowledge Representation: Analyzing Scene Context for Cognitive Agents with a Tiny Language Model
por: Burgess, Mark
Publicado: (2025)
por: Burgess, Mark
Publicado: (2025)
Globally Interpretable Classifiers via Boolean Formulas with Dynamic Propositions
por: Jaakkola, Reijo, et al.
Publicado: (2024)
por: Jaakkola, Reijo, et al.
Publicado: (2024)
ProofBridge: Auto-Formalization of Natural Language Proofs in Lean via Joint Embeddings
por: Jana, Prithwish, et al.
Publicado: (2025)
por: Jana, Prithwish, et al.
Publicado: (2025)
Anticipating Oblivious Opponents in Stochastic Games
por: Kalat, Shadi Tasdighi, et al.
Publicado: (2024)
por: Kalat, Shadi Tasdighi, et al.
Publicado: (2024)
Considerations on Approaches and Metrics in Automated Theorem Generation/Finding in Geometry
por: Quaresma, Pedro, et al.
Publicado: (2024)
por: Quaresma, Pedro, et al.
Publicado: (2024)
Sutra: Tensor-Op RNNs as a Compilation Target for Vector Symbolic Architectures
por: Leonhart, Emma
Publicado: (2026)
por: Leonhart, Emma
Publicado: (2026)
AI Art is Theft: Labour, Extraction, and Exploitation, Or, On the Dangers of Stochastic Pollocks
por: Goetze, Trystan S.
Publicado: (2024)
por: Goetze, Trystan S.
Publicado: (2024)
Active Automata Learning with Advice
por: Fica, Michał, et al.
Publicado: (2025)
por: Fica, Michał, et al.
Publicado: (2025)
Uncertainty-Aware Reward Discounting for Mitigating Reward Hacking
por: Singha, Disha
Publicado: (2026)
por: Singha, Disha
Publicado: (2026)
Learning How to Cube
por: Erata, Ferhat, et al.
Publicado: (2026)
por: Erata, Ferhat, et al.
Publicado: (2026)
CountPath: Automating Fragment Counting in Digital Pathology
por: Vieira, Ana Beatriz, et al.
Publicado: (2025)
por: Vieira, Ana Beatriz, et al.
Publicado: (2025)
Locality, Consistency, and the Tractability Frontier
por: Simas, Tristan
Publicado: (2026)
por: Simas, Tristan
Publicado: (2026)
Reward Model Interpretability via Optimal and Pessimal Tokens
por: Christian, Brian, et al.
Publicado: (2025)
por: Christian, Brian, et al.
Publicado: (2025)
Complex Event Recognition with Symbolic Register Transducers: Extended Technical Report
por: Alevizos, Elias, et al.
Publicado: (2024)
por: Alevizos, Elias, et al.
Publicado: (2024)
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy
por: Lan, Guangchen, et al.
Publicado: (2026)
por: Lan, Guangchen, et al.
Publicado: (2026)
Ejemplares similares
-
Beyond Sliding Windows: Learning to Manage Memory in Non-Markovian Environments
por: Tasse, Geraud Nangue, et al.
Publicado: (2025) -
Skill Machines: Temporal Logic Skill Composition in Reinforcement Learning
por: Tasse, Geraud Nangue, et al.
Publicado: (2022) -
Reward Machines for Deep RL in Noisy and Uncertain Environments
por: Li, Andrew C., et al.
Publicado: (2024) -
Argumentation and Machine Learning
por: Rago, Antonio, et al.
Publicado: (2024) -
Unsupervised Automata Learning via Discrete Optimization
por: Lutz, Simon, et al.
Publicado: (2023)