Proto Successor Measure: Representing the Behavior Space of an RL Agent
Fuente:
arXiv
Guardado en:
| Autores principales: | Agarwal, Siddhant, Sikchi, Harshit, Stone, Peter, Zhang, Amy |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Regularized Latent Dynamics Prediction is a Strong Baseline For Behavioral Foundation Models
por: Jajoo, Pranaya, et al.
Publicado: (2026)
por: Jajoo, Pranaya, et al.
Publicado: (2026)
Dual RL: Unification and New Methods for Reinforcement and Imitation Learning
por: Sikchi, Harshit, et al.
Publicado: (2023)
por: Sikchi, Harshit, et al.
Publicado: (2023)
RLZero: Direct Policy Inference from Language Without In-Domain Supervision
por: Sikchi, Harshit, et al.
Publicado: (2024)
por: Sikchi, Harshit, et al.
Publicado: (2024)
A Dual Approach to Imitation Learning from Observations with Offline Datasets
por: Sikchi, Harshit, et al.
Publicado: (2024)
por: Sikchi, Harshit, et al.
Publicado: (2024)
An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning
por: Xu, Haoran, et al.
Publicado: (2025)
por: Xu, Haoran, et al.
Publicado: (2025)
Contrastive Preference Learning: Learning from Human Feedback without RL
por: Hejna, Joey, et al.
Publicado: (2023)
por: Hejna, Joey, et al.
Publicado: (2023)
Fast Adaptation with Behavioral Foundation Models
por: Sikchi, Harshit, et al.
Publicado: (2025)
por: Sikchi, Harshit, et al.
Publicado: (2025)
SMORE: Score Models for Offline Goal-Conditioned Reinforcement Learning
por: Sikchi, Harshit, et al.
Publicado: (2023)
por: Sikchi, Harshit, et al.
Publicado: (2023)
SLAC: Simulation-Pretrained Latent Action Space for Whole-Body Real-World RL
por: Hu, Jiaheng, et al.
Publicado: (2025)
por: Hu, Jiaheng, et al.
Publicado: (2025)
A Distributional Analogue to the Successor Representation
por: Wiltzer, Harley, et al.
Publicado: (2024)
por: Wiltzer, Harley, et al.
Publicado: (2024)
Null Counterfactual Factor Interactions for Goal-Conditioned Reinforcement Learning
por: Chuck, Caleb, et al.
Publicado: (2025)
por: Chuck, Caleb, et al.
Publicado: (2025)
Quality-Diversity Actor-Critic: Learning High-Performing and Diverse Behaviors via Value and Successor Features Critics
por: Grillotti, Luca, et al.
Publicado: (2024)
por: Grillotti, Luca, et al.
Publicado: (2024)
Ensemble Successor Representations for Task Generalization in Offline-to-Online Reinforcement Learning
por: Wang, Changhong, et al.
Publicado: (2024)
por: Wang, Changhong, et al.
Publicado: (2024)
Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
por: Rafailov, Rafael, et al.
Publicado: (2024)
por: Rafailov, Rafael, et al.
Publicado: (2024)
Decoupling Exploration and Exploitation for Unsupervised Pre-training with Successor Features
por: Kim, JaeYoon, et al.
Publicado: (2024)
por: Kim, JaeYoon, et al.
Publicado: (2024)
Stacked Universal Successor Feature Approximators for Safety in Reinforcement Learning
por: Cannon, Ian, et al.
Publicado: (2024)
por: Cannon, Ian, et al.
Publicado: (2024)
VoiceAgentBench: Are Voice Assistants ready for agentic tasks?
por: Jain, Dhruv, et al.
Publicado: (2025)
por: Jain, Dhruv, et al.
Publicado: (2025)
Design Considerations in Offline Preference-based RL
por: Agarwal, Alekh, et al.
Publicado: (2025)
por: Agarwal, Alekh, et al.
Publicado: (2025)
CREStE: Scalable Mapless Navigation with Internet Scale Priors and Counterfactual Guidance
por: Zhang, Arthur, et al.
Publicado: (2025)
por: Zhang, Arthur, et al.
Publicado: (2025)
Non-Adversarial Inverse Reinforcement Learning via Successor Feature Matching
por: Jain, Arnav Kumar, et al.
Publicado: (2024)
por: Jain, Arnav Kumar, et al.
Publicado: (2024)
Understanding the Countably Infinite: Neural Network Models of the Successor Function and its Acquisition
por: Gupta, Vima, et al.
Publicado: (2023)
por: Gupta, Vima, et al.
Publicado: (2023)
Generalized Graph Transformer Variational Autoencoder
por: Karki, Siddhant
Publicado: (2025)
por: Karki, Siddhant
Publicado: (2025)
ArenaRL: Scaling RL for Open-Ended Agents via Tournament-based Relative Ranking
por: Zhang, Qiang, et al.
Publicado: (2026)
por: Zhang, Qiang, et al.
Publicado: (2026)
Hypformer: Exploring Efficient Transformer Fully in Hyperbolic Space
por: Yang, Menglin, et al.
Publicado: (2024)
por: Yang, Menglin, et al.
Publicado: (2024)
MoECollab: Democratizing LLM Development Through Collaborative Mixture of Experts
por: Harshit
Publicado: (2025)
por: Harshit
Publicado: (2025)
Explaining RL Decisions with Trajectories
por: Deshmukh, Shripad Vilasrao, et al.
Publicado: (2023)
por: Deshmukh, Shripad Vilasrao, et al.
Publicado: (2023)
Non-Linear Reinforcement Learning in Large Action Spaces: Structural Conditions and Sample-efficiency of Posterior Sampling
por: Agarwal, Alekh, et al.
Publicado: (2022)
por: Agarwal, Alekh, et al.
Publicado: (2022)
RL-ACRGNet: Reinforcement Learning-Based Chest Radiology Report Generation Network
por: Meena, Yogesh Kumar, et al.
Publicado: (2026)
por: Meena, Yogesh Kumar, et al.
Publicado: (2026)
Deep RL Needs Deep Behavior Analysis: Exploring Implicit Planning by Model-Free Agents in Open-Ended Environments
por: Simmons-Edler, Riley, et al.
Publicado: (2025)
por: Simmons-Edler, Riley, et al.
Publicado: (2025)
Language Models Represent Space and Time
por: Gurnee, Wes, et al.
Publicado: (2023)
por: Gurnee, Wes, et al.
Publicado: (2023)
Meta-RL Induces Exploration in Language Agents
por: Jiang, Yulun, et al.
Publicado: (2025)
por: Jiang, Yulun, et al.
Publicado: (2025)
Accuracy-Constrained CNN Pruning for Efficient and Reliable EEG-Based Seizure Detection
por: K, Mounvik, et al.
Publicado: (2025)
por: K, Mounvik, et al.
Publicado: (2025)
ProtoGCD: Unified and Unbiased Prototype Learning for Generalized Category Discovery
por: Ma, Shijie, et al.
Publicado: (2025)
por: Ma, Shijie, et al.
Publicado: (2025)
Learning Temporal Distances: Contrastive Successor Features Can Provide a Metric Structure for Decision-Making
por: Myers, Vivek, et al.
Publicado: (2024)
por: Myers, Vivek, et al.
Publicado: (2024)
ProtoNAM: Prototypical Neural Additive Models for Interpretable Deep Tabular Learning
por: Xiong, Guangzhi, et al.
Publicado: (2024)
por: Xiong, Guangzhi, et al.
Publicado: (2024)
Adapting the Behavior of Reinforcement Learning Agents to Changing Action Spaces and Reward Functions
por: de la Rosa, Raul, et al.
Publicado: (2026)
por: de la Rosa, Raul, et al.
Publicado: (2026)
MobileRL: Online Agentic Reinforcement Learning for Mobile GUI Agents
por: Xu, Yifan, et al.
Publicado: (2025)
por: Xu, Yifan, et al.
Publicado: (2025)
ProgAgent:A Continual RL Agent with Progress-Aware Rewards
por: Tan, Jinzhou, et al.
Publicado: (2026)
por: Tan, Jinzhou, et al.
Publicado: (2026)
Learning Successor Features with Distributed Hebbian Temporal Memory
por: Dzhivelikian, Evgenii, et al.
Publicado: (2023)
por: Dzhivelikian, Evgenii, et al.
Publicado: (2023)
Data-driven Discovery with Large Generative Models
por: Majumder, Bodhisattwa Prasad, et al.
Publicado: (2024)
por: Majumder, Bodhisattwa Prasad, et al.
Publicado: (2024)
Ejemplares similares
-
Regularized Latent Dynamics Prediction is a Strong Baseline For Behavioral Foundation Models
por: Jajoo, Pranaya, et al.
Publicado: (2026) -
Dual RL: Unification and New Methods for Reinforcement and Imitation Learning
por: Sikchi, Harshit, et al.
Publicado: (2023) -
RLZero: Direct Policy Inference from Language Without In-Domain Supervision
por: Sikchi, Harshit, et al.
Publicado: (2024) -
A Dual Approach to Imitation Learning from Observations with Offline Datasets
por: Sikchi, Harshit, et al.
Publicado: (2024) -
An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning
por: Xu, Haoran, et al.
Publicado: (2025)