Guardado en:
| Autores principales: | Corrado, Nicholas E., Hanna, Josiah P. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2508.01049 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
On-Policy Policy Gradient Reinforcement Learning Without On-Policy Sampling
por: Corrado, Nicholas E., et al.
Publicado: (2023)
por: Corrado, Nicholas E., et al.
Publicado: (2023)
Distributionally Robust Multi-Task Reinforcement Learning via Adaptive Task Sampling
por: Corrado, Nicholas E., et al.
Publicado: (2026)
por: Corrado, Nicholas E., et al.
Publicado: (2026)
Understanding when Dynamics-Invariant Data Augmentations Benefit Model-Free Reinforcement Learning Updates
por: Corrado, Nicholas E., et al.
Publicado: (2023)
por: Corrado, Nicholas E., et al.
Publicado: (2023)
Guided Data Augmentation for Offline Reinforcement Learning and Imitation Learning
por: Corrado, Nicholas E., et al.
Publicado: (2023)
por: Corrado, Nicholas E., et al.
Publicado: (2023)
Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation
por: Zhou, Hongyi, et al.
Publicado: (2025)
por: Zhou, Hongyi, et al.
Publicado: (2025)
SaVeR: Optimal Data Collection Strategy for Safe Policy Evaluation in Tabular MDP
por: Mukherjee, Subhojyoti, et al.
Publicado: (2024)
por: Mukherjee, Subhojyoti, et al.
Publicado: (2024)
Adaptive Exploration for Data-Efficient General Value Function Evaluations
por: Jain, Arushi, et al.
Publicado: (2024)
por: Jain, Arushi, et al.
Publicado: (2024)
SPEED: Experimental Design for Policy Evaluation in Linear Heteroscedastic Bandits
por: Mukherjee, Subhojyoti, et al.
Publicado: (2023)
por: Mukherjee, Subhojyoti, et al.
Publicado: (2023)
When Can Model-Free Reinforcement Learning be Enough for Thinking?
por: Hanna, Josiah P., et al.
Publicado: (2025)
por: Hanna, Josiah P., et al.
Publicado: (2025)
Pretraining Decision Transformers with Reward Prediction for In-Context Multi-task Structured Bandit Learning
por: Mukherjee, Subhojyoti, et al.
Publicado: (2024)
por: Mukherjee, Subhojyoti, et al.
Publicado: (2024)
SMAT: Staged Multi-Agent Training for Co-Adaptive Exoskeleton Control
por: Yuan, Yifei, et al.
Publicado: (2026)
por: Yuan, Yifei, et al.
Publicado: (2026)
Policy and World Modeling Co-Training for Language Agents
por: Lu, Ning, et al.
Publicado: (2026)
por: Lu, Ning, et al.
Publicado: (2026)
Stable Offline Value Function Learning with Bisimulation-based Representations
por: Pavse, Brahma S., et al.
Publicado: (2024)
por: Pavse, Brahma S., et al.
Publicado: (2024)
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training
por: Fakoor, Rasool, et al.
Publicado: (2026)
por: Fakoor, Rasool, et al.
Publicado: (2026)
Agent-Agnostic Centralized Training for Decentralized Multi-Agent Cooperative Driving
por: Yan, Shengchao, et al.
Publicado: (2024)
por: Yan, Shengchao, et al.
Publicado: (2024)
Learning to Stabilize Online Reinforcement Learning in Unbounded State Spaces
por: Pavse, Brahma S., et al.
Publicado: (2023)
por: Pavse, Brahma S., et al.
Publicado: (2023)
Adaptive Sample Sharing for Multi Agent Linear Bandits
por: Cherkaoui, Hamza, et al.
Publicado: (2023)
por: Cherkaoui, Hamza, et al.
Publicado: (2023)
AutoMixAlign: Adaptive Data Mixing for Multi-Task Preference Optimization in LLMs
por: Corrado, Nicholas E., et al.
Publicado: (2025)
por: Corrado, Nicholas E., et al.
Publicado: (2025)
An Empirical Study on the Power of Future Prediction in Partially Observable Environments
por: Kwon, Jeongyeol, et al.
Publicado: (2024)
por: Kwon, Jeongyeol, et al.
Publicado: (2024)
Adaptive Federated LoRA in Heterogeneous Wireless Networks with Independent Sampling
por: Hou, Yanzhao, et al.
Publicado: (2025)
por: Hou, Yanzhao, et al.
Publicado: (2025)
Co2PO: Coordinated Constrained Policy Optimization for Multi-Agent RL
por: Patel, Shrenik, et al.
Publicado: (2026)
por: Patel, Shrenik, et al.
Publicado: (2026)
MAT-Agent: Adaptive Multi-Agent Training Optimization
por: Zhang, Jusheng, et al.
Publicado: (2025)
por: Zhang, Jusheng, et al.
Publicado: (2025)
Centralized Permutation Equivariant Policy for Cooperative Multi-Agent Reinforcement Learning
por: Xu, Zhuofan, et al.
Publicado: (2025)
por: Xu, Zhuofan, et al.
Publicado: (2025)
Reliable Self-Harm Risk Screening via Adaptive Multi-Agent LLM Systems
por: Karnam, Meghana, et al.
Publicado: (2026)
por: Karnam, Meghana, et al.
Publicado: (2026)
Reinforcement Learning via Auxiliary Task Distillation
por: Harish, Abhinav Narayan, et al.
Publicado: (2024)
por: Harish, Abhinav Narayan, et al.
Publicado: (2024)
Adaptive Federated Learning in Heterogeneous Wireless Networks with Independent Sampling
por: Geng, Jiaxiang, et al.
Publicado: (2024)
por: Geng, Jiaxiang, et al.
Publicado: (2024)
An Introduction to Centralized Training for Decentralized Execution in Cooperative Multi-Agent Reinforcement Learning
por: Amato, Christopher
Publicado: (2024)
por: Amato, Christopher
Publicado: (2024)
Action-Graph Policies: Learning Action Co-dependencies in Multi-Agent Reinforcement Learning
por: Gupta, Nikunj, et al.
Publicado: (2026)
por: Gupta, Nikunj, et al.
Publicado: (2026)
Adaptive Gradient Normalization and Independent Sampling for (Stochastic) Generalized-Smooth Optimization
por: Yang, Yufeng, et al.
Publicado: (2024)
por: Yang, Yufeng, et al.
Publicado: (2024)
Sparsely Multimodal Data Fusion
por: Bjorgaard, Josiah
Publicado: (2024)
por: Bjorgaard, Josiah
Publicado: (2024)
Co-Optimizing Reconfigurable Environments and Policies for Decentralized Multi-Agent Navigation
por: Gao, Zhan, et al.
Publicado: (2024)
por: Gao, Zhan, et al.
Publicado: (2024)
Offline Multi-Agent Reinforcement Learning via In-Sample Sequential Policy Optimization
por: Liu, Zongkai, et al.
Publicado: (2024)
por: Liu, Zongkai, et al.
Publicado: (2024)
Rollout-Training Co-Design for Efficient LLM-Based Multi-Agent Reinforcement Learning
por: Jiang, Zhida, et al.
Publicado: (2026)
por: Jiang, Zhida, et al.
Publicado: (2026)
CoFi-PGMA: Counterfactual Policy Gradients under Filtered Feedback for Multi-Agent LLMs
por: Tong, Stela, et al.
Publicado: (2026)
por: Tong, Stela, et al.
Publicado: (2026)
Decentralized Transformers with Centralized Aggregation are Sample-Efficient Multi-Agent World Models
por: Zhang, Yang, et al.
Publicado: (2024)
por: Zhang, Yang, et al.
Publicado: (2024)
Reliability-Aware Adaptive Self-Consistency for Efficient Sampling in LLM Reasoning
por: Kim, Junseok, et al.
Publicado: (2026)
por: Kim, Junseok, et al.
Publicado: (2026)
Fully Independent Communication in Multi-Agent Reinforcement Learning
por: Pina, Rafael, et al.
Publicado: (2024)
por: Pina, Rafael, et al.
Publicado: (2024)
Emergent Coordination and Phase Structure in Independent Multi-Agent Reinforcement Learning
por: Yamaguchi, Azusa
Publicado: (2025)
por: Yamaguchi, Azusa
Publicado: (2025)
Balanced Training of Energy-Based Models with Adaptive Flow Sampling
por: Grenioux, Louis, et al.
Publicado: (2023)
por: Grenioux, Louis, et al.
Publicado: (2023)
Approximate Global Convergence of Independent Learning in Multi-Agent Systems
por: Jin, Ruiyang, et al.
Publicado: (2024)
por: Jin, Ruiyang, et al.
Publicado: (2024)
Ejemplares similares
-
On-Policy Policy Gradient Reinforcement Learning Without On-Policy Sampling
por: Corrado, Nicholas E., et al.
Publicado: (2023) -
Distributionally Robust Multi-Task Reinforcement Learning via Adaptive Task Sampling
por: Corrado, Nicholas E., et al.
Publicado: (2026) -
Understanding when Dynamics-Invariant Data Augmentations Benefit Model-Free Reinforcement Learning Updates
por: Corrado, Nicholas E., et al.
Publicado: (2023) -
Guided Data Augmentation for Offline Reinforcement Learning and Imitation Learning
por: Corrado, Nicholas E., et al.
Publicado: (2023) -
Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation
por: Zhou, Hongyi, et al.
Publicado: (2025)