Conservative Bias in Multi-Teacher Learning: Why Agents Prefer Low-Reward Advisors
Fuente:
arXiv
Guardado en:
| Autores principales: | Mesto, Maher, Cruz, Francisco |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Predictive Traffic Rule Compliance using Reinforcement Learning
por: Huang, Yanliang, et al.
Publicado: (2025)
por: Huang, Yanliang, et al.
Publicado: (2025)
Enhancing Heterogeneous Multi-Agent Cooperation in Decentralized MARL via GNN-driven Intrinsic Rewards
por: Monon, Jahir Sadik, et al.
Publicado: (2024)
por: Monon, Jahir Sadik, et al.
Publicado: (2024)
Energy-Efficient Quadruped Locomotion with Compliant Feet
por: Pal, Pramod, et al.
Publicado: (2026)
por: Pal, Pramod, et al.
Publicado: (2026)
Robo-CSK-Organizer: Commonsense Knowledge to Organize Detected Objects for Multipurpose Robots
por: Hidalgo, Rafael, et al.
Publicado: (2024)
por: Hidalgo, Rafael, et al.
Publicado: (2024)
Atomic-Probe Governance for Skill Updates in Compositional Robot Policies
por: Qin, Xue, et al.
Publicado: (2026)
por: Qin, Xue, et al.
Publicado: (2026)
Bimanual Robot Manipulation via Multi-Agent In-Context Learning
por: Palma, Alessio, et al.
Publicado: (2026)
por: Palma, Alessio, et al.
Publicado: (2026)
MTDrive: Multi-turn Interactive Reinforcement Learning for Autonomous Driving
por: Li, Xidong, et al.
Publicado: (2026)
por: Li, Xidong, et al.
Publicado: (2026)
Decentralized Aerial Manipulation of a Cable-Suspended Load using Multi-Agent Reinforcement Learning
por: Zeng, Jack, et al.
Publicado: (2025)
por: Zeng, Jack, et al.
Publicado: (2025)
Simulation-Based Counterfactual Causal Discovery on Real World Driver Behaviour
por: Howard, Rhys, et al.
Publicado: (2023)
por: Howard, Rhys, et al.
Publicado: (2023)
CaMeRL: Collision-Aware and Memory-Enhanced Reinforcement Learning for UAV Navigation in Multi-Scale Obstacle Environments
por: Hong, Hong, et al.
Publicado: (2026)
por: Hong, Hong, et al.
Publicado: (2026)
Cortex 2.0: Grounding World Models in Real-World Industrial Deployment
por: Aida, Adriana, et al.
Publicado: (2026)
por: Aida, Adriana, et al.
Publicado: (2026)
Sim-to-reality adaptation for Deep Reinforcement Learning applied to an underwater docking application
por: Chaarani, Alaaeddine, et al.
Publicado: (2026)
por: Chaarani, Alaaeddine, et al.
Publicado: (2026)
Magnet-Based Soft Robotic Skin Using a 3D-Printed Multi-Lattice Structure and CNN-Based Tactile Super-Resolution
por: Bang, Yunseong, et al.
Publicado: (2026)
por: Bang, Yunseong, et al.
Publicado: (2026)
DRAE: Dynamic Retrieval-Augmented Expert Networks for Lifelong Learning and Task Adaptation in Robotics
por: Long, Yayu, et al.
Publicado: (2025)
por: Long, Yayu, et al.
Publicado: (2025)
ScoRe-Flow: Complete Distributional Control via Score-Based Reinforcement Learning for Flow Matching
por: Qiu, Xiaotian, et al.
Publicado: (2026)
por: Qiu, Xiaotian, et al.
Publicado: (2026)
Towards a Robust Soft Baby Robot With Rich Interaction Ability for Advanced Machine Learning Algorithms
por: Alhakami, Mohannad, et al.
Publicado: (2024)
por: Alhakami, Mohannad, et al.
Publicado: (2024)
Generating Causal Explanations of Vehicular Agent Behavioural Interactions with Learnt Reward Profiles
por: Howard, Rhys, et al.
Publicado: (2025)
por: Howard, Rhys, et al.
Publicado: (2025)
Improving Agent Behaviors with RL Fine-tuning for Autonomous Driving
por: Peng, Zhenghao, et al.
Publicado: (2024)
por: Peng, Zhenghao, et al.
Publicado: (2024)
FCBV-Net: Category-Level Robotic Garment Smoothing via Feature-Conditioned Bimanual Value Prediction
por: Daba, Mohammed, et al.
Publicado: (2025)
por: Daba, Mohammed, et al.
Publicado: (2025)
Preventing Robotic Jailbreaking via Multimodal Domain Adaptation
por: Marchiori, Francesco, et al.
Publicado: (2025)
por: Marchiori, Francesco, et al.
Publicado: (2025)
Affordance-Aware Interactive Decision-Making and Execution for Ambiguous Instructions
por: Xu, Hengxuan, et al.
Publicado: (2026)
por: Xu, Hengxuan, et al.
Publicado: (2026)
SPACeR: Self-Play Anchoring with Centralized Reference Models
por: Chang, Wei-Jer, et al.
Publicado: (2025)
por: Chang, Wei-Jer, et al.
Publicado: (2025)
CUPID: Curating Data your Robot Loves with Influence Functions
por: Agia, Christopher, et al.
Publicado: (2025)
por: Agia, Christopher, et al.
Publicado: (2025)
SaiVLA-0: Cerebrum--Pons--Cerebellum Tripartite Architecture for Compute-Aware Vision-Language-Action
por: Shi, Xiang, et al.
Publicado: (2026)
por: Shi, Xiang, et al.
Publicado: (2026)
When Does Adaptive Guidance Help? Belief-Aware Privileged Distillation for Autonomous Driving Under Partial Observability
por: Haklidir, Mehmet
Publicado: (2026)
por: Haklidir, Mehmet
Publicado: (2026)
Evaluating Temporal Observation-Based Causal Discovery Techniques Applied to Road Driver Behaviour
por: Howard, Rhys, et al.
Publicado: (2023)
por: Howard, Rhys, et al.
Publicado: (2023)
MEReQ: Max-Ent Residual-Q Inverse RL for Sample-Efficient Alignment from Intervention
por: Chen, Yuxin, et al.
Publicado: (2024)
por: Chen, Yuxin, et al.
Publicado: (2024)
GammaZero: Learning To Guide POMDP Belief Space Search With Graph Representations
por: Mangannavar, Rajesh, et al.
Publicado: (2025)
por: Mangannavar, Rajesh, et al.
Publicado: (2025)
The Shortcomings of Force-from-Motion in Robot Learning
por: Aljalbout, Elie, et al.
Publicado: (2024)
por: Aljalbout, Elie, et al.
Publicado: (2024)
Building Minimal and Reusable Causal State Abstractions for Reinforcement Learning
por: Wang, Zizhao, et al.
Publicado: (2024)
por: Wang, Zizhao, et al.
Publicado: (2024)
RoboPack: Learning Tactile-Informed Dynamics Models for Dense Packing
por: Ai, Bo, et al.
Publicado: (2024)
por: Ai, Bo, et al.
Publicado: (2024)
Multi-Agent Pathfinding with Non-Unit Integer Edge Costs via Enhanced Conflict-Based Search and Graph Discretization
por: Fan, Hongkai, et al.
Publicado: (2026)
por: Fan, Hongkai, et al.
Publicado: (2026)
Tulip Agent -- Enabling LLM-Based Agents to Solve Tasks Using Large Tool Libraries
por: Ocker, Felix, et al.
Publicado: (2024)
por: Ocker, Felix, et al.
Publicado: (2024)
LiloDriver: A Lifelong Learning Framework for Closed-loop Motion Planning in Long-tail Autonomous Driving Scenarios
por: Yao, Huaiyuan, et al.
Publicado: (2025)
por: Yao, Huaiyuan, et al.
Publicado: (2025)
RoboGrind: Intuitive and Interactive Surface Treatment with Industrial Robots
por: Alt, Benjamin, et al.
Publicado: (2024)
por: Alt, Benjamin, et al.
Publicado: (2024)
Selective Progress-Aware Querying for Human-in-the-Loop Reinforcement Learning
por: Muraleedharan, Anujith, et al.
Publicado: (2025)
por: Muraleedharan, Anujith, et al.
Publicado: (2025)
Dream to Fly: Model-Based Reinforcement Learning for Vision-Based Drone Flight
por: Romero, Angel, et al.
Publicado: (2025)
por: Romero, Angel, et al.
Publicado: (2025)
The Reality Gap in Robotics: Challenges, Solutions, and Best Practices
por: Aljalbout, Elie, et al.
Publicado: (2025)
por: Aljalbout, Elie, et al.
Publicado: (2025)
Achieving Scalable Robot Autonomy via neurosymbolic planning using lightweight local LLM
por: Attolino, Nicholas, et al.
Publicado: (2025)
por: Attolino, Nicholas, et al.
Publicado: (2025)
A Framework for Neurosymbolic Robot Action Planning using Large Language Models
por: Capitanelli, Alessio, et al.
Publicado: (2023)
por: Capitanelli, Alessio, et al.
Publicado: (2023)
Ejemplares similares
-
Predictive Traffic Rule Compliance using Reinforcement Learning
por: Huang, Yanliang, et al.
Publicado: (2025) -
Enhancing Heterogeneous Multi-Agent Cooperation in Decentralized MARL via GNN-driven Intrinsic Rewards
por: Monon, Jahir Sadik, et al.
Publicado: (2024) -
Energy-Efficient Quadruped Locomotion with Compliant Feet
por: Pal, Pramod, et al.
Publicado: (2026) -
Robo-CSK-Organizer: Commonsense Knowledge to Organize Detected Objects for Multipurpose Robots
por: Hidalgo, Rafael, et al.
Publicado: (2024) -
Atomic-Probe Governance for Skill Updates in Compositional Robot Policies
por: Qin, Xue, et al.
Publicado: (2026)