Getting By Goal Misgeneralization With a Little Help From a Mentor
Fuente:
arXiv
Salvato in:
| Autori principali: | Trinh, Tu, Danesh, Mohamad H., Khanh, Nguyen X., Plaut, Benjamin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Probabilities of Chat LLMs Are Miscalibrated but Still Predict Correctness on Multiple-Choice Q&A
di: Plaut, Benjamin, et al.
Pubblicazione: (2024)
di: Plaut, Benjamin, et al.
Pubblicazione: (2024)
YRC-Bench: A Benchmark for Learning to Coordinate with Experts
di: Danesh, Mohamad H., et al.
Pubblicazione: (2025)
di: Danesh, Mohamad H., et al.
Pubblicazione: (2025)
Avoiding Catastrophe in Online Learning by Asking for Help
di: Plaut, Benjamin, et al.
Pubblicazione: (2024)
di: Plaut, Benjamin, et al.
Pubblicazione: (2024)
Safe Learning Under Irreversible Dynamics via Asking for Help
di: Plaut, Benjamin, et al.
Pubblicazione: (2025)
di: Plaut, Benjamin, et al.
Pubblicazione: (2025)
Learning When Not to Learn: Risk-Sensitive Abstention in Bandits with Unbounded Rewards
di: Liaw, Sarah, et al.
Pubblicazione: (2025)
di: Liaw, Sarah, et al.
Pubblicazione: (2025)
Safety Training Persists Through Helpfulness Optimization in LLM Agents
di: Plaut, Benjamin
Pubblicazione: (2026)
di: Plaut, Benjamin
Pubblicazione: (2026)
Contextual Pre-planning on Reward Machine Abstractions for Enhanced Transfer in Deep Reinforcement Learning
di: Azran, Guy, et al.
Pubblicazione: (2023)
di: Azran, Guy, et al.
Pubblicazione: (2023)
Towards Layer-Wise Personalized Federated Learning: Adaptive Layer Disentanglement via Conflicting Gradients
di: Nguyen, Minh Duong, et al.
Pubblicazione: (2024)
di: Nguyen, Minh Duong, et al.
Pubblicazione: (2024)
Mitigating Goal Misgeneralization via Minimax Regret
di: Sadek, Karim Abdel, et al.
Pubblicazione: (2025)
di: Sadek, Karim Abdel, et al.
Pubblicazione: (2025)
Taming the Tail in Class-Conditional GANs: Knowledge Sharing via Unconditional Training at Lower Resolutions
di: Khorram, Saeed, et al.
Pubblicazione: (2024)
di: Khorram, Saeed, et al.
Pubblicazione: (2024)
Drift Q-Learning
di: Houssaini, Anas, et al.
Pubblicazione: (2026)
di: Houssaini, Anas, et al.
Pubblicazione: (2026)
Reinforcement Learning from LLM Feedback to Counteract Goal Misgeneralization
di: Barj, Houda Nait El, et al.
Pubblicazione: (2024)
di: Barj, Houda Nait El, et al.
Pubblicazione: (2024)
ML-Driven Approaches to Combat Medicare Fraud: Advances in Class Imbalance Solutions, Feature Engineering, Adaptive Learning, and Business Impact
di: Farahmandazad, Dorsa, et al.
Pubblicazione: (2025)
di: Farahmandazad, Dorsa, et al.
Pubblicazione: (2025)
Toward Cost-efficient Adaptive Clinical Trials in Knee Osteoarthritis with Reinforcement Learning
di: Nguyen, Khanh, et al.
Pubblicazione: (2024)
di: Nguyen, Khanh, et al.
Pubblicazione: (2024)
Position: Capability Control Should be a Separate Goal From Alignment
di: Siddiqui, Shoaib Ahmed, et al.
Pubblicazione: (2026)
di: Siddiqui, Shoaib Ahmed, et al.
Pubblicazione: (2026)
METIS: Mentoring Engine for Thoughtful Inquiry & Solutions
di: Kumar, Abhinav Rajeev, et al.
Pubblicazione: (2026)
di: Kumar, Abhinav Rajeev, et al.
Pubblicazione: (2026)
SeqBattNet: A Discrete-State Physics-Informed Neural Network with Aging Adaptation for Battery Modeling
di: Tran, Khoa, et al.
Pubblicazione: (2025)
di: Tran, Khoa, et al.
Pubblicazione: (2025)
OGBench: Benchmarking Offline Goal-Conditioned RL
di: Park, Seohong, et al.
Pubblicazione: (2024)
di: Park, Seohong, et al.
Pubblicazione: (2024)
Choice Between Partial Trajectories: Disentangling Goals from Beliefs
di: Marklund, Henrik, et al.
Pubblicazione: (2024)
di: Marklund, Henrik, et al.
Pubblicazione: (2024)
Cross-Modality Controlled Molecule Generation with Diffusion Language Model
di: Zhang, Yunzhe, et al.
Pubblicazione: (2025)
di: Zhang, Yunzhe, et al.
Pubblicazione: (2025)
Reconciling Spatial and Temporal Abstractions for Goal Representation
di: Zadem, Mehdi, et al.
Pubblicazione: (2024)
di: Zadem, Mehdi, et al.
Pubblicazione: (2024)
Phase Transitions in Driven Informational Systems: A Two-Field Perspective on Learning Theory and Non-Equilibrium Chemistry
di: Khanh, Truong Xuan
Pubblicazione: (2026)
di: Khanh, Truong Xuan
Pubblicazione: (2026)
Foundations of Artificial Intelligence Frameworks: Notion and Limits of AGI
di: Bui, Khanh Gia
Pubblicazione: (2025)
di: Bui, Khanh Gia
Pubblicazione: (2025)
Accelerating Goal-Conditioned RL Algorithms and Research
di: Bortkiewicz, Michał, et al.
Pubblicazione: (2024)
di: Bortkiewicz, Michał, et al.
Pubblicazione: (2024)
Self-Supervised Goal-Reaching Results in Multi-Agent Cooperation and Exploration
di: Nimonkar, Chirayu, et al.
Pubblicazione: (2025)
di: Nimonkar, Chirayu, et al.
Pubblicazione: (2025)
HyBattNet: Hybrid Framework for Predicting the Remaining Useful Life of Lithium-Ion Batteries
di: Tran, Khoa, et al.
Pubblicazione: (2025)
di: Tran, Khoa, et al.
Pubblicazione: (2025)
How Far Can Fairness Constraints Help Recover From Biased Data?
di: Sharma, Mohit, et al.
Pubblicazione: (2023)
di: Sharma, Mohit, et al.
Pubblicazione: (2023)
Dense and Diverse Goal Coverage in Multi Goal Reinforcement Learning
di: Singh, Sagalpreet, et al.
Pubblicazione: (2025)
di: Singh, Sagalpreet, et al.
Pubblicazione: (2025)
Mode-Dependent Rectification for Stable PPO Training
di: Mohamad, Mohamad, et al.
Pubblicazione: (2026)
di: Mohamad, Mohamad, et al.
Pubblicazione: (2026)
HRLAIF: Improvements in Helpfulness and Harmlessness in Open-domain Reinforcement Learning From AI Feedback
di: Li, Ang, et al.
Pubblicazione: (2024)
di: Li, Ang, et al.
Pubblicazione: (2024)
Test-time Diverse Reasoning by Riemannian Activation Steering
di: Khanh, Ly Tran Ho, et al.
Pubblicazione: (2025)
di: Khanh, Ly Tran Ho, et al.
Pubblicazione: (2025)
HIQL: Offline Goal-Conditioned RL with Latent States as Actions
di: Park, Seohong, et al.
Pubblicazione: (2023)
di: Park, Seohong, et al.
Pubblicazione: (2023)
Autonomous Learning From Success and Failure: Goal-Conditioned Supervised Learning with Negative Feedback
di: Zhang, Zeqiang, et al.
Pubblicazione: (2025)
di: Zhang, Zeqiang, et al.
Pubblicazione: (2025)
Minimizing Collateral Damage in Activation Steering
di: Nguyen, Tam, et al.
Pubblicazione: (2026)
di: Nguyen, Tam, et al.
Pubblicazione: (2026)
Measuring Goal-Directedness
di: MacDermott, Matt, et al.
Pubblicazione: (2024)
di: MacDermott, Matt, et al.
Pubblicazione: (2024)
Dual Goal Representations
di: Park, Seohong, et al.
Pubblicazione: (2025)
di: Park, Seohong, et al.
Pubblicazione: (2025)
Stabilizing Contrastive RL: Techniques for Robotic Goal Reaching from Offline Data
di: Zheng, Chongyi, et al.
Pubblicazione: (2023)
di: Zheng, Chongyi, et al.
Pubblicazione: (2023)
Goal Exploration via Adaptive Skill Distribution for Goal-Conditioned Reinforcement Learning
di: Wu, Lisheng, et al.
Pubblicazione: (2024)
di: Wu, Lisheng, et al.
Pubblicazione: (2024)
Chain-of-Goals Hierarchical Policy for Long-Horizon Offline Goal-Conditioned RL
di: Choi, Jinwoo, et al.
Pubblicazione: (2026)
di: Choi, Jinwoo, et al.
Pubblicazione: (2026)
Proposing Hierarchical Goal-Conditioned Policy Planning in Multi-Goal Reinforcement Learning
di: Rens, Gavin B.
Pubblicazione: (2025)
di: Rens, Gavin B.
Pubblicazione: (2025)
Documenti analoghi
-
Probabilities of Chat LLMs Are Miscalibrated but Still Predict Correctness on Multiple-Choice Q&A
di: Plaut, Benjamin, et al.
Pubblicazione: (2024) -
YRC-Bench: A Benchmark for Learning to Coordinate with Experts
di: Danesh, Mohamad H., et al.
Pubblicazione: (2025) -
Avoiding Catastrophe in Online Learning by Asking for Help
di: Plaut, Benjamin, et al.
Pubblicazione: (2024) -
Safe Learning Under Irreversible Dynamics via Asking for Help
di: Plaut, Benjamin, et al.
Pubblicazione: (2025) -
Learning When Not to Learn: Risk-Sensitive Abstention in Bandits with Unbounded Rewards
di: Liaw, Sarah, et al.
Pubblicazione: (2025)