Efficient Skill Discovery via Regret-Aware Optimization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, He, Zhou, Ming, Zhai, Shaopeng, Sun, Ying, Xiong, Hui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
No Regrets: Investigating and Improving Regret Approximations for Curriculum Discovery
von: Rutherford, Alexander, et al.
Veröffentlicht: (2024)
von: Rutherford, Alexander, et al.
Veröffentlicht: (2024)
Regret-Based Federated Causal Discovery with Unknown Interventions
von: Baldo, Federico, et al.
Veröffentlicht: (2025)
von: Baldo, Federico, et al.
Veröffentlicht: (2025)
Skill Weaving: Efficient LLM Improvement via Modular Skillpacks
von: Li, Zhuo, et al.
Veröffentlicht: (2026)
von: Li, Zhuo, et al.
Veröffentlicht: (2026)
Reference Grounded Skill Discovery
von: Rho, Seungeun, et al.
Veröffentlicht: (2025)
von: Rho, Seungeun, et al.
Veröffentlicht: (2025)
Agentic Skill Discovery
von: Zhao, Xufeng, et al.
Veröffentlicht: (2024)
von: Zhao, Xufeng, et al.
Veröffentlicht: (2024)
Evolutionary Task Discovery: Advancing Reasoning Frontiers via Skill Composition and Complexity Scaling
von: Ye, Liqin, et al.
Veröffentlicht: (2026)
von: Ye, Liqin, et al.
Veröffentlicht: (2026)
SkillOrchestra: Learning to Route Agents via Skill Transfer
von: Wang, Jiayu, et al.
Veröffentlicht: (2026)
von: Wang, Jiayu, et al.
Veröffentlicht: (2026)
Provably Efficient Exploration in Reward Machines with Low Regret
von: Bourel, Hippolyte, et al.
Veröffentlicht: (2024)
von: Bourel, Hippolyte, et al.
Veröffentlicht: (2024)
Variational Offline Multi-agent Skill Discovery
von: Chen, Jiayu, et al.
Veröffentlicht: (2024)
von: Chen, Jiayu, et al.
Veröffentlicht: (2024)
Variance-Dependent Regret Lower Bounds for Contextual Bandits
von: He, Jiafan, et al.
Veröffentlicht: (2025)
von: He, Jiafan, et al.
Veröffentlicht: (2025)
ReSkill: Reconciling Skill Creation with Policy Optimization in Agentic RL
von: He, Zelin, et al.
Veröffentlicht: (2026)
von: He, Zelin, et al.
Veröffentlicht: (2026)
Proposer-Agent-Evaluator(PAE): Autonomous Skill Discovery For Foundation Model Internet Agents
von: Zhou, Yifei, et al.
Veröffentlicht: (2024)
von: Zhou, Yifei, et al.
Veröffentlicht: (2024)
Focused Skill Discovery: Learning to Control Specific State Variables while Minimizing Side Effects
von: Carr, Jonathan Colaço, et al.
Veröffentlicht: (2025)
von: Carr, Jonathan Colaço, et al.
Veröffentlicht: (2025)
UCPO: Uncertainty-Aware Policy Optimization
von: Zeng, Xianzhou, et al.
Veröffentlicht: (2026)
von: Zeng, Xianzhou, et al.
Veröffentlicht: (2026)
PROWL: Prioritized Regret-Driven Optimization for World Model Learning
von: Güzel, Ahmet H., et al.
Veröffentlicht: (2026)
von: Güzel, Ahmet H., et al.
Veröffentlicht: (2026)
Regret-Guided Search Control for Efficient Learning in AlphaZero
von: Tsai, Yun-Jui, et al.
Veröffentlicht: (2026)
von: Tsai, Yun-Jui, et al.
Veröffentlicht: (2026)
DRPO: Efficient Reasoning via Decoupled Reward Policy Optimization
von: Li, Gang, et al.
Veröffentlicht: (2025)
von: Li, Gang, et al.
Veröffentlicht: (2025)
Leveraging Human Feedback for Semantically-Relevant Skill Discovery
von: Hussonnois, Maxence, et al.
Veröffentlicht: (2026)
von: Hussonnois, Maxence, et al.
Veröffentlicht: (2026)
POETS: Uncertainty-Aware LLM Optimization via Compute-Efficient Policy Ensembles
von: Menet, Nicolas, et al.
Veröffentlicht: (2026)
von: Menet, Nicolas, et al.
Veröffentlicht: (2026)
Efficient Differentiable Causal Discovery via Reliable Super-Structure Learning
von: Ma, Pingchuan, et al.
Veröffentlicht: (2026)
von: Ma, Pingchuan, et al.
Veröffentlicht: (2026)
A Regret Perspective on Online Multiple Testing
von: Hao, Qingyang, et al.
Veröffentlicht: (2026)
von: Hao, Qingyang, et al.
Veröffentlicht: (2026)
OptSkills: Learning Generalizable Optimization Skills from Problem Archetypes via Cluster-Based Distillation
von: Yang, Haochen, et al.
Veröffentlicht: (2026)
von: Yang, Haochen, et al.
Veröffentlicht: (2026)
Beyond Shallow Behavior: Task-Efficient Value-Based Multi-Task Offline MARL via Skill Discovery
von: Wang, Xun, et al.
Veröffentlicht: (2025)
von: Wang, Xun, et al.
Veröffentlicht: (2025)
SUSD: Structured Unsupervised Skill Discovery through State Factorization
von: Hosseini, Seyed Mohammad Hadi, et al.
Veröffentlicht: (2026)
von: Hosseini, Seyed Mohammad Hadi, et al.
Veröffentlicht: (2026)
Towards Generalizable PDE Dynamics Forecasting via Physics-Guided Invariant Learning
von: Li, Siyang, et al.
Veröffentlicht: (2025)
von: Li, Siyang, et al.
Veröffentlicht: (2025)
Adversarial Environment Design via Regret-Guided Diffusion Models
von: Chung, Hojun, et al.
Veröffentlicht: (2024)
von: Chung, Hojun, et al.
Veröffentlicht: (2024)
Reasoning without Regret
von: Chitra, Tarun
Veröffentlicht: (2025)
von: Chitra, Tarun
Veröffentlicht: (2025)
GoldenStart: Q-Guided Priors and Entropy Control for Distilling Flow Policies
von: Zhang, He, et al.
Veröffentlicht: (2026)
von: Zhang, He, et al.
Veröffentlicht: (2026)
Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback
von: Yang, Yongjin, et al.
Veröffentlicht: (2025)
von: Yang, Yongjin, et al.
Veröffentlicht: (2025)
Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill?
von: Fan, Chenrui, et al.
Veröffentlicht: (2025)
von: Fan, Chenrui, et al.
Veröffentlicht: (2025)
OSC: Hardware Efficient W4A4 Quantization via Outlier Separation in Channel Dimension
von: Zhang, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Zhang, Zhiyuan, et al.
Veröffentlicht: (2026)
Unsupervised Skill Discovery as Exploration for Learning Agile Locomotion
von: Rho, Seungeun, et al.
Veröffentlicht: (2025)
von: Rho, Seungeun, et al.
Veröffentlicht: (2025)
Goal Discovery with Causal Capacity for Efficient Reinforcement Learning
von: Yu, Yan, et al.
Veröffentlicht: (2025)
von: Yu, Yan, et al.
Veröffentlicht: (2025)
Variance-Dependent Regret Bounds for Non-stationary Linear Bandits
von: Wang, Zhiyong, et al.
Veröffentlicht: (2024)
von: Wang, Zhiyong, et al.
Veröffentlicht: (2024)
Language Guided Skill Discovery
von: Rho, Seungeun, et al.
Veröffentlicht: (2024)
von: Rho, Seungeun, et al.
Veröffentlicht: (2024)
Efficient Discovery of Approximate Causal Abstractions via Neural Mechanism Sparsification
von: Asiaee, Amir
Veröffentlicht: (2026)
von: Asiaee, Amir
Veröffentlicht: (2026)
CATCH: Channel-Aware multivariate Time Series Anomaly Detection via Frequency Patching
von: Wu, Xingjian, et al.
Veröffentlicht: (2024)
von: Wu, Xingjian, et al.
Veröffentlicht: (2024)
MMG2Skill: Can Agents Distill In-the-Wild Guides into Self-Evolving Skills?
von: Che, Xinyu, et al.
Veröffentlicht: (2026)
von: Che, Xinyu, et al.
Veröffentlicht: (2026)
Benign Overfitting in Adversarial Training for Vision Transformers
von: Zhang, Jiaming, et al.
Veröffentlicht: (2026)
von: Zhang, Jiaming, et al.
Veröffentlicht: (2026)
Differentiable Constraint-Based Causal Discovery
von: Zhou, Jincheng, et al.
Veröffentlicht: (2025)
von: Zhou, Jincheng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
No Regrets: Investigating and Improving Regret Approximations for Curriculum Discovery
von: Rutherford, Alexander, et al.
Veröffentlicht: (2024) -
Regret-Based Federated Causal Discovery with Unknown Interventions
von: Baldo, Federico, et al.
Veröffentlicht: (2025) -
Skill Weaving: Efficient LLM Improvement via Modular Skillpacks
von: Li, Zhuo, et al.
Veröffentlicht: (2026) -
Reference Grounded Skill Discovery
von: Rho, Seungeun, et al.
Veröffentlicht: (2025) -
Agentic Skill Discovery
von: Zhao, Xufeng, et al.
Veröffentlicht: (2024)