CoPS: Empowering LLM Agents with Provable Cross-Task Experience Sharing
Fuente:
arXiv
Salvato in:
| Autori principali: | Yang, Chen, Zhao, Chenyang, Gu, Quanquan, Zhou, Dongruo |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Uncertainty-Aware Reward-Free Exploration with General Function Approximation
di: Zhang, Junkai, et al.
Pubblicazione: (2024)
di: Zhang, Junkai, et al.
Pubblicazione: (2024)
Provable Zero-Shot Generalization in Offline Reinforcement Learning
di: Wang, Zhiyong, et al.
Pubblicazione: (2025)
di: Wang, Zhiyong, et al.
Pubblicazione: (2025)
How to Provably Improve Return Conditioned Supervised Learning?
di: Liu, Zhishuai, et al.
Pubblicazione: (2025)
di: Liu, Zhishuai, et al.
Pubblicazione: (2025)
Sample and Computationally Efficient Continuous-Time Reinforcement Learning with General Function Approximation
di: Zhao, Runze, et al.
Pubblicazione: (2025)
di: Zhao, Runze, et al.
Pubblicazione: (2025)
Provably Robust Adaptation for Language-Empowered Foundation Models
di: Lai, Yuni, et al.
Pubblicazione: (2025)
di: Lai, Yuni, et al.
Pubblicazione: (2025)
Provably Efficient Representation Selection in Low-rank Markov Decision Processes: From Online to Offline RL
di: Zhang, Weitong, et al.
Pubblicazione: (2021)
di: Zhang, Weitong, et al.
Pubblicazione: (2021)
Variance-Dependent Regret Lower Bounds for Contextual Bandits
di: He, Jiafan, et al.
Pubblicazione: (2025)
di: He, Jiafan, et al.
Pubblicazione: (2025)
Creative Agents: Empowering Agents with Imagination for Creative Tasks
di: Cai, Penglin, et al.
Pubblicazione: (2023)
di: Cai, Penglin, et al.
Pubblicazione: (2023)
Global Convergence and Rich Feature Learning in $L$-Layer Infinite-Width Neural Networks under $μ$P Parametrization
di: Chen, Zixiang, et al.
Pubblicazione: (2025)
di: Chen, Zixiang, et al.
Pubblicazione: (2025)
Training LLM Agents to Empower Humans
di: Ellis, Evan, et al.
Pubblicazione: (2025)
di: Ellis, Evan, et al.
Pubblicazione: (2025)
Empowering Time Series Forecasting with LLM-Agents
di: Yeh, Chin-Chia Michael, et al.
Pubblicazione: (2025)
di: Yeh, Chin-Chia Michael, et al.
Pubblicazione: (2025)
Enhancing LLM Agents for Code Generation with Possibility and Pass-rate Prioritized Experience Replay
di: Chen, Yuyang, et al.
Pubblicazione: (2024)
di: Chen, Yuyang, et al.
Pubblicazione: (2024)
Variance-Dependent Regret Bounds for Non-stationary Linear Bandits
di: Wang, Zhiyong, et al.
Pubblicazione: (2024)
di: Wang, Zhiyong, et al.
Pubblicazione: (2024)
Understanding Transferable Representation Learning and Zero-shot Transfer in CLIP
di: Chen, Zixiang, et al.
Pubblicazione: (2023)
di: Chen, Zixiang, et al.
Pubblicazione: (2023)
Co-Exploration and Co-Exploitation via Shared Structure in Multi-Task Bandits
di: Mukherjee, Sumantrak, et al.
Pubblicazione: (2025)
di: Mukherjee, Sumantrak, et al.
Pubblicazione: (2025)
Return Augmented Decision Transformer for Off-Dynamics Reinforcement Learning
di: Wang, Ruhan, et al.
Pubblicazione: (2024)
di: Wang, Ruhan, et al.
Pubblicazione: (2024)
Nearly Minimax Optimal Regret for Learning Linear Mixture Stochastic Shortest Path
di: Di, Qiwei, et al.
Pubblicazione: (2024)
di: Di, Qiwei, et al.
Pubblicazione: (2024)
On the Limits of Test-Time Compute: Sequential Reward Filtering for Better Inference
di: Yu, Yue, et al.
Pubblicazione: (2025)
di: Yu, Yue, et al.
Pubblicazione: (2025)
Diffusion Language Models Can Perform Many Tasks with Scaling and Instruction-Finetuning
di: Ye, Jiasheng, et al.
Pubblicazione: (2023)
di: Ye, Jiasheng, et al.
Pubblicazione: (2023)
Transformers Trained via Gradient Descent Can Provably Learn a Class of Teacher Models
di: Zhang, Chenyang, et al.
Pubblicazione: (2026)
di: Zhang, Chenyang, et al.
Pubblicazione: (2026)
Robust Layerwise Scaling Rules by Proper Weight Decay Tuning
di: Fan, Zhiyuan, et al.
Pubblicazione: (2025)
di: Fan, Zhiyuan, et al.
Pubblicazione: (2025)
ACT: Empowering Decision Transformer with Dynamic Programming via Advantage Conditioning
di: Gao, Chen-Xiao, et al.
Pubblicazione: (2023)
di: Gao, Chen-Xiao, et al.
Pubblicazione: (2023)
Shared Parameter Subspaces and Cross-Task Linearity in Emergently Misaligned Behavior
di: Arturi, Daniel Aarao Reis, et al.
Pubblicazione: (2025)
di: Arturi, Daniel Aarao Reis, et al.
Pubblicazione: (2025)
Reason for Future, Act for Now: A Principled Framework for Autonomous LLM Agents with Provable Sample Efficiency
di: Liu, Zhihan, et al.
Pubblicazione: (2023)
di: Liu, Zhihan, et al.
Pubblicazione: (2023)
Centralized Reward Agent for Knowledge Sharing and Transfer in Multi-Task Reinforcement Learning
di: Ma, Haozhe, et al.
Pubblicazione: (2024)
di: Ma, Haozhe, et al.
Pubblicazione: (2024)
On the Convergence of Adaptive Gradient Methods for Nonconvex Optimization
di: Zhou, Dongruo, et al.
Pubblicazione: (2018)
di: Zhou, Dongruo, et al.
Pubblicazione: (2018)
Provable In-Context Vector Arithmetic via Retrieving Task Concepts
di: Bu, Dake, et al.
Pubblicazione: (2025)
di: Bu, Dake, et al.
Pubblicazione: (2025)
On the Design of KL-Regularized Policy Gradient Algorithms for LLM Reasoning
di: Zhang, Yifan, et al.
Pubblicazione: (2025)
di: Zhang, Yifan, et al.
Pubblicazione: (2025)
Accelerated Preference Optimization for Large Language Model Alignment
di: He, Jiafan, et al.
Pubblicazione: (2024)
di: He, Jiafan, et al.
Pubblicazione: (2024)
FiCoTS: Fine-to-Coarse LLM-Enhanced Hierarchical Cross-Modality Interaction for Time Series Forecasting
di: Lyu, Yafei, et al.
Pubblicazione: (2025)
di: Lyu, Yafei, et al.
Pubblicazione: (2025)
Verifiability-First Agents: Provable Observability and Lightweight Audit Agents for Controlling Autonomous LLM Systems
di: Gupta, Abhivansh
Pubblicazione: (2025)
di: Gupta, Abhivansh
Pubblicazione: (2025)
Provable Differentially Private Computation of the Cross-Attention Mechanism
di: Ke, Yekun, et al.
Pubblicazione: (2024)
di: Ke, Yekun, et al.
Pubblicazione: (2024)
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration
di: Zhao, Heyang, et al.
Pubblicazione: (2025)
di: Zhao, Heyang, et al.
Pubblicazione: (2025)
Rephrase and Respond: Let Large Language Models Ask Better Questions for Themselves
di: Deng, Yihe, et al.
Pubblicazione: (2023)
di: Deng, Yihe, et al.
Pubblicazione: (2023)
On the Emergence of Cross-Task Linearity in the Pretraining-Finetuning Paradigm
di: Zhou, Zhanpeng, et al.
Pubblicazione: (2024)
di: Zhou, Zhanpeng, et al.
Pubblicazione: (2024)
Towards a Sharp Analysis of Offline Policy Learning for $f$-Divergence-Regularized Contextual Bandits
di: Zhao, Qingyue, et al.
Pubblicazione: (2025)
di: Zhao, Qingyue, et al.
Pubblicazione: (2025)
On the Optimal Sample Complexity of Offline Multi-Armed Bandits with KL Regularization
di: Ji, Kaixuan, et al.
Pubblicazione: (2026)
di: Ji, Kaixuan, et al.
Pubblicazione: (2026)
Near-Optimal Regret for KL-Regularized Multi-Armed Bandits
di: Ji, Kaixuan, et al.
Pubblicazione: (2026)
di: Ji, Kaixuan, et al.
Pubblicazione: (2026)
When Does Multi-Agent RL Improve LLM Workflows? Workflow, Scale, and Policy-Sharing Tradeoffs
di: Zeng, Yifan, et al.
Pubblicazione: (2026)
di: Zeng, Yifan, et al.
Pubblicazione: (2026)
Efficient UAV Swarm-Based Multi-Task Federated Learning with Dynamic Task Knowledge Sharing
di: Yang, Yubo, et al.
Pubblicazione: (2025)
di: Yang, Yubo, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Uncertainty-Aware Reward-Free Exploration with General Function Approximation
di: Zhang, Junkai, et al.
Pubblicazione: (2024) -
Provable Zero-Shot Generalization in Offline Reinforcement Learning
di: Wang, Zhiyong, et al.
Pubblicazione: (2025) -
How to Provably Improve Return Conditioned Supervised Learning?
di: Liu, Zhishuai, et al.
Pubblicazione: (2025) -
Sample and Computationally Efficient Continuous-Time Reinforcement Learning with General Function Approximation
di: Zhao, Runze, et al.
Pubblicazione: (2025) -
Provably Robust Adaptation for Language-Empowered Foundation Models
di: Lai, Yuni, et al.
Pubblicazione: (2025)