Learning the Target Network in Function Space
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Asadi, Kavosh, Liu, Yao, Sabach, Shoham, Yin, Ming, Fakoor, Rasool |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TAIL: Task-specific Adapters for Imitation Learning with Large Pretrained Models
von: Liu, Zuxin, et al.
Veröffentlicht: (2023)
von: Liu, Zuxin, et al.
Veröffentlicht: (2023)
Displacement-Resistant Extensions of DPO with Nonconvex $f$-Divergences
von: Pipano, Idan, et al.
Veröffentlicht: (2026)
von: Pipano, Idan, et al.
Veröffentlicht: (2026)
C2-DPO: Constrained Controlled Direct Preference Optimization
von: Asadi, Kavosh, et al.
Veröffentlicht: (2025)
von: Asadi, Kavosh, et al.
Veröffentlicht: (2025)
Budgeting Counterfactual for Offline RL
von: Liu, Yao, et al.
Veröffentlicht: (2023)
von: Liu, Yao, et al.
Veröffentlicht: (2023)
Offline Learning and Forgetting for Reasoning with Large Language Models
von: Ni, Tianwei, et al.
Veröffentlicht: (2025)
von: Ni, Tianwei, et al.
Veröffentlicht: (2025)
From Demonstrations to Rewards: Alignment Without Explicit Human Preferences
von: Zeng, Siliang, et al.
Veröffentlicht: (2025)
von: Zeng, Siliang, et al.
Veröffentlicht: (2025)
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training
von: Fakoor, Rasool, et al.
Veröffentlicht: (2026)
von: Fakoor, Rasool, et al.
Veröffentlicht: (2026)
Adjoint sharding for very long context training of state space models
von: Xu, Xingzi, et al.
Veröffentlicht: (2025)
von: Xu, Xingzi, et al.
Veröffentlicht: (2025)
EXTRACT: Efficient Policy Learning by Extracting Transferable Robot Skills from Offline Data
von: Zhang, Jesse, et al.
Veröffentlicht: (2024)
von: Zhang, Jesse, et al.
Veröffentlicht: (2024)
Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe
von: Thekumparampil, Kiran Koshy, et al.
Veröffentlicht: (2024)
von: Thekumparampil, Kiran Koshy, et al.
Veröffentlicht: (2024)
Adaptive Online Learning with LSTM Networks for Energy Price Prediction
von: Salihoglu, Salih, et al.
Veröffentlicht: (2025)
von: Salihoglu, Salih, et al.
Veröffentlicht: (2025)
PAIR-Former: Budgeted Relational Multi-Instance Learning for Functional miRNA Target Prediction
von: Yin, Jiaqi, et al.
Veröffentlicht: (2026)
von: Yin, Jiaqi, et al.
Veröffentlicht: (2026)
Balancing Coverage and Draft Latency in Vocabulary Trimming for Faster Speculative Decoding
von: Shoham, Ofir Ben
Veröffentlicht: (2026)
von: Shoham, Ofir Ben
Veröffentlicht: (2026)
Understanding the Challenges in Iterative Generative Optimization with LLMs
von: Nie, Allen, et al.
Veröffentlicht: (2026)
von: Nie, Allen, et al.
Veröffentlicht: (2026)
EIDOS: Latent-Space Predictive Learning for Time Series Foundation Models
von: Zhou, Xinxing, et al.
Veröffentlicht: (2026)
von: Zhou, Xinxing, et al.
Veröffentlicht: (2026)
CEL: A Continual Learning Model for Disease Outbreak Prediction by Leveraging Domain Adaptation via Elastic Weight Consolidation
von: Aslam, Saba, et al.
Veröffentlicht: (2024)
von: Aslam, Saba, et al.
Veröffentlicht: (2024)
Understanding Understanding: A Pragmatic Framework Motivated by Large Language Models
von: Leyton-Brown, Kevin, et al.
Veröffentlicht: (2024)
von: Leyton-Brown, Kevin, et al.
Veröffentlicht: (2024)
CPLLM: Clinical Prediction with Large Language Models
von: Shoham, Ofir Ben, et al.
Veröffentlicht: (2023)
von: Shoham, Ofir Ben, et al.
Veröffentlicht: (2023)
Directional-Clamp PPO
von: Karpel, Gilad, et al.
Veröffentlicht: (2025)
von: Karpel, Gilad, et al.
Veröffentlicht: (2025)
Target Networks and Over-parameterization Stabilize Off-policy Bootstrapping with Function Approximation
von: Che, Fengdi, et al.
Veröffentlicht: (2024)
von: Che, Fengdi, et al.
Veröffentlicht: (2024)
In-Context Learning for Pure Exploration in Continuous Spaces
von: Russo, Alessio, et al.
Veröffentlicht: (2026)
von: Russo, Alessio, et al.
Veröffentlicht: (2026)
Beyond Scaleup: Knowledge-aware Parsimony Learning from Deep Networks
von: Yao, Quanming, et al.
Veröffentlicht: (2024)
von: Yao, Quanming, et al.
Veröffentlicht: (2024)
Learning to Reason Efficiently with Discounted Reinforcement Learning
von: Ayoub, Alex, et al.
Veröffentlicht: (2025)
von: Ayoub, Alex, et al.
Veröffentlicht: (2025)
Minimalist Softmax Attention Provably Learns Constrained Boolean Functions
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2025)
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2025)
Adaptive $Q$-Network: On-the-fly Target Selection for Deep Reinforcement Learning
von: Vincent, Théo, et al.
Veröffentlicht: (2024)
von: Vincent, Théo, et al.
Veröffentlicht: (2024)
FSP-Laplace: Function-Space Priors for the Laplace Approximation in Bayesian Deep Learning
von: Cinquin, Tristan, et al.
Veröffentlicht: (2024)
von: Cinquin, Tristan, et al.
Veröffentlicht: (2024)
Adapting the Behavior of Reinforcement Learning Agents to Changing Action Spaces and Reward Functions
von: de la Rosa, Raul, et al.
Veröffentlicht: (2026)
von: de la Rosa, Raul, et al.
Veröffentlicht: (2026)
On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures
von: Yin, Ming, et al.
Veröffentlicht: (2025)
von: Yin, Ming, et al.
Veröffentlicht: (2025)
Towards Instance-wise Personalized Federated Learning via Semi-Implicit Bayesian Prompt Tuning
von: Ye, Tiandi, et al.
Veröffentlicht: (2025)
von: Ye, Tiandi, et al.
Veröffentlicht: (2025)
A comparative study of deep learning and ensemble learning to extend the horizon of traffic forecasting
von: Zheng, Xiao, et al.
Veröffentlicht: (2025)
von: Zheng, Xiao, et al.
Veröffentlicht: (2025)
Variable Assignment Invariant Neural Networks for Learning Logic Programs
von: Phua, Yin Jun, et al.
Veröffentlicht: (2024)
von: Phua, Yin Jun, et al.
Veröffentlicht: (2024)
Embedding Knowledge Graph in Function Spaces
von: Teyou, Louis Mozart Kamdem, et al.
Veröffentlicht: (2024)
von: Teyou, Louis Mozart Kamdem, et al.
Veröffentlicht: (2024)
High-order Regularization for Machine Learning and Learning-based Control
von: Liu, Xinghua, et al.
Veröffentlicht: (2025)
von: Liu, Xinghua, et al.
Veröffentlicht: (2025)
APAR: Modeling Irregular Target Functions in Tabular Regression via Arithmetic-Aware Pre-Training and Adaptive-Regularized Fine-Tuning
von: Wu, Hong-Wei, et al.
Veröffentlicht: (2024)
von: Wu, Hong-Wei, et al.
Veröffentlicht: (2024)
HiF-DTA: Hierarchical Feature Learning Network for Drug-Target Affinity Prediction
von: Li, Minghui, et al.
Veröffentlicht: (2025)
von: Li, Minghui, et al.
Veröffentlicht: (2025)
Relaxing Continuous Constraints of Equivariant Graph Neural Networks for Physical Dynamics Learning
von: Zheng, Zinan, et al.
Veröffentlicht: (2024)
von: Zheng, Zinan, et al.
Veröffentlicht: (2024)
Brain Network Classification Based on Graph Contrastive Learning and Graph Transformer
von: Zhu, ZhiTeng, et al.
Veröffentlicht: (2025)
von: Zhu, ZhiTeng, et al.
Veröffentlicht: (2025)
Offline Policy Evaluation for Reinforcement Learning with Adaptively Collected Data
von: Madhow, Sunil, et al.
Veröffentlicht: (2023)
von: Madhow, Sunil, et al.
Veröffentlicht: (2023)
Target-Aligned Reinforcement Learning
von: Pleiss, Leonard S., et al.
Veröffentlicht: (2026)
von: Pleiss, Leonard S., et al.
Veröffentlicht: (2026)
A Unified Kernel for Neural Network Learning
von: Zhang, Shao-Qun, et al.
Veröffentlicht: (2024)
von: Zhang, Shao-Qun, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
TAIL: Task-specific Adapters for Imitation Learning with Large Pretrained Models
von: Liu, Zuxin, et al.
Veröffentlicht: (2023) -
Displacement-Resistant Extensions of DPO with Nonconvex $f$-Divergences
von: Pipano, Idan, et al.
Veröffentlicht: (2026) -
C2-DPO: Constrained Controlled Direct Preference Optimization
von: Asadi, Kavosh, et al.
Veröffentlicht: (2025) -
Budgeting Counterfactual for Offline RL
von: Liu, Yao, et al.
Veröffentlicht: (2023) -
Offline Learning and Forgetting for Reasoning with Large Language Models
von: Ni, Tianwei, et al.
Veröffentlicht: (2025)