Concurrent Learning with Aggregated States via Randomized Least Squares Value Iteration
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Yan, Bai, Qinxun, Zhang, Yiteng, Dong, Shi, Dimakopoulou, Maria, Sun, Qi, Zhou, Zhengyuan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Functional Critics Are Essential for Actor-Critic: From Off-Policy Stability to Efficient Exploration
by: Bai, Qinxun, et al.
Published: (2025)
by: Bai, Qinxun, et al.
Published: (2025)
Transformers Don't In-Context Learn Least Squares Regression
by: Hill, Joshua, et al.
Published: (2025)
by: Hill, Joshua, et al.
Published: (2025)
APFL: Analytic Personalized Federated Learning via Dual-Stream Least Squares
by: Fan, Kejia, et al.
Published: (2025)
by: Fan, Kejia, et al.
Published: (2025)
Randomized Least Squares Value Iteration itself is Joint Differentially Private
by: Lu, Haiyang, et al.
Published: (2026)
by: Lu, Haiyang, et al.
Published: (2026)
Large Legislative Models: Towards Efficient AI Policymaking in Economic Simulations
by: Gasztowtt, Henry, et al.
Published: (2024)
by: Gasztowtt, Henry, et al.
Published: (2024)
Stopping Criteria for Value Iteration on Concurrent Stochastic Reachability and Safety Games
by: Grobelna, Marta, et al.
Published: (2025)
by: Grobelna, Marta, et al.
Published: (2025)
Ordinary Least Squares is a Special Case of Transformer
by: Tan, Xiaojun, et al.
Published: (2026)
by: Tan, Xiaojun, et al.
Published: (2026)
Towards Safe Reinforcement Learning via Constraining Conditional Value-at-Risk
by: Ying, Chengyang, et al.
Published: (2022)
by: Ying, Chengyang, et al.
Published: (2022)
Stateful Evidence-Driven Retrieval-Augmented Generation with Iterative Reasoning
by: Dong, Qi, et al.
Published: (2026)
by: Dong, Qi, et al.
Published: (2026)
Deep-Learning-Aided Alternating Least Squares for Tensor CP Decomposition and Its Application to Massive MIMO Channel Estimation
by: Gong, Xiao, et al.
Published: (2023)
by: Gong, Xiao, et al.
Published: (2023)
KernelSHAP-IQ: Weighted Least-Square Optimization for Shapley Interactions
by: Fumagalli, Fabian, et al.
Published: (2024)
by: Fumagalli, Fabian, et al.
Published: (2024)
Neural Value Iteration
by: You, Yang, et al.
Published: (2025)
by: You, Yang, et al.
Published: (2025)
Pessimistic Value Iteration for Multi-Task Data Sharing in Offline Reinforcement Learning
by: Bai, Chenjia, et al.
Published: (2024)
by: Bai, Chenjia, et al.
Published: (2024)
SLIM: Sim-to-Real Legged Instructive Manipulation via Long-Horizon Visuomotor Learning
by: Zhang, Haichao, et al.
Published: (2025)
by: Zhang, Haichao, et al.
Published: (2025)
Toward Stable Value Alignment: Introducing Independent Modules for Consistent Value Guidance
by: Chen, Wenhao, et al.
Published: (2026)
by: Chen, Wenhao, et al.
Published: (2026)
Analytic Subspace Routing: How Recursive Least Squares Works in Continual Learning of Large Language Model
by: Tong, Kai, et al.
Published: (2025)
by: Tong, Kai, et al.
Published: (2025)
Adaptively Learning to Select-Rank in Online Platforms
by: Wang, Jingyuan, et al.
Published: (2024)
by: Wang, Jingyuan, et al.
Published: (2024)
ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning
by: Xu, Ziqiang, et al.
Published: (2025)
by: Xu, Ziqiang, et al.
Published: (2025)
Homomorphic Mappings for Value-Preserving State Aggregation in Markov Decision Processes
by: Zhao, Shuo, et al.
Published: (2025)
by: Zhao, Shuo, et al.
Published: (2025)
AdvantageFlow: Advantage-Weighted Least Squares for RL in Flow Models
by: Kveton, Branislav, et al.
Published: (2026)
by: Kveton, Branislav, et al.
Published: (2026)
Diverse Randomized Value Functions: A Provably Pessimistic Approach for Offline Reinforcement Learning
by: Yu, Xudong, et al.
Published: (2024)
by: Yu, Xudong, et al.
Published: (2024)
Explaining Reinforcement Learning: A Counterfactual Shapley Values Approach
by: Shi, Yiwei, et al.
Published: (2024)
by: Shi, Yiwei, et al.
Published: (2024)
Entropy-regularized Point-based Value Iteration
by: Delecki, Harrison, et al.
Published: (2024)
by: Delecki, Harrison, et al.
Published: (2024)
Highway Value Iteration Networks
by: Wang, Yuhui, et al.
Published: (2024)
by: Wang, Yuhui, et al.
Published: (2024)
Circuit-Aware SAT Solving: Guiding CDCL via Conditional Probabilities
by: Zhu, Jiaying, et al.
Published: (2025)
by: Zhu, Jiaying, et al.
Published: (2025)
Point-Based Value Iteration for POMDPs with Neural Perception Mechanisms
by: Yan, Rui, et al.
Published: (2023)
by: Yan, Rui, et al.
Published: (2023)
Divide-Then-Aggregate: An Efficient Tool Learning Method via Parallel Tool Invocation
by: Zhu, Dongsheng, et al.
Published: (2025)
by: Zhu, Dongsheng, et al.
Published: (2025)
Supply Chain Optimization via Generative Simulation and Iterative Decision Policies
by: Bai, Haoyue, et al.
Published: (2025)
by: Bai, Haoyue, et al.
Published: (2025)
Transformer-Squared: Self-adaptive LLMs
by: Sun, Qi, et al.
Published: (2025)
by: Sun, Qi, et al.
Published: (2025)
BQSched: A Non-intrusive Scheduler for Batch Concurrent Queries via Reinforcement Learning
by: Xu, Chenhao, et al.
Published: (2025)
by: Xu, Chenhao, et al.
Published: (2025)
NPSolver: Neural Poisson Solver with Iterative Physics Supervision
by: Zeng, Bocheng, et al.
Published: (2026)
by: Zeng, Bocheng, et al.
Published: (2026)
Boosting LLM via Learning from Data Iteratively and Selectively
by: Jia, Qi, et al.
Published: (2024)
by: Jia, Qi, et al.
Published: (2024)
Provable Distributional Value Iteration under Partial Observability
by: Preuett III, Larry, et al.
Published: (2025)
by: Preuett III, Larry, et al.
Published: (2025)
Long-Horizon Q-Learning: Accurate Value Learning via n-Step Inequalities
by: Abraham, Armaan A., et al.
Published: (2026)
by: Abraham, Armaan A., et al.
Published: (2026)
Offline Reinforcement Learning: Role of State Aggregation and Trajectory Data
by: Jia, Zeyu, et al.
Published: (2024)
by: Jia, Zeyu, et al.
Published: (2024)
From "Weak" Signals to Strong Models: Preference Delta Aggregation with LoRA Merging
by: Sun, Qi, et al.
Published: (2026)
by: Sun, Qi, et al.
Published: (2026)
CoPRIS: Efficient and Stable Reinforcement Learning via Concurrency-Controlled Partial Rollout with Importance Sampling
by: Qu, Zekai, et al.
Published: (2025)
by: Qu, Zekai, et al.
Published: (2025)
Pessimistic Nonlinear Least-Squares Value Iteration for Offline Reinforcement Learning
by: Di, Qiwei, et al.
Published: (2023)
by: Di, Qiwei, et al.
Published: (2023)
SVD Based Least Squares for X-Ray Pneumonia Classification Using Deep Features
by: Erdogan, Mete, et al.
Published: (2025)
by: Erdogan, Mete, et al.
Published: (2025)
Optimal Bayesian Stopping for Efficient Inference of Consistent LLM Answers
by: Huang, Jingkai, et al.
Published: (2026)
by: Huang, Jingkai, et al.
Published: (2026)
Similar Items
-
Functional Critics Are Essential for Actor-Critic: From Off-Policy Stability to Efficient Exploration
by: Bai, Qinxun, et al.
Published: (2025) -
Transformers Don't In-Context Learn Least Squares Regression
by: Hill, Joshua, et al.
Published: (2025) -
APFL: Analytic Personalized Federated Learning via Dual-Stream Least Squares
by: Fan, Kejia, et al.
Published: (2025) -
Randomized Least Squares Value Iteration itself is Joint Differentially Private
by: Lu, Haiyang, et al.
Published: (2026) -
Large Legislative Models: Towards Efficient AI Policymaking in Economic Simulations
by: Gasztowtt, Henry, et al.
Published: (2024)