Using Common Random Numbers for Simulation-based Planning with Rollouts
Fuente:
arXiv
Guardado en:
| Autores principales: | Yadav, Sandarbh, Maliakkal, Frederic J, Khadilkar, Harshad, Kalyanakrishnan, Shivaram |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A View of the Certainty-Equivalence Method for PAC RL as an Application of the Trajectory Tree Method
por: Kalyanakrishnan, Shivaram, et al.
Publicado: (2025)
por: Kalyanakrishnan, Shivaram, et al.
Publicado: (2025)
Transformers are Expressive, But Are They Expressive Enough for Regression?
por: Nath, Swaroop, et al.
Publicado: (2024)
por: Nath, Swaroop, et al.
Publicado: (2024)
Efficiency Boost in Decentralized Optimization: Reimagining Neighborhood Aggregation with Minimal Overhead
por: Kalwar, Durgesh, et al.
Publicado: (2025)
por: Kalwar, Durgesh, et al.
Publicado: (2025)
AEGIS: An Agent for Extraction and Geographic Identification in Scholarly Proceedings
por: Vishesh, Om, et al.
Publicado: (2025)
por: Vishesh, Om, et al.
Publicado: (2025)
On-line Learning in Tree MDPs by Treating Policies as Bandit Arms
por: Shah, Anvay, et al.
Publicado: (2026)
por: Shah, Anvay, et al.
Publicado: (2026)
A Meta Reinforcement Learning Approach to Goals-Based Wealth Management
por: Das, Sanjiv R., et al.
Publicado: (2026)
por: Das, Sanjiv R., et al.
Publicado: (2026)
Image-Based Malware Classification Using QR and Aztec Codes
por: Khadilkar, Atharva, et al.
Publicado: (2024)
por: Khadilkar, Atharva, et al.
Publicado: (2024)
Howard's Policy Iteration is Subexponential for Deterministic Markov Decision Problems with Rewards of Fixed Bit-size and Arbitrary Discount Factor
por: Mukherjee, Dibyangshu, et al.
Publicado: (2025)
por: Mukherjee, Dibyangshu, et al.
Publicado: (2025)
Efficient Computation of Blackwell Optimal Policies using Rational Functions
por: Mukherjee, Dibyangshu, et al.
Publicado: (2025)
por: Mukherjee, Dibyangshu, et al.
Publicado: (2025)
DiffusionRollout: Uncertainty-Aware Rollout Planning in Long-Horizon PDE Solving
por: Yoo, Seungwoo, et al.
Publicado: (2026)
por: Yoo, Seungwoo, et al.
Publicado: (2026)
Causal-Counterfactual RAG: The Integration of Causal-Counterfactual Reasoning into RAG
por: Khadilkar, Harshad, et al.
Publicado: (2025)
por: Khadilkar, Harshad, et al.
Publicado: (2025)
Redistributing Rewards Across Time and Agents for Multi-Agent Reinforcement Learning
por: Kapoor, Aditya, et al.
Publicado: (2025)
por: Kapoor, Aditya, et al.
Publicado: (2025)
Sustainable AI Training via Hardware-Software Co-Design on NVIDIA, AMD, and Emerging GPU Architectures
por: Makin, Yashasvi, et al.
Publicado: (2025)
por: Makin, Yashasvi, et al.
Publicado: (2025)
Quantization Undoes Alignment: Bias Emergence in Compressed LLMs Across Models and Precision Levels
por: Rath, Plawan Kumar, et al.
Publicado: (2026)
por: Rath, Plawan Kumar, et al.
Publicado: (2026)
Agent-Temporal Credit Assignment for Optimal Policy Preservation in Sparse Multi-Agent Reinforcement Learning
por: Kapoor, Aditya, et al.
Publicado: (2024)
por: Kapoor, Aditya, et al.
Publicado: (2024)
Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR
por: Wang, Tao, et al.
Publicado: (2026)
por: Wang, Tao, et al.
Publicado: (2026)
Weight Pruning Amplifies Bias: A Multi-Method Study of Compressed LLMs for Edge AI
por: Rath, Plawan Kumar, et al.
Publicado: (2026)
por: Rath, Plawan Kumar, et al.
Publicado: (2026)
Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning
por: Xu, Yixuan Even, et al.
Publicado: (2025)
por: Xu, Yixuan Even, et al.
Publicado: (2025)
ROAST: Rollout-based On-distribution Activation Steering Technique
por: Su, Xuanbo, et al.
Publicado: (2026)
por: Su, Xuanbo, et al.
Publicado: (2026)
DeepClean: Integrated Distortion Identification and Algorithm Selection for Rectifying Image Corruptions
por: Kapoor, Aditya, et al.
Publicado: (2024)
por: Kapoor, Aditya, et al.
Publicado: (2024)
On Rollouts in Model-Based Reinforcement Learning
por: Frauenknecht, Bernd, et al.
Publicado: (2025)
por: Frauenknecht, Bernd, et al.
Publicado: (2025)
Learning from the Right Rollouts: Data Attribution for PPO-based LLM Post-Training
por: Shu, Dong, et al.
Publicado: (2026)
por: Shu, Dong, et al.
Publicado: (2026)
Leveraging Domain Knowledge for Efficient Reward Modelling in RLHF: A Case-Study in E-Commerce Opinion Summarization
por: Nath, Swaroop, et al.
Publicado: (2024)
por: Nath, Swaroop, et al.
Publicado: (2024)
Guided Cooperation in Hierarchical Reinforcement Learning via Model-based Rollout
por: Wang, Haoran, et al.
Publicado: (2023)
por: Wang, Haoran, et al.
Publicado: (2023)
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning
por: Liu, Wenpu, et al.
Publicado: (2026)
por: Liu, Wenpu, et al.
Publicado: (2026)
QuRL: Efficient Reinforcement Learning with Quantized Rollout
por: Li, Yuhang, et al.
Publicado: (2026)
por: Li, Yuhang, et al.
Publicado: (2026)
Controlling Transient Amplification Improves Long-horizon Rollouts
por: Pervez, Adeel, et al.
Publicado: (2026)
por: Pervez, Adeel, et al.
Publicado: (2026)
Optimistic Model Rollouts for Pessimistic Offline Policy Optimization
por: Zhai, Yuanzhao, et al.
Publicado: (2024)
por: Zhai, Yuanzhao, et al.
Publicado: (2024)
Maximum Entropy Exploration Without the Rollouts
por: Adamczyk, Jacob, et al.
Publicado: (2026)
por: Adamczyk, Jacob, et al.
Publicado: (2026)
Variational Secret Common Randomness Extraction
por: Li, Xinyang, et al.
Publicado: (2025)
por: Li, Xinyang, et al.
Publicado: (2025)
Decoding Speculative Decoding
por: Yan, Minghao, et al.
Publicado: (2024)
por: Yan, Minghao, et al.
Publicado: (2024)
PolyThrottle: Energy-efficient Neural Network Inference on Edge Devices
por: Yan, Minghao, et al.
Publicado: (2023)
por: Yan, Minghao, et al.
Publicado: (2023)
Horizon Imagination: Efficient On-Policy Rollout in Diffusion World Models
por: Cohen, Lior, et al.
Publicado: (2026)
por: Cohen, Lior, et al.
Publicado: (2026)
Heddle: A Distributed Orchestration System for Agentic RL Rollout
por: Zhang, Zili, et al.
Publicado: (2026)
por: Zhang, Zili, et al.
Publicado: (2026)
Memory-Conditioned Flow-Matching for Stable Autoregressive PDE Rollouts
por: Armegioiu, Victor
Publicado: (2026)
por: Armegioiu, Victor
Publicado: (2026)
On the Trade-off between the Number of Nodes and the Number of Trees in a Random Forest
por: Akutsu, Tatsuya, et al.
Publicado: (2023)
por: Akutsu, Tatsuya, et al.
Publicado: (2023)
Selective Rollout: Mid-Trajectory Termination for Multi-Sample Agent RL
por: Zhai, Zhiyuan, et al.
Publicado: (2026)
por: Zhai, Zhiyuan, et al.
Publicado: (2026)
Improving Sampling Efficiency in RLVR through Adaptive Rollout and Response Reuse
por: Zhang, Yuheng, et al.
Publicado: (2025)
por: Zhang, Yuheng, et al.
Publicado: (2025)
ImagineBench: Evaluating Reinforcement Learning with Large Language Model Rollouts
por: Pang, Jing-Cheng, et al.
Publicado: (2025)
por: Pang, Jing-Cheng, et al.
Publicado: (2025)
Model-Agnostic Knowledge Guided Correction for Improved Neural Surrogate Rollout
por: Srikishan, Bharat, et al.
Publicado: (2025)
por: Srikishan, Bharat, et al.
Publicado: (2025)
Ejemplares similares
-
A View of the Certainty-Equivalence Method for PAC RL as an Application of the Trajectory Tree Method
por: Kalyanakrishnan, Shivaram, et al.
Publicado: (2025) -
Transformers are Expressive, But Are They Expressive Enough for Regression?
por: Nath, Swaroop, et al.
Publicado: (2024) -
Efficiency Boost in Decentralized Optimization: Reimagining Neighborhood Aggregation with Minimal Overhead
por: Kalwar, Durgesh, et al.
Publicado: (2025) -
AEGIS: An Agent for Extraction and Geographic Identification in Scholarly Proceedings
por: Vishesh, Om, et al.
Publicado: (2025) -
On-line Learning in Tree MDPs by Treating Policies as Bandit Arms
por: Shah, Anvay, et al.
Publicado: (2026)