Efficient Exploration at Scale
Fuente:
arXiv
Saved in:
| Main Authors: | Asghari, Seyed Mohammad, Chute, Chris, Dwaracherla, Vikranth, Lu, Xiuyuan, Jafarnia, Mehdi, Minden, Victor, Wen, Zheng, Van Roy, Benjamin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient Exploration for LLMs
by: Dwaracherla, Vikranth, et al.
Published: (2024)
by: Dwaracherla, Vikranth, et al.
Published: (2024)
RLHF and IIA: Perverse Incentives
by: Xu, Wanqiao, et al.
Published: (2023)
by: Xu, Wanqiao, et al.
Published: (2023)
Exploration Unbound
by: Arumugam, Dilip, et al.
Published: (2024)
by: Arumugam, Dilip, et al.
Published: (2024)
Satisficing Exploration for Deep Reinforcement Learning
by: Arumugam, Dilip, et al.
Published: (2024)
by: Arumugam, Dilip, et al.
Published: (2024)
Information-Theoretic Foundations for Neural Scaling Laws
by: Jeon, Hong Jun, et al.
Published: (2024)
by: Jeon, Hong Jun, et al.
Published: (2024)
Securing Healthcare with Deep Learning: A CNN-Based Model for medical IoT Threat Detection
by: Mohamadi, Alireza, et al.
Published: (2024)
by: Mohamadi, Alireza, et al.
Published: (2024)
Choice Between Partial Trajectories: Disentangling Goals from Beliefs
by: Marklund, Henrik, et al.
Published: (2024)
by: Marklund, Henrik, et al.
Published: (2024)
Efficient Adversarial Attacks on High-dimensional Offline Bandits
by: Hosseini, Seyed Mohammad Hadi, et al.
Published: (2026)
by: Hosseini, Seyed Mohammad Hadi, et al.
Published: (2026)
Provably Efficient Exploration in Reward Machines with Low Regret
by: Bourel, Hippolyte, et al.
Published: (2024)
by: Bourel, Hippolyte, et al.
Published: (2024)
Information-Theoretic Foundations for Machine Learning
by: Jeon, Hong Jun, et al.
Published: (2024)
by: Jeon, Hong Jun, et al.
Published: (2024)
Aligning AI Agents via Information-Directed Sampling
by: Jeon, Hong Jun, et al.
Published: (2024)
by: Jeon, Hong Jun, et al.
Published: (2024)
Cognify: Supercharging Gen-AI Workflows With Hierarchical Autotuning
by: He, Zijian, et al.
Published: (2025)
by: He, Zijian, et al.
Published: (2025)
SOE: Sample-Efficient Robot Policy Self-Improvement via On-Manifold Exploration
by: Jin, Yang, et al.
Published: (2025)
by: Jin, Yang, et al.
Published: (2025)
Efficient Reinforcement Learning for Large Language Models with Intrinsic Exploration
by: Sun, Yan, et al.
Published: (2025)
by: Sun, Yan, et al.
Published: (2025)
Capacity-Constrained Continual Learning
by: Wen, Zheng, et al.
Published: (2025)
by: Wen, Zheng, et al.
Published: (2025)
Consequentialist Objectives and Catastrophe
by: Marklund, Henrik, et al.
Published: (2026)
by: Marklund, Henrik, et al.
Published: (2026)
Maintaining Plasticity in Continual Learning via Regenerative Regularization
by: Kumar, Saurabh, et al.
Published: (2023)
by: Kumar, Saurabh, et al.
Published: (2023)
Misalignment from Treating Means as Ends
by: Marklund, Henrik, et al.
Published: (2025)
by: Marklund, Henrik, et al.
Published: (2025)
More Efficient Randomized Exploration for Reinforcement Learning via Approximate Sampling
by: Ishfaq, Haque, et al.
Published: (2024)
by: Ishfaq, Haque, et al.
Published: (2024)
Scaling Algorithm Distillation for Continuous Control with Mamba
by: Beaussant, Samuel, et al.
Published: (2025)
by: Beaussant, Samuel, et al.
Published: (2025)
SUSD: Structured Unsupervised Skill Discovery through State Factorization
by: Hosseini, Seyed Mohammad Hadi, et al.
Published: (2026)
by: Hosseini, Seyed Mohammad Hadi, et al.
Published: (2026)
Guided Exploration for Efficient Relational Model Learning
by: Feng, Annie, et al.
Published: (2025)
by: Feng, Annie, et al.
Published: (2025)
Efficient Exploration for Iterative Nash Preference Optimization
by: Nan, Tianlong, et al.
Published: (2026)
by: Nan, Tianlong, et al.
Published: (2026)
COVID-19 Probability Prediction Using Machine Learning: An Infectious Approach
by: Ilani, Mohsen Asghari, et al.
Published: (2024)
by: Ilani, Mohsen Asghari, et al.
Published: (2024)
Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems
by: Carr, Jonathan Colaço, et al.
Published: (2026)
by: Carr, Jonathan Colaço, et al.
Published: (2026)
The Need for a Big World Simulator: A Scientific Challenge for Continual Learning
by: Kumar, Saurabh, et al.
Published: (2024)
by: Kumar, Saurabh, et al.
Published: (2024)
RIZE: Adaptive Regularization for Imitation Learning
by: Karimi, Adib, et al.
Published: (2025)
by: Karimi, Adib, et al.
Published: (2025)
Provably Efficient Exploration in Inverse Constrained Reinforcement Learning
by: Yue, Bo, et al.
Published: (2024)
by: Yue, Bo, et al.
Published: (2024)
Toward Efficient Exploration by Large Language Model Agents
by: Arumugam, Dilip, et al.
Published: (2025)
by: Arumugam, Dilip, et al.
Published: (2025)
NAVIX: Scaling MiniGrid Environments with JAX
by: Pignatelli, Eduardo, et al.
Published: (2024)
by: Pignatelli, Eduardo, et al.
Published: (2024)
TSENOR: Highly-Efficient Algorithm for Finding Transposable N:M Sparse Masks
by: Meng, Xiang, et al.
Published: (2025)
by: Meng, Xiang, et al.
Published: (2025)
SEQR: Secure and Efficient QR-based LoRA Routing
by: Fleshman, William, et al.
Published: (2025)
by: Fleshman, William, et al.
Published: (2025)
ML-Master: Towards AI-for-AI via Integration of Exploration and Reasoning
by: Liu, Zexi, et al.
Published: (2025)
by: Liu, Zexi, et al.
Published: (2025)
Roamify: Designing and Evaluating an LLM Based Google Chrome Extension for Personalised Itinerary Planning
by: Udandarao, Vikranth, et al.
Published: (2025)
by: Udandarao, Vikranth, et al.
Published: (2025)
Adaptive Exploration for Data-Efficient General Value Function Evaluations
by: Jain, Arushi, et al.
Published: (2024)
by: Jain, Arushi, et al.
Published: (2024)
Sample Efficient Preference Alignment in LLMs via Active Exploration
by: Mehta, Viraj, et al.
Published: (2023)
by: Mehta, Viraj, et al.
Published: (2023)
GeoChemAD: Benchmarking Unsupervised Geochemical Anomaly Detection for Mineral Exploration
by: Ding, Yihao, et al.
Published: (2026)
by: Ding, Yihao, et al.
Published: (2026)
The Impact of Quantization and Pruning on Deep Reinforcement Learning Models
by: Lu, Heng, et al.
Published: (2024)
by: Lu, Heng, et al.
Published: (2024)
SensorBench: Benchmarking LLMs in Coding-Based Sensor Processing
by: Quan, Pengrui, et al.
Published: (2024)
by: Quan, Pengrui, et al.
Published: (2024)
EEG-Based Consumer Behaviour Prediction: An Exploration from Classical Machine Learning to Graph Neural Networks
by: Afshar, Mohammad Parsa, et al.
Published: (2025)
by: Afshar, Mohammad Parsa, et al.
Published: (2025)
Similar Items
-
Efficient Exploration for LLMs
by: Dwaracherla, Vikranth, et al.
Published: (2024) -
RLHF and IIA: Perverse Incentives
by: Xu, Wanqiao, et al.
Published: (2023) -
Exploration Unbound
by: Arumugam, Dilip, et al.
Published: (2024) -
Satisficing Exploration for Deep Reinforcement Learning
by: Arumugam, Dilip, et al.
Published: (2024) -
Information-Theoretic Foundations for Neural Scaling Laws
by: Jeon, Hong Jun, et al.
Published: (2024)