Active Policy Improvement from Multiple Black-box Oracles
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Xuefeng, Yoneda, Takuma, Wang, Chaoqi, Walter, Matthew R., Chen, Yuxin |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Blending Imitation and Reinforcement Learning for Robust Policy Improvement
by: Liu, Xuefeng, et al.
Published: (2023)
by: Liu, Xuefeng, et al.
Published: (2023)
Contextual Active Model Selection
by: Liu, Xuefeng, et al.
Published: (2022)
by: Liu, Xuefeng, et al.
Published: (2022)
Instruction Learning Paradigms: A Dual Perspective on White-box and Black-box LLMs
by: Ren, Yanwei, et al.
Published: (2025)
by: Ren, Yanwei, et al.
Published: (2025)
From Black-box to Causal-box: Towards Building More Interpretable Models
by: Hwang, Inwoo, et al.
Published: (2025)
by: Hwang, Inwoo, et al.
Published: (2025)
BoSS: A Best-of-Strategies Selector as an Oracle for Deep Active Learning
by: Huseljic, Denis, et al.
Published: (2026)
by: Huseljic, Denis, et al.
Published: (2026)
DREAM: Domain-agnostic Reverse Engineering Attributes of Black-box Model
by: Li, Rongqing, et al.
Published: (2024)
by: Li, Rongqing, et al.
Published: (2024)
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
by: Wang, Chaoqi, et al.
Published: (2025)
by: Wang, Chaoqi, et al.
Published: (2025)
Improving Autoregressive Training with Dynamic Oracles
by: Yang, Jianing, et al.
Published: (2024)
by: Yang, Jianing, et al.
Published: (2024)
Localized Graph-Based Neural Dynamics Models for Terrain Manipulation
by: Liu, Chaoqi, et al.
Published: (2025)
by: Liu, Chaoqi, et al.
Published: (2025)
Monte Carlo Tree Search based Space Transfer for Black-box Optimization
by: Wang, Shukuan, et al.
Published: (2024)
by: Wang, Shukuan, et al.
Published: (2024)
BELLA: Black box model Explanations by Local Linear Approximations
by: Radulovic, Nedeljko, et al.
Published: (2023)
by: Radulovic, Nedeljko, et al.
Published: (2023)
Black-box Uncertainty Quantification Method for LLM-as-a-Judge
by: Wagner, Nico, et al.
Published: (2024)
by: Wagner, Nico, et al.
Published: (2024)
On Transfer-based Universal Attacks in Pure Black-box Setting
by: Jalwana, Mohammad A. A. K., et al.
Published: (2025)
by: Jalwana, Mohammad A. A. K., et al.
Published: (2025)
Hyperband-based Bayesian Optimization for Black-box Prompt Selection
by: Schneider, Lennart, et al.
Published: (2024)
by: Schneider, Lennart, et al.
Published: (2024)
Task-free Adaptive Meta Black-box Optimization
by: Wang, Chao, et al.
Published: (2026)
by: Wang, Chao, et al.
Published: (2026)
Provable and Practical In-Context Policy Optimization for Self-Improvement
by: Yu, Tianrun, et al.
Published: (2026)
by: Yu, Tianrun, et al.
Published: (2026)
Self-Improvement Imitation with Biologically Guided Search for Protein Design Under Oracle Budgets
by: Khanna, Ashima, et al.
Published: (2026)
by: Khanna, Ashima, et al.
Published: (2026)
A Generative Approach to Surrogate-based Black-box Attacks
by: Moraffah, Raha, et al.
Published: (2024)
by: Moraffah, Raha, et al.
Published: (2024)
Zer0-Jack: A Memory-efficient Gradient-based Jailbreaking Method for Black-box Multi-modal Large Language Models
by: Chen, Tiejin, et al.
Published: (2024)
by: Chen, Tiejin, et al.
Published: (2024)
Generative Modeling from Black-box Corruptions via Self-Consistent Stochastic Interpolants
by: Modi, Chirag, et al.
Published: (2025)
by: Modi, Chirag, et al.
Published: (2025)
Building Trust in Black-box Optimization: A Comprehensive Framework for Explainability
by: Nezami, Nazanin, et al.
Published: (2024)
by: Nezami, Nazanin, et al.
Published: (2024)
BoLT: A Benchmark to Democratize Black-box Optimization Research for Expensive LLM Tasks
by: Chew, Ruth Wan Theng, et al.
Published: (2026)
by: Chew, Ruth Wan Theng, et al.
Published: (2026)
Go Beyond Black-box Policies: Rethinking the Design of Learning Agent for Interpretable and Verifiable HVAC Control
by: An, Zhiyu, et al.
Published: (2024)
by: An, Zhiyu, et al.
Published: (2024)
Prompting Policies for Multi-step Reasoning and Tool-Use in Black-box LLMs with Iterative Distillation of Experience
by: Sayana, Krishna, et al.
Published: (2026)
by: Sayana, Krishna, et al.
Published: (2026)
Multi-Modal Manipulation via Multi-Modal Policy Consensus
by: Chen, Haonan, et al.
Published: (2025)
by: Chen, Haonan, et al.
Published: (2025)
Knowledge Editing on Black-box Large Language Models
by: Song, Xiaoshuai, et al.
Published: (2024)
by: Song, Xiaoshuai, et al.
Published: (2024)
Adaptive Divergence Regularized Policy Optimization for Fine-tuning Generative Models
by: Fan, Jiajun, et al.
Published: (2025)
by: Fan, Jiajun, et al.
Published: (2025)
AMIGO: Agentic Multi-Image Grounding Oracle Benchmark
by: Wang, Min, et al.
Published: (2026)
by: Wang, Min, et al.
Published: (2026)
Towards LLM-guided Causal Explainability for Black-box Text Classifiers
by: Bhattacharjee, Amrita, et al.
Published: (2023)
by: Bhattacharjee, Amrita, et al.
Published: (2023)
A General Black-box Adversarial Attack on Graph-based Fake News Detectors
by: Zhu, Peican, et al.
Published: (2024)
by: Zhu, Peican, et al.
Published: (2024)
Deep SPI: Safe Policy Improvement via World Models
by: Delgrange, Florent, et al.
Published: (2025)
by: Delgrange, Florent, et al.
Published: (2025)
Generalized Policy Improvement Algorithms with Theoretically Supported Sample Reuse
by: Queeney, James, et al.
Published: (2022)
by: Queeney, James, et al.
Published: (2022)
Going Beyond Heuristics by Imposing Policy Improvement as a Constraint
by: Lee, Chi-Chang, et al.
Published: (2025)
by: Lee, Chi-Chang, et al.
Published: (2025)
Transformers Provably Implement In-Context Reinforcement Learning with Policy Improvement
by: Liang, Haodong, et al.
Published: (2026)
by: Liang, Haodong, et al.
Published: (2026)
Learning to Correct for QA Reasoning with Black-box LLMs
by: Kim, Jaehyung, et al.
Published: (2024)
by: Kim, Jaehyung, et al.
Published: (2024)
Auto-Discovery-Bench: Diagnosing Structured State Tracking in Oracle-Guided Discovery
by: Chen, Tingting, et al.
Published: (2025)
by: Chen, Tingting, et al.
Published: (2025)
Fast Direct: Query-Efficient Online Black-box Guidance for Diffusion-model Target Generation
by: Tan, Kim Yong, et al.
Published: (2025)
by: Tan, Kim Yong, et al.
Published: (2025)
Curvature Dynamic Black-box Attack: revisiting adversarial robustness via dynamic curvature estimation
by: Sun, Peiran
Published: (2025)
by: Sun, Peiran
Published: (2025)
R.I.P.: A Simple Black-box Attack on Continual Test-time Adaptation
by: Hoang, Trung-Hieu, et al.
Published: (2024)
by: Hoang, Trung-Hieu, et al.
Published: (2024)
Sample-Efficient Policy Space Response Oracles with Joint Experience Best Response
by: Bighashdel, Ariyan, et al.
Published: (2026)
by: Bighashdel, Ariyan, et al.
Published: (2026)
Similar Items
-
Blending Imitation and Reinforcement Learning for Robust Policy Improvement
by: Liu, Xuefeng, et al.
Published: (2023) -
Contextual Active Model Selection
by: Liu, Xuefeng, et al.
Published: (2022) -
Instruction Learning Paradigms: A Dual Perspective on White-box and Black-box LLMs
by: Ren, Yanwei, et al.
Published: (2025) -
From Black-box to Causal-box: Towards Building More Interpretable Models
by: Hwang, Inwoo, et al.
Published: (2025) -
BoSS: A Best-of-Strategies Selector as an Oracle for Deep Active Learning
by: Huseljic, Denis, et al.
Published: (2026)