Saved in:
| Main Authors: | Luo, Fan-Ming, Cao, Xingchen, Qin, Rong-Jun, Yu, Yang |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2206.00238 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Minimax-Optimal Reward-Agnostic Exploration in Reinforcement Learning
by: Li, Gen, et al.
Published: (2023)
by: Li, Gen, et al.
Published: (2023)
Improved Bounds for Reward-Agnostic and Reward-Free Exploration
by: Ridel, Oran, et al.
Published: (2026)
by: Ridel, Oran, et al.
Published: (2026)
Discriminative Representation Learning for Clinical Prediction
by: Zhang, Yang, et al.
Published: (2026)
by: Zhang, Yang, et al.
Published: (2026)
Provable Reward-Agnostic Preference-Based Reinforcement Learning
by: Zhan, Wenhao, et al.
Published: (2023)
by: Zhan, Wenhao, et al.
Published: (2023)
Environment Agnostic Goal-Conditioning, A Study of Reward-Free Autonomous Learning
by: Åström, Hampus, et al.
Published: (2025)
by: Åström, Hampus, et al.
Published: (2025)
Inverse Reinforcement Learning with Dynamic Reward Scaling for LLM Alignment
by: Cheng, Ruoxi, et al.
Published: (2025)
by: Cheng, Ruoxi, et al.
Published: (2025)
Uncertainty-Penalized Reinforcement Learning from Human Feedback with Diverse Reward LoRA Ensembles
by: Zhai, Yuanzhao, et al.
Published: (2023)
by: Zhai, Yuanzhao, et al.
Published: (2023)
Extrinsicaly Rewarded Soft Q Imitation Learning with Discriminator
by: Furuyama, Ryoma, et al.
Published: (2024)
by: Furuyama, Ryoma, et al.
Published: (2024)
PerPO: Perceptual Preference Optimization via Discriminative Rewarding
by: Zhu, Zining, et al.
Published: (2025)
by: Zhu, Zining, et al.
Published: (2025)
Reward-Agnostic Prompt Optimization for Text-to-Image Diffusion Models
by: Kim, Semin, et al.
Published: (2025)
by: Kim, Semin, et al.
Published: (2025)
Pre-Trained Policy Discriminators are General Reward Models
by: Dou, Shihan, et al.
Published: (2025)
by: Dou, Shihan, et al.
Published: (2025)
Physics-Informed Regularization for Domain-Agnostic Dynamical System Modeling
by: Huang, Zijie, et al.
Published: (2024)
by: Huang, Zijie, et al.
Published: (2024)
Transfer Learning for Nonparametric Contextual Dynamic Pricing
by: Wang, Fan, et al.
Published: (2025)
by: Wang, Fan, et al.
Published: (2025)
Efficient Recurrent Off-Policy RL Requires a Context-Encoder-Specific Learning Rate
by: Luo, Fan-Ming, et al.
Published: (2024)
by: Luo, Fan-Ming, et al.
Published: (2024)
Learning Discriminative Dynamics with Label Corruption for Noisy Label Detection
by: Kim, Suyeon, et al.
Published: (2024)
by: Kim, Suyeon, et al.
Published: (2024)
Boosting Adversarial Transferability via Ensemble Non-Attention
by: Zou, Yipeng, et al.
Published: (2025)
by: Zou, Yipeng, et al.
Published: (2025)
Isolation-based Spherical Ensemble Representations for Anomaly Detection
by: Cao, Yang, et al.
Published: (2025)
by: Cao, Yang, et al.
Published: (2025)
Bridging Generative and Discriminative Noisy-Label Learning via Direction-Agnostic EM Formulation
by: Liu, Fengbei, et al.
Published: (2023)
by: Liu, Fengbei, et al.
Published: (2023)
Diversity-Aware Agnostic Ensemble of Sharpness Minimizers
by: Bui, Anh, et al.
Published: (2024)
by: Bui, Anh, et al.
Published: (2024)
Reward Model Ensembles Help Mitigate Overoptimization
by: Coste, Thomas, et al.
Published: (2023)
by: Coste, Thomas, et al.
Published: (2023)
Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
by: Eisenstein, Jacob, et al.
Published: (2023)
by: Eisenstein, Jacob, et al.
Published: (2023)
Centralized Reward Agent for Knowledge Sharing and Transfer in Multi-Task Reinforcement Learning
by: Ma, Haozhe, et al.
Published: (2024)
by: Ma, Haozhe, et al.
Published: (2024)
Geometric Understanding of Discriminability and Transferability for Visual Domain Adaptation
by: Luo, You-Wei, et al.
Published: (2024)
by: Luo, You-Wei, et al.
Published: (2024)
Scalable Ensembling For Mitigating Reward Overoptimisation
by: Ahmed, Ahmed M., et al.
Published: (2024)
by: Ahmed, Ahmed M., et al.
Published: (2024)
Beyond Discriminant Patterns: On the Robustness of Decision Rule Ensembles
by: Du, Xin, et al.
Published: (2021)
by: Du, Xin, et al.
Published: (2021)
Data-Agnostic Cardinality Learning from Imperfect Workloads
by: Wu, Peizhi, et al.
Published: (2025)
by: Wu, Peizhi, et al.
Published: (2025)
Provably Efficient Reward Transfer in Reinforcement Learning with Discrete Markov Decision Processes
by: Vora, Kevin, et al.
Published: (2025)
by: Vora, Kevin, et al.
Published: (2025)
RuleExplorer: A Scalable Matrix Visualization for Understanding Tree Ensemble Classifiers
by: Li, Zhen, et al.
Published: (2024)
by: Li, Zhen, et al.
Published: (2024)
DelTA: Discriminative Token Credit Assignment for Reinforcement Learning from Verifiable Rewards
by: Zhang, Kaiyi, et al.
Published: (2026)
by: Zhang, Kaiyi, et al.
Published: (2026)
R3: Robust Rubric-Agnostic Reward Models
by: Anugraha, David, et al.
Published: (2025)
by: Anugraha, David, et al.
Published: (2025)
On the Transferability and Discriminability of Repersentation Learning in Unsupervised Domain Adaptation
by: Qiang, Wenwen, et al.
Published: (2025)
by: Qiang, Wenwen, et al.
Published: (2025)
Tree-based Ensemble Learning for Out-of-distribution Detection
by: Shen, Zhaiming, et al.
Published: (2024)
by: Shen, Zhaiming, et al.
Published: (2024)
Direct Discriminative Optimization: Your Likelihood-Based Visual Generative Model is Secretly a GAN Discriminator
by: Zheng, Kaiwen, et al.
Published: (2025)
by: Zheng, Kaiwen, et al.
Published: (2025)
Discriminative Ordering Through Ensemble Consensus
by: Ohl, Louis, et al.
Published: (2025)
by: Ohl, Louis, et al.
Published: (2025)
Stochastic Voronoi Ensembles for Anomaly Detection
by: Cao, Yang, et al.
Published: (2026)
by: Cao, Yang, et al.
Published: (2026)
Kullback-Leibler Maillard Sampling for Multi-armed Bandits with Bounded Rewards
by: Qin, Hao, et al.
Published: (2023)
by: Qin, Hao, et al.
Published: (2023)
FunPRM: Function-as-Step Process Reward Model with Meta Reward Correction for Code Generation
by: Zhang, Ruiyi, et al.
Published: (2026)
by: Zhang, Ruiyi, et al.
Published: (2026)
Towards Better Spherical Sliced-Wasserstein Distance Learning with Data-Adaptive Discriminative Projection Direction
by: Zhang, Hongliang, et al.
Published: (2024)
by: Zhang, Hongliang, et al.
Published: (2024)
Enhancing One-Shot Federated Learning Through Data and Ensemble Co-Boosting
by: Dai, Rong, et al.
Published: (2024)
by: Dai, Rong, et al.
Published: (2024)
Focal Reward: Balanced Reinforcement Learning under Rubric-Based Rewards
by: Huang, Yu, et al.
Published: (2026)
by: Huang, Yu, et al.
Published: (2026)
Similar Items
-
Minimax-Optimal Reward-Agnostic Exploration in Reinforcement Learning
by: Li, Gen, et al.
Published: (2023) -
Improved Bounds for Reward-Agnostic and Reward-Free Exploration
by: Ridel, Oran, et al.
Published: (2026) -
Discriminative Representation Learning for Clinical Prediction
by: Zhang, Yang, et al.
Published: (2026) -
Provable Reward-Agnostic Preference-Based Reinforcement Learning
by: Zhan, Wenhao, et al.
Published: (2023) -
Environment Agnostic Goal-Conditioning, A Study of Reward-Free Autonomous Learning
by: Åström, Hampus, et al.
Published: (2025)