Guardado en:
| Autores principales: | Luo, Fan-Ming, Cao, Xingchen, Qin, Rong-Jun, Yu, Yang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2022
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2206.00238 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Minimax-Optimal Reward-Agnostic Exploration in Reinforcement Learning
por: Li, Gen, et al.
Publicado: (2023)
por: Li, Gen, et al.
Publicado: (2023)
Improved Bounds for Reward-Agnostic and Reward-Free Exploration
por: Ridel, Oran, et al.
Publicado: (2026)
por: Ridel, Oran, et al.
Publicado: (2026)
Discriminative Representation Learning for Clinical Prediction
por: Zhang, Yang, et al.
Publicado: (2026)
por: Zhang, Yang, et al.
Publicado: (2026)
Provable Reward-Agnostic Preference-Based Reinforcement Learning
por: Zhan, Wenhao, et al.
Publicado: (2023)
por: Zhan, Wenhao, et al.
Publicado: (2023)
Environment Agnostic Goal-Conditioning, A Study of Reward-Free Autonomous Learning
por: Åström, Hampus, et al.
Publicado: (2025)
por: Åström, Hampus, et al.
Publicado: (2025)
Inverse Reinforcement Learning with Dynamic Reward Scaling for LLM Alignment
por: Cheng, Ruoxi, et al.
Publicado: (2025)
por: Cheng, Ruoxi, et al.
Publicado: (2025)
Uncertainty-Penalized Reinforcement Learning from Human Feedback with Diverse Reward LoRA Ensembles
por: Zhai, Yuanzhao, et al.
Publicado: (2023)
por: Zhai, Yuanzhao, et al.
Publicado: (2023)
Extrinsicaly Rewarded Soft Q Imitation Learning with Discriminator
por: Furuyama, Ryoma, et al.
Publicado: (2024)
por: Furuyama, Ryoma, et al.
Publicado: (2024)
PerPO: Perceptual Preference Optimization via Discriminative Rewarding
por: Zhu, Zining, et al.
Publicado: (2025)
por: Zhu, Zining, et al.
Publicado: (2025)
Reward-Agnostic Prompt Optimization for Text-to-Image Diffusion Models
por: Kim, Semin, et al.
Publicado: (2025)
por: Kim, Semin, et al.
Publicado: (2025)
Pre-Trained Policy Discriminators are General Reward Models
por: Dou, Shihan, et al.
Publicado: (2025)
por: Dou, Shihan, et al.
Publicado: (2025)
Physics-Informed Regularization for Domain-Agnostic Dynamical System Modeling
por: Huang, Zijie, et al.
Publicado: (2024)
por: Huang, Zijie, et al.
Publicado: (2024)
Transfer Learning for Nonparametric Contextual Dynamic Pricing
por: Wang, Fan, et al.
Publicado: (2025)
por: Wang, Fan, et al.
Publicado: (2025)
Efficient Recurrent Off-Policy RL Requires a Context-Encoder-Specific Learning Rate
por: Luo, Fan-Ming, et al.
Publicado: (2024)
por: Luo, Fan-Ming, et al.
Publicado: (2024)
Learning Discriminative Dynamics with Label Corruption for Noisy Label Detection
por: Kim, Suyeon, et al.
Publicado: (2024)
por: Kim, Suyeon, et al.
Publicado: (2024)
Boosting Adversarial Transferability via Ensemble Non-Attention
por: Zou, Yipeng, et al.
Publicado: (2025)
por: Zou, Yipeng, et al.
Publicado: (2025)
Isolation-based Spherical Ensemble Representations for Anomaly Detection
por: Cao, Yang, et al.
Publicado: (2025)
por: Cao, Yang, et al.
Publicado: (2025)
Bridging Generative and Discriminative Noisy-Label Learning via Direction-Agnostic EM Formulation
por: Liu, Fengbei, et al.
Publicado: (2023)
por: Liu, Fengbei, et al.
Publicado: (2023)
Diversity-Aware Agnostic Ensemble of Sharpness Minimizers
por: Bui, Anh, et al.
Publicado: (2024)
por: Bui, Anh, et al.
Publicado: (2024)
Reward Model Ensembles Help Mitigate Overoptimization
por: Coste, Thomas, et al.
Publicado: (2023)
por: Coste, Thomas, et al.
Publicado: (2023)
Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
por: Eisenstein, Jacob, et al.
Publicado: (2023)
por: Eisenstein, Jacob, et al.
Publicado: (2023)
Centralized Reward Agent for Knowledge Sharing and Transfer in Multi-Task Reinforcement Learning
por: Ma, Haozhe, et al.
Publicado: (2024)
por: Ma, Haozhe, et al.
Publicado: (2024)
Geometric Understanding of Discriminability and Transferability for Visual Domain Adaptation
por: Luo, You-Wei, et al.
Publicado: (2024)
por: Luo, You-Wei, et al.
Publicado: (2024)
Scalable Ensembling For Mitigating Reward Overoptimisation
por: Ahmed, Ahmed M., et al.
Publicado: (2024)
por: Ahmed, Ahmed M., et al.
Publicado: (2024)
Beyond Discriminant Patterns: On the Robustness of Decision Rule Ensembles
por: Du, Xin, et al.
Publicado: (2021)
por: Du, Xin, et al.
Publicado: (2021)
Data-Agnostic Cardinality Learning from Imperfect Workloads
por: Wu, Peizhi, et al.
Publicado: (2025)
por: Wu, Peizhi, et al.
Publicado: (2025)
Provably Efficient Reward Transfer in Reinforcement Learning with Discrete Markov Decision Processes
por: Vora, Kevin, et al.
Publicado: (2025)
por: Vora, Kevin, et al.
Publicado: (2025)
RuleExplorer: A Scalable Matrix Visualization for Understanding Tree Ensemble Classifiers
por: Li, Zhen, et al.
Publicado: (2024)
por: Li, Zhen, et al.
Publicado: (2024)
DelTA: Discriminative Token Credit Assignment for Reinforcement Learning from Verifiable Rewards
por: Zhang, Kaiyi, et al.
Publicado: (2026)
por: Zhang, Kaiyi, et al.
Publicado: (2026)
R3: Robust Rubric-Agnostic Reward Models
por: Anugraha, David, et al.
Publicado: (2025)
por: Anugraha, David, et al.
Publicado: (2025)
On the Transferability and Discriminability of Repersentation Learning in Unsupervised Domain Adaptation
por: Qiang, Wenwen, et al.
Publicado: (2025)
por: Qiang, Wenwen, et al.
Publicado: (2025)
Tree-based Ensemble Learning for Out-of-distribution Detection
por: Shen, Zhaiming, et al.
Publicado: (2024)
por: Shen, Zhaiming, et al.
Publicado: (2024)
Direct Discriminative Optimization: Your Likelihood-Based Visual Generative Model is Secretly a GAN Discriminator
por: Zheng, Kaiwen, et al.
Publicado: (2025)
por: Zheng, Kaiwen, et al.
Publicado: (2025)
Discriminative Ordering Through Ensemble Consensus
por: Ohl, Louis, et al.
Publicado: (2025)
por: Ohl, Louis, et al.
Publicado: (2025)
Stochastic Voronoi Ensembles for Anomaly Detection
por: Cao, Yang, et al.
Publicado: (2026)
por: Cao, Yang, et al.
Publicado: (2026)
Kullback-Leibler Maillard Sampling for Multi-armed Bandits with Bounded Rewards
por: Qin, Hao, et al.
Publicado: (2023)
por: Qin, Hao, et al.
Publicado: (2023)
FunPRM: Function-as-Step Process Reward Model with Meta Reward Correction for Code Generation
por: Zhang, Ruiyi, et al.
Publicado: (2026)
por: Zhang, Ruiyi, et al.
Publicado: (2026)
Towards Better Spherical Sliced-Wasserstein Distance Learning with Data-Adaptive Discriminative Projection Direction
por: Zhang, Hongliang, et al.
Publicado: (2024)
por: Zhang, Hongliang, et al.
Publicado: (2024)
Enhancing One-Shot Federated Learning Through Data and Ensemble Co-Boosting
por: Dai, Rong, et al.
Publicado: (2024)
por: Dai, Rong, et al.
Publicado: (2024)
Focal Reward: Balanced Reinforcement Learning under Rubric-Based Rewards
por: Huang, Yu, et al.
Publicado: (2026)
por: Huang, Yu, et al.
Publicado: (2026)
Ejemplares similares
-
Minimax-Optimal Reward-Agnostic Exploration in Reinforcement Learning
por: Li, Gen, et al.
Publicado: (2023) -
Improved Bounds for Reward-Agnostic and Reward-Free Exploration
por: Ridel, Oran, et al.
Publicado: (2026) -
Discriminative Representation Learning for Clinical Prediction
por: Zhang, Yang, et al.
Publicado: (2026) -
Provable Reward-Agnostic Preference-Based Reinforcement Learning
por: Zhan, Wenhao, et al.
Publicado: (2023) -
Environment Agnostic Goal-Conditioning, A Study of Reward-Free Autonomous Learning
por: Åström, Hampus, et al.
Publicado: (2025)