Your Policy Regularizer is Secretly an Adversary
Fuente:
arXiv
Saved in:
| Main Authors: | Brekelmans, Rob, Genewein, Tim, Grau-Moya, Jordi, Delétang, Grégoire, Kunesch, Markus, Legg, Shane, Ortega, Pedro |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Amortized Planning with Large-Scale Transformers: A Case Study on Chess
by: Ruoss, Anian, et al.
Published: (2024)
by: Ruoss, Anian, et al.
Published: (2024)
Learning Universal Predictors
by: Grau-Moya, Jordi, et al.
Published: (2024)
by: Grau-Moya, Jordi, et al.
Published: (2024)
Why is prompting hard? Understanding prompts on binary sequence predictors
by: Wenliang, Li Kevin, et al.
Published: (2025)
by: Wenliang, Li Kevin, et al.
Published: (2025)
Understanding Prompt Tuning and In-Context Learning via Meta-Learning
by: Genewein, Tim, et al.
Published: (2025)
by: Genewein, Tim, et al.
Published: (2025)
Language Modeling Is Compression
by: Delétang, Grégoire, et al.
Published: (2023)
by: Delétang, Grégoire, et al.
Published: (2023)
Variational Representations of Annealing Paths: Bregman Information under Monotonic Embedding
by: Brekelmans, Rob, et al.
Published: (2022)
by: Brekelmans, Rob, et al.
Published: (2022)
Measurement-Aligned Sampling for Inverse Problem
by: Zhang, Shaorong, et al.
Published: (2025)
by: Zhang, Shaorong, et al.
Published: (2025)
LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities
by: Schmied, Thomas, et al.
Published: (2025)
by: Schmied, Thomas, et al.
Published: (2025)
Probabilistic Inference in Language Models via Twisted Sequential Monte Carlo
by: Zhao, Stephen, et al.
Published: (2024)
by: Zhao, Stephen, et al.
Published: (2024)
Reducing the Probability of Undesirable Outputs in Language Models Using Probabilistic Inference
by: Zhao, Stephen, et al.
Published: (2025)
by: Zhao, Stephen, et al.
Published: (2025)
Distributional Bellman Operators over Mean Embeddings
by: Wenliang, Li Kevin, et al.
Published: (2023)
by: Wenliang, Li Kevin, et al.
Published: (2023)
Compression via Pre-trained Transformers: A Study on Byte-Level Multimodal Data
by: Heurtel-Depeiges, David, et al.
Published: (2024)
by: Heurtel-Depeiges, David, et al.
Published: (2024)
Partition Tree Weighting for Non-Stationary Stochastic Bandits
by: Veness, Joel, et al.
Published: (2025)
by: Veness, Joel, et al.
Published: (2025)
Improving Mutual Information Estimation with Annealed and Energy-Based Bounds
by: Brekelmans, Rob, et al.
Published: (2023)
by: Brekelmans, Rob, et al.
Published: (2023)
Generation Order and Parallel Decoding in Masked Diffusion Models: An Information-Theoretic Perspective
by: Zhang, Shaorong, et al.
Published: (2026)
by: Zhang, Shaorong, et al.
Published: (2026)
Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning
by: Li, Shangzhe, et al.
Published: (2026)
by: Li, Shangzhe, et al.
Published: (2026)
LMAct: A Benchmark for In-Context Imitation Learning with Long Multimodal Demonstrations
by: Ruoss, Anian, et al.
Published: (2024)
by: Ruoss, Anian, et al.
Published: (2024)
A Computational Framework for Solving Wasserstein Lagrangian Flows
by: Neklyudov, Kirill, et al.
Published: (2023)
by: Neklyudov, Kirill, et al.
Published: (2023)
Your Transformer is Secretly Linear
by: Razzhigaev, Anton, et al.
Published: (2024)
by: Razzhigaev, Anton, et al.
Published: (2024)
Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer
by: Liu, Zhihan, et al.
Published: (2024)
by: Liu, Zhihan, et al.
Published: (2024)
Generalized Flow Matching for Transition Dynamics Modeling
by: Wang, Haibo, et al.
Published: (2024)
by: Wang, Haibo, et al.
Published: (2024)
It Takes Two: Your GRPO Is Secretly DPO
by: Wu, Yihong, et al.
Published: (2025)
by: Wu, Yihong, et al.
Published: (2025)
Annealed Importance Sampling with q-Paths
by: Brekelmans, Rob, et al.
Published: (2020)
by: Brekelmans, Rob, et al.
Published: (2020)
SGD and Weight Decay Secretly Minimize the Rank of Your Neural Network
by: Galanti, Tomer, et al.
Published: (2022)
by: Galanti, Tomer, et al.
Published: (2022)
Your Classifier Can Be Secretly a Likelihood-Based OOD Detector
by: Burapacheep, Jirayu, et al.
Published: (2024)
by: Burapacheep, Jirayu, et al.
Published: (2024)
Discrete Stochastic Localization for Non-autoregressive Generation
by: Wu, Yunshu, et al.
Published: (2026)
by: Wu, Yunshu, et al.
Published: (2026)
Variance Covariance Regularization Enforces Pairwise Independence in Self-Supervised Representations
by: Mialon, Grégoire, et al.
Published: (2022)
by: Mialon, Grégoire, et al.
Published: (2022)
Your Dense Retriever is Secretly an Expeditious Reasoner
by: Zhang, Yichi, et al.
Published: (2025)
by: Zhang, Yichi, et al.
Published: (2025)
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model
by: Wang, Xiyao, et al.
Published: (2025)
by: Wang, Xiyao, et al.
Published: (2025)
Multiverse: Your Language Models Secretly Decide How to Parallelize and Merge Generation
by: Yang, Xinyu, et al.
Published: (2025)
by: Yang, Xinyu, et al.
Published: (2025)
From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function
by: Rafailov, Rafael, et al.
Published: (2024)
by: Rafailov, Rafael, et al.
Published: (2024)
Your Pre-trained LLM is Secretly an Unsupervised Confidence Calibrator
by: Luo, Beier, et al.
Published: (2025)
by: Luo, Beier, et al.
Published: (2025)
Your Mixture-of-Experts LLM Is Secretly an Embedding Model For Free
by: Li, Ziyue, et al.
Published: (2024)
by: Li, Ziyue, et al.
Published: (2024)
Your Diffusion Model is Secretly a Noise Classifier and Benefits from Contrastive Training
by: Wu, Yunshu, et al.
Published: (2024)
by: Wu, Yunshu, et al.
Published: (2024)
Scalable Spatio-Temporal SE(3) Diffusion for Long-Horizon Protein Dynamics
by: Shoghi, Nima, et al.
Published: (2026)
by: Shoghi, Nima, et al.
Published: (2026)
Learning Overspecified Gaussian Mixtures Exponentially Fast with the EM Algorithm
by: Assylbekov, Zhenisbek, et al.
Published: (2025)
by: Assylbekov, Zhenisbek, et al.
Published: (2025)
Your Learned Constraint is Secretly a Backward Reachable Tube
by: Qadri, Mohamad, et al.
Published: (2025)
by: Qadri, Mohamad, et al.
Published: (2025)
Discrete Stochastic Localization for Non-autoregressive Generation
by: Wu, Yunshu, et al.
Published: (2026)
by: Wu, Yunshu, et al.
Published: (2026)
Representation Convergence: Mutual Distillation is Secretly a Form of Regularization
by: Xie, Zhengpeng, et al.
Published: (2025)
by: Xie, Zhengpeng, et al.
Published: (2025)
Zeroth-Order Optimization is Secretly Single-Step Policy Optimization
by: Qiu, Junbin, et al.
Published: (2025)
by: Qiu, Junbin, et al.
Published: (2025)
Similar Items
-
Amortized Planning with Large-Scale Transformers: A Case Study on Chess
by: Ruoss, Anian, et al.
Published: (2024) -
Learning Universal Predictors
by: Grau-Moya, Jordi, et al.
Published: (2024) -
Why is prompting hard? Understanding prompts on binary sequence predictors
by: Wenliang, Li Kevin, et al.
Published: (2025) -
Understanding Prompt Tuning and In-Context Learning via Meta-Learning
by: Genewein, Tim, et al.
Published: (2025) -
Language Modeling Is Compression
by: Delétang, Grégoire, et al.
Published: (2023)