Advantage Alignment Algorithms
Fuente:
arXiv
Salvato in:
| Autori principali: | Duque, Juan Agustin, Aghajohari, Milad, Cooijmans, Tim, Ciuca, Razvan, Zhang, Tianyu, Gidel, Gauthier, Courville, Aaron |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
LOQA: Learning with Opponent Q-Learning Awareness
di: Aghajohari, Milad, et al.
Pubblicazione: (2024)
di: Aghajohari, Milad, et al.
Pubblicazione: (2024)
Best Response Shaping
di: Aghajohari, Milad, et al.
Pubblicazione: (2024)
di: Aghajohari, Milad, et al.
Pubblicazione: (2024)
Towards Sustainable Investment Policies Informed by Opponent Shaping
di: Duque, Juan Agustin, et al.
Pubblicazione: (2026)
di: Duque, Juan Agustin, et al.
Pubblicazione: (2026)
Learning Robust Social Strategies with Large Language Models
di: Piche, Dereck, et al.
Pubblicazione: (2025)
di: Piche, Dereck, et al.
Pubblicazione: (2025)
The Markovian Thinker: Architecture-Agnostic Linear Scaling of Reasoning
di: Aghajohari, Milad, et al.
Pubblicazione: (2025)
di: Aghajohari, Milad, et al.
Pubblicazione: (2025)
VinePPO: Refining Credit Assignment in RL Training of LLMs
di: Kazemnejad, Amirhossein, et al.
Pubblicazione: (2024)
di: Kazemnejad, Amirhossein, et al.
Pubblicazione: (2024)
Why Open Source? A Game-Theoretic Analysis of the AI Race
di: Mladenovic, Andjela, et al.
Pubblicazione: (2026)
di: Mladenovic, Andjela, et al.
Pubblicazione: (2026)
Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space
di: Schwinn, Leo, et al.
Pubblicazione: (2024)
di: Schwinn, Leo, et al.
Pubblicazione: (2024)
Sarah Frank-Wolfe: Methods for Constrained Optimization with Best Rates and Practical Features
di: Beznosikov, Aleksandr, et al.
Pubblicazione: (2023)
di: Beznosikov, Aleksandr, et al.
Pubblicazione: (2023)
Omega: Optimistic EMA Gradients
di: Ramirez, Juan, et al.
Pubblicazione: (2023)
di: Ramirez, Juan, et al.
Pubblicazione: (2023)
Performative Prediction with Neural Networks
di: Mofakhami, Mehrnaz, et al.
Pubblicazione: (2023)
di: Mofakhami, Mehrnaz, et al.
Pubblicazione: (2023)
Adversarial Alignment for LLMs Requires Simpler, Reproducible, and More Measurable Objectives
di: Schwinn, Leo, et al.
Pubblicazione: (2025)
di: Schwinn, Leo, et al.
Pubblicazione: (2025)
Proving Linear Mode Connectivity of Neural Networks via Optimal Transport
di: Ferbach, Damien, et al.
Pubblicazione: (2023)
di: Ferbach, Damien, et al.
Pubblicazione: (2023)
General Causal Imputation via Synthetic Interventions
di: Jiralerspong, Marco, et al.
Pubblicazione: (2024)
di: Jiralerspong, Marco, et al.
Pubblicazione: (2024)
Expected flow networks in stochastic environments and two-player zero-sum games
di: Jiralerspong, Marco, et al.
Pubblicazione: (2023)
di: Jiralerspong, Marco, et al.
Pubblicazione: (2023)
Discrete Compositional Generation via General Soft Operators and Robust Reinforcement Learning
di: Jiralerspong, Marco, et al.
Pubblicazione: (2025)
di: Jiralerspong, Marco, et al.
Pubblicazione: (2025)
When is Momentum Extragradient Optimal? A Polynomial-Based Analysis
di: Kim, Junhyung Lyle, et al.
Pubblicazione: (2022)
di: Kim, Junhyung Lyle, et al.
Pubblicazione: (2022)
On the Stability of Iterative Retraining of Generative Models on their own Data
di: Bertrand, Quentin, et al.
Pubblicazione: (2023)
di: Bertrand, Quentin, et al.
Pubblicazione: (2023)
In-Context Learning Can Re-learn Forbidden Tasks
di: Xhonneux, Sophie, et al.
Pubblicazione: (2024)
di: Xhonneux, Sophie, et al.
Pubblicazione: (2024)
Efficient Adversarial Training in LLMs with Continuous Attacks
di: Xhonneux, Sophie, et al.
Pubblicazione: (2024)
di: Xhonneux, Sophie, et al.
Pubblicazione: (2024)
Dimension-adapted Momentum Outscales SGD
di: Ferbach, Damien, et al.
Pubblicazione: (2025)
di: Ferbach, Damien, et al.
Pubblicazione: (2025)
Logarithmic-time Schedules for Scaling Language Models with Momentum
di: Ferbach, Damien, et al.
Pubblicazione: (2026)
di: Ferbach, Damien, et al.
Pubblicazione: (2026)
Tight Lower Bounds and Improved Convergence in Performative Prediction
di: Khorsandi, Pedram, et al.
Pubblicazione: (2024)
di: Khorsandi, Pedram, et al.
Pubblicazione: (2024)
Asynchronous Algorithmic Alignment with Cocycles
di: Dudzik, Andrew, et al.
Pubblicazione: (2023)
di: Dudzik, Andrew, et al.
Pubblicazione: (2023)
Versatile Energy-Based Probabilistic Models for High Energy Physics
di: Cheng, Taoli, et al.
Pubblicazione: (2023)
di: Cheng, Taoli, et al.
Pubblicazione: (2023)
A Generative Approach to LLM Harmfulness Mitigation with Red Flag Tokens
di: Dobre, David, et al.
Pubblicazione: (2025)
di: Dobre, David, et al.
Pubblicazione: (2025)
Neuroplastic Expansion in Deep Reinforcement Learning
di: Liu, Jiashun, et al.
Pubblicazione: (2024)
di: Liu, Jiashun, et al.
Pubblicazione: (2024)
The Curse of Diversity in Ensemble-Based Exploration
di: Lin, Zhixuan, et al.
Pubblicazione: (2024)
di: Lin, Zhixuan, et al.
Pubblicazione: (2024)
In value-based deep reinforcement learning, a pruned network is a good network
di: Obando-Ceron, Johan, et al.
Pubblicazione: (2024)
di: Obando-Ceron, Johan, et al.
Pubblicazione: (2024)
Compositional Discrete Latent Code for High Fidelity, Productive Diffusion Models
di: Lavoie, Samuel, et al.
Pubblicazione: (2025)
di: Lavoie, Samuel, et al.
Pubblicazione: (2025)
GenRL: Multimodal-foundation world models for generalization in embodied agents
di: Mazzaglia, Pietro, et al.
Pubblicazione: (2024)
di: Mazzaglia, Pietro, et al.
Pubblicazione: (2024)
Self-Consuming Generative Models with Curated Data Provably Optimize Human Preferences
di: Ferbach, Damien, et al.
Pubblicazione: (2024)
di: Ferbach, Damien, et al.
Pubblicazione: (2024)
Solving Hidden Monotone Variational Inequalities with Surrogate Losses
di: D'Orazio, Ryan, et al.
Pubblicazione: (2024)
di: D'Orazio, Ryan, et al.
Pubblicazione: (2024)
Not All LLM Reasoners Are Created Equal
di: Hosseini, Arian, et al.
Pubblicazione: (2024)
di: Hosseini, Arian, et al.
Pubblicazione: (2024)
Synaptic Weight Distributions Depend on the Geometry of Plasticity
di: Pogodin, Roman, et al.
Pubblicazione: (2023)
di: Pogodin, Roman, et al.
Pubblicazione: (2023)
Latent Space Representations of Neural Algorithmic Reasoners
di: Mirjanić, Vladimir V., et al.
Pubblicazione: (2023)
di: Mirjanić, Vladimir V., et al.
Pubblicazione: (2023)
Feature Likelihood Divergence: Evaluating the Generalization of Generative Models Using Samples
di: Jiralerspong, Marco, et al.
Pubblicazione: (2023)
di: Jiralerspong, Marco, et al.
Pubblicazione: (2023)
Performative Prediction on Games and Mechanism Design
di: Góis, António, et al.
Pubblicazione: (2024)
di: Góis, António, et al.
Pubblicazione: (2024)
Explaining and Preventing Alignment Collapse in Iterative RLHF
di: Gauthier, Etienne, et al.
Pubblicazione: (2026)
di: Gauthier, Etienne, et al.
Pubblicazione: (2026)
Scattered Mixture-of-Experts Implementation
di: Tan, Shawn, et al.
Pubblicazione: (2024)
di: Tan, Shawn, et al.
Pubblicazione: (2024)
Documenti analoghi
-
LOQA: Learning with Opponent Q-Learning Awareness
di: Aghajohari, Milad, et al.
Pubblicazione: (2024) -
Best Response Shaping
di: Aghajohari, Milad, et al.
Pubblicazione: (2024) -
Towards Sustainable Investment Policies Informed by Opponent Shaping
di: Duque, Juan Agustin, et al.
Pubblicazione: (2026) -
Learning Robust Social Strategies with Large Language Models
di: Piche, Dereck, et al.
Pubblicazione: (2025) -
The Markovian Thinker: Architecture-Agnostic Linear Scaling of Reasoning
di: Aghajohari, Milad, et al.
Pubblicazione: (2025)