Discovering Preference Optimization Algorithms with and for Large Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Lu, Chris, Holt, Samuel, Fanconi, Claudio, Chan, Alex J., Foerster, Jakob, van der Schaar, Mihaela, Lange, Robert Tjarko |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Tiny Autoregressive Recursive Models
por: Rauba, Paulius, et al.
Publicado: (2026)
por: Rauba, Paulius, et al.
Publicado: (2026)
Cascaded Language Models for Cost-effective Human-AI Decision-Making
por: Fanconi, Claudio, et al.
Publicado: (2025)
por: Fanconi, Claudio, et al.
Publicado: (2025)
Discovering Temporally-Aware Reinforcement Learning Algorithms
por: Jackson, Matthew Thomas, et al.
Publicado: (2024)
por: Jackson, Matthew Thomas, et al.
Publicado: (2024)
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes
por: Kobalczyk, Katarzyna, et al.
Publicado: (2024)
por: Kobalczyk, Katarzyna, et al.
Publicado: (2024)
Dense Reward for Free in Reinforcement Learning from Human Feedback
por: Chan, Alex J., et al.
Publicado: (2024)
por: Chan, Alex J., et al.
Publicado: (2024)
Learning Reasoning Rewards from Expert Demonstrations with Inverse Reinforcement Learning
por: Fanconi, Claudio, et al.
Publicado: (2025)
por: Fanconi, Claudio, et al.
Publicado: (2025)
L2MAC: Large Language Model Automatic Computer for Extensive Code Generation
por: Holt, Samuel, et al.
Publicado: (2023)
por: Holt, Samuel, et al.
Publicado: (2023)
Discovering Quality-Diversity Algorithms via Meta-Black-Box Optimization
por: Faldor, Maxence, et al.
Publicado: (2025)
por: Faldor, Maxence, et al.
Publicado: (2025)
Automatically Learning Hybrid Digital Twins of Dynamical Systems
por: Holt, Samuel, et al.
Publicado: (2024)
por: Holt, Samuel, et al.
Publicado: (2024)
Behaviour Distillation
por: Lupu, Andrei, et al.
Publicado: (2024)
por: Lupu, Andrei, et al.
Publicado: (2024)
Deep Hierarchical Learning with Nested Subspace Networks for Large Language Models
por: Rauba, Paulius, et al.
Publicado: (2025)
por: Rauba, Paulius, et al.
Publicado: (2025)
Matchmaker: Self-Improving Large Language Model Programs for Schema Matching
por: Seedat, Nabeel, et al.
Publicado: (2024)
por: Seedat, Nabeel, et al.
Publicado: (2024)
G-Sim: Generative Simulations with Large Language Models and Gradient-Free Calibration
por: Holt, Samuel, et al.
Publicado: (2025)
por: Holt, Samuel, et al.
Publicado: (2025)
Deep Generative Symbolic Regression
por: Holt, Samuel, et al.
Publicado: (2023)
por: Holt, Samuel, et al.
Publicado: (2023)
Inverse-RLignment: Large Language Model Alignment from Demonstrations through Inverse Reinforcement Learning
por: Sun, Hao, et al.
Publicado: (2024)
por: Sun, Hao, et al.
Publicado: (2024)
Large Language Models to Enhance Bayesian Optimization
por: Liu, Tennison, et al.
Publicado: (2024)
por: Liu, Tennison, et al.
Publicado: (2024)
Preference Learning for AI Alignment: a Causal Perspective
por: Kobalczyk, Katarzyna, et al.
Publicado: (2025)
por: Kobalczyk, Katarzyna, et al.
Publicado: (2025)
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities
por: Sun, Hao, et al.
Publicado: (2025)
por: Sun, Hao, et al.
Publicado: (2025)
The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
por: Lu, Chris, et al.
Publicado: (2024)
por: Lu, Chris, et al.
Publicado: (2024)
Language Bottleneck Models for Qualitative Knowledge State Modeling
por: Berthon, Antonin, et al.
Publicado: (2025)
por: Berthon, Antonin, et al.
Publicado: (2025)
Simulating Viva Voce Examinations to Evaluate Clinical Reasoning in Large Language Models
por: Chiu, Christopher, et al.
Publicado: (2025)
por: Chiu, Christopher, et al.
Publicado: (2025)
Discovering Minimal Reinforcement Learning Environments
por: Liesen, Jarek, et al.
Publicado: (2024)
por: Liesen, Jarek, et al.
Publicado: (2024)
Why Tabular Foundation Models Should Be a Research Priority
por: van Breugel, Boris, et al.
Publicado: (2024)
por: van Breugel, Boris, et al.
Publicado: (2024)
Large Language Models As Evolution Strategies
por: Lange, Robert Tjarko, et al.
Publicado: (2024)
por: Lange, Robert Tjarko, et al.
Publicado: (2024)
ODE Discovery for Longitudinal Heterogeneous Treatment Effects Inference
por: Kacprzyk, Krzysztof, et al.
Publicado: (2024)
por: Kacprzyk, Krzysztof, et al.
Publicado: (2024)
On Error Propagation of Diffusion Models
por: Li, Yangming, et al.
Publicado: (2023)
por: Li, Yangming, et al.
Publicado: (2023)
Retrieval Augmented Thought Process for Private Data Handling in Healthcare
por: Pouplin, Thomas, et al.
Publicado: (2024)
por: Pouplin, Thomas, et al.
Publicado: (2024)
Nonparametric LLM Evaluation from Preference Data
por: Frauen, Dennis, et al.
Publicado: (2026)
por: Frauen, Dennis, et al.
Publicado: (2026)
Context-Aware Testing: A New Paradigm for Model Testing with Large Language Models
por: Rauba, Paulius, et al.
Publicado: (2024)
por: Rauba, Paulius, et al.
Publicado: (2024)
Risk-Sensitive Diffusion: Robustly Optimizing Diffusion Models with Noisy Samples
por: Li, Yangming, et al.
Publicado: (2024)
por: Li, Yangming, et al.
Publicado: (2024)
Knowledge-Informed Kernel State Reconstruction from Heterogeneous Partial Observations
por: Muscarnera, Luca, et al.
Publicado: (2026)
por: Muscarnera, Luca, et al.
Publicado: (2026)
Shape Arithmetic Expressions: Advancing Scientific Discovery Beyond Closed-Form Equations
por: Kacprzyk, Krzysztof, et al.
Publicado: (2024)
por: Kacprzyk, Krzysztof, et al.
Publicado: (2024)
No Equations Needed: Learning System Dynamics Without Relying on Closed-Form ODEs
por: Kacprzyk, Krzysztof, et al.
Publicado: (2025)
por: Kacprzyk, Krzysztof, et al.
Publicado: (2025)
Towards Automated Knowledge Integration From Human-Interpretable Representations
por: Kobalczyk, Katarzyna, et al.
Publicado: (2024)
por: Kobalczyk, Katarzyna, et al.
Publicado: (2024)
Autoformulation of Mathematical Optimization Models Using LLMs
por: Astorga, Nicolás, et al.
Publicado: (2024)
por: Astorga, Nicolás, et al.
Publicado: (2024)
The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search
por: Yamada, Yutaro, et al.
Publicado: (2025)
por: Yamada, Yutaro, et al.
Publicado: (2025)
LaTable: Towards Large Tabular Models
por: van Breugel, Boris, et al.
Publicado: (2024)
por: van Breugel, Boris, et al.
Publicado: (2024)
Not All Explanations for Deep Learning Phenomena Are Equally Valuable
por: Jeffares, Alan, et al.
Publicado: (2025)
por: Jeffares, Alan, et al.
Publicado: (2025)
Active Timepoint Selection for Learning Measure-Valued Trajectories
por: Huynh, Nicolas, et al.
Publicado: (2026)
por: Huynh, Nicolas, et al.
Publicado: (2026)
Hyperparameter Trajectory Inference with Conditional Lagrangian Optimal Transport
por: Amad, Harry, et al.
Publicado: (2026)
por: Amad, Harry, et al.
Publicado: (2026)
Ejemplares similares
-
Tiny Autoregressive Recursive Models
por: Rauba, Paulius, et al.
Publicado: (2026) -
Cascaded Language Models for Cost-effective Human-AI Decision-Making
por: Fanconi, Claudio, et al.
Publicado: (2025) -
Discovering Temporally-Aware Reinforcement Learning Algorithms
por: Jackson, Matthew Thomas, et al.
Publicado: (2024) -
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes
por: Kobalczyk, Katarzyna, et al.
Publicado: (2024) -
Dense Reward for Free in Reinforcement Learning from Human Feedback
por: Chan, Alex J., et al.
Publicado: (2024)