Discovering Preference Optimization Algorithms with and for Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lu, Chris, Holt, Samuel, Fanconi, Claudio, Chan, Alex J., Foerster, Jakob, van der Schaar, Mihaela, Lange, Robert Tjarko |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Tiny Autoregressive Recursive Models
von: Rauba, Paulius, et al.
Veröffentlicht: (2026)
von: Rauba, Paulius, et al.
Veröffentlicht: (2026)
Cascaded Language Models for Cost-effective Human-AI Decision-Making
von: Fanconi, Claudio, et al.
Veröffentlicht: (2025)
von: Fanconi, Claudio, et al.
Veröffentlicht: (2025)
Discovering Temporally-Aware Reinforcement Learning Algorithms
von: Jackson, Matthew Thomas, et al.
Veröffentlicht: (2024)
von: Jackson, Matthew Thomas, et al.
Veröffentlicht: (2024)
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes
von: Kobalczyk, Katarzyna, et al.
Veröffentlicht: (2024)
von: Kobalczyk, Katarzyna, et al.
Veröffentlicht: (2024)
Dense Reward for Free in Reinforcement Learning from Human Feedback
von: Chan, Alex J., et al.
Veröffentlicht: (2024)
von: Chan, Alex J., et al.
Veröffentlicht: (2024)
Learning Reasoning Rewards from Expert Demonstrations with Inverse Reinforcement Learning
von: Fanconi, Claudio, et al.
Veröffentlicht: (2025)
von: Fanconi, Claudio, et al.
Veröffentlicht: (2025)
L2MAC: Large Language Model Automatic Computer for Extensive Code Generation
von: Holt, Samuel, et al.
Veröffentlicht: (2023)
von: Holt, Samuel, et al.
Veröffentlicht: (2023)
Discovering Quality-Diversity Algorithms via Meta-Black-Box Optimization
von: Faldor, Maxence, et al.
Veröffentlicht: (2025)
von: Faldor, Maxence, et al.
Veröffentlicht: (2025)
Automatically Learning Hybrid Digital Twins of Dynamical Systems
von: Holt, Samuel, et al.
Veröffentlicht: (2024)
von: Holt, Samuel, et al.
Veröffentlicht: (2024)
Behaviour Distillation
von: Lupu, Andrei, et al.
Veröffentlicht: (2024)
von: Lupu, Andrei, et al.
Veröffentlicht: (2024)
Deep Hierarchical Learning with Nested Subspace Networks for Large Language Models
von: Rauba, Paulius, et al.
Veröffentlicht: (2025)
von: Rauba, Paulius, et al.
Veröffentlicht: (2025)
Matchmaker: Self-Improving Large Language Model Programs for Schema Matching
von: Seedat, Nabeel, et al.
Veröffentlicht: (2024)
von: Seedat, Nabeel, et al.
Veröffentlicht: (2024)
G-Sim: Generative Simulations with Large Language Models and Gradient-Free Calibration
von: Holt, Samuel, et al.
Veröffentlicht: (2025)
von: Holt, Samuel, et al.
Veröffentlicht: (2025)
Deep Generative Symbolic Regression
von: Holt, Samuel, et al.
Veröffentlicht: (2023)
von: Holt, Samuel, et al.
Veröffentlicht: (2023)
Inverse-RLignment: Large Language Model Alignment from Demonstrations through Inverse Reinforcement Learning
von: Sun, Hao, et al.
Veröffentlicht: (2024)
von: Sun, Hao, et al.
Veröffentlicht: (2024)
Large Language Models to Enhance Bayesian Optimization
von: Liu, Tennison, et al.
Veröffentlicht: (2024)
von: Liu, Tennison, et al.
Veröffentlicht: (2024)
Preference Learning for AI Alignment: a Causal Perspective
von: Kobalczyk, Katarzyna, et al.
Veröffentlicht: (2025)
von: Kobalczyk, Katarzyna, et al.
Veröffentlicht: (2025)
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities
von: Sun, Hao, et al.
Veröffentlicht: (2025)
von: Sun, Hao, et al.
Veröffentlicht: (2025)
The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
von: Lu, Chris, et al.
Veröffentlicht: (2024)
von: Lu, Chris, et al.
Veröffentlicht: (2024)
Language Bottleneck Models for Qualitative Knowledge State Modeling
von: Berthon, Antonin, et al.
Veröffentlicht: (2025)
von: Berthon, Antonin, et al.
Veröffentlicht: (2025)
Simulating Viva Voce Examinations to Evaluate Clinical Reasoning in Large Language Models
von: Chiu, Christopher, et al.
Veröffentlicht: (2025)
von: Chiu, Christopher, et al.
Veröffentlicht: (2025)
Discovering Minimal Reinforcement Learning Environments
von: Liesen, Jarek, et al.
Veröffentlicht: (2024)
von: Liesen, Jarek, et al.
Veröffentlicht: (2024)
Why Tabular Foundation Models Should Be a Research Priority
von: van Breugel, Boris, et al.
Veröffentlicht: (2024)
von: van Breugel, Boris, et al.
Veröffentlicht: (2024)
Large Language Models As Evolution Strategies
von: Lange, Robert Tjarko, et al.
Veröffentlicht: (2024)
von: Lange, Robert Tjarko, et al.
Veröffentlicht: (2024)
ODE Discovery for Longitudinal Heterogeneous Treatment Effects Inference
von: Kacprzyk, Krzysztof, et al.
Veröffentlicht: (2024)
von: Kacprzyk, Krzysztof, et al.
Veröffentlicht: (2024)
On Error Propagation of Diffusion Models
von: Li, Yangming, et al.
Veröffentlicht: (2023)
von: Li, Yangming, et al.
Veröffentlicht: (2023)
Retrieval Augmented Thought Process for Private Data Handling in Healthcare
von: Pouplin, Thomas, et al.
Veröffentlicht: (2024)
von: Pouplin, Thomas, et al.
Veröffentlicht: (2024)
Nonparametric LLM Evaluation from Preference Data
von: Frauen, Dennis, et al.
Veröffentlicht: (2026)
von: Frauen, Dennis, et al.
Veröffentlicht: (2026)
Context-Aware Testing: A New Paradigm for Model Testing with Large Language Models
von: Rauba, Paulius, et al.
Veröffentlicht: (2024)
von: Rauba, Paulius, et al.
Veröffentlicht: (2024)
Risk-Sensitive Diffusion: Robustly Optimizing Diffusion Models with Noisy Samples
von: Li, Yangming, et al.
Veröffentlicht: (2024)
von: Li, Yangming, et al.
Veröffentlicht: (2024)
Knowledge-Informed Kernel State Reconstruction from Heterogeneous Partial Observations
von: Muscarnera, Luca, et al.
Veröffentlicht: (2026)
von: Muscarnera, Luca, et al.
Veröffentlicht: (2026)
Shape Arithmetic Expressions: Advancing Scientific Discovery Beyond Closed-Form Equations
von: Kacprzyk, Krzysztof, et al.
Veröffentlicht: (2024)
von: Kacprzyk, Krzysztof, et al.
Veröffentlicht: (2024)
No Equations Needed: Learning System Dynamics Without Relying on Closed-Form ODEs
von: Kacprzyk, Krzysztof, et al.
Veröffentlicht: (2025)
von: Kacprzyk, Krzysztof, et al.
Veröffentlicht: (2025)
Towards Automated Knowledge Integration From Human-Interpretable Representations
von: Kobalczyk, Katarzyna, et al.
Veröffentlicht: (2024)
von: Kobalczyk, Katarzyna, et al.
Veröffentlicht: (2024)
Autoformulation of Mathematical Optimization Models Using LLMs
von: Astorga, Nicolás, et al.
Veröffentlicht: (2024)
von: Astorga, Nicolás, et al.
Veröffentlicht: (2024)
The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search
von: Yamada, Yutaro, et al.
Veröffentlicht: (2025)
von: Yamada, Yutaro, et al.
Veröffentlicht: (2025)
LaTable: Towards Large Tabular Models
von: van Breugel, Boris, et al.
Veröffentlicht: (2024)
von: van Breugel, Boris, et al.
Veröffentlicht: (2024)
Not All Explanations for Deep Learning Phenomena Are Equally Valuable
von: Jeffares, Alan, et al.
Veröffentlicht: (2025)
von: Jeffares, Alan, et al.
Veröffentlicht: (2025)
Active Timepoint Selection for Learning Measure-Valued Trajectories
von: Huynh, Nicolas, et al.
Veröffentlicht: (2026)
von: Huynh, Nicolas, et al.
Veröffentlicht: (2026)
Hyperparameter Trajectory Inference with Conditional Lagrangian Optimal Transport
von: Amad, Harry, et al.
Veröffentlicht: (2026)
von: Amad, Harry, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Tiny Autoregressive Recursive Models
von: Rauba, Paulius, et al.
Veröffentlicht: (2026) -
Cascaded Language Models for Cost-effective Human-AI Decision-Making
von: Fanconi, Claudio, et al.
Veröffentlicht: (2025) -
Discovering Temporally-Aware Reinforcement Learning Algorithms
von: Jackson, Matthew Thomas, et al.
Veröffentlicht: (2024) -
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes
von: Kobalczyk, Katarzyna, et al.
Veröffentlicht: (2024) -
Dense Reward for Free in Reinforcement Learning from Human Feedback
von: Chan, Alex J., et al.
Veröffentlicht: (2024)