Skip to content
VuFind
  • Login
    • English
    • Deutsch
    • Español
    • Français
    • Italiano
Advanced
  • Cite this
  • Text this
  • Email this
  • Print
  • Export Record
    • Export to RefWorks
    • Export to EndNoteWeb
    • Export to EndNote
  • Save to List
  • Permanent link
Cover Image

Saved in:
Bibliographic Details
Main Authors: Richemond, Pierre Harvey, Tang, Yunhao, Guo, Daniel, Calandriello, Daniele, Azar, Mohammad Gheshlaghi, Rafailov, Rafael, Pires, Bernardo Avila, Tarassov, Eugene, Spangher, Lucas, Ellsworth, Will, Severyn, Aliaksei, Mallinson, Jonathan, Shani, Lior, Shamir, Gil, Joshi, Rishabh, Liu, Tianqi, Munos, Remi, Piot, Bilal
Format: Preprint
Published: 2024
Subjects:
Machine Learning
Artificial Intelligence
Online Access:https://arxiv.org/abs/2405.19107
Tags: Add Tag
No Tags, Be the first to tag this record!
  • Holdings
  • Description
  • Table of Contents
  • Comments
  • Similar Items
  • Staff View

Internet

https://arxiv.org/abs/2405.19107

Similar Items

  • Generalized Preference Optimization: A Unified Approach to Offline Alignment
    by: Tang, Yunhao, et al.
    Published: (2024)
  • Human Alignment of Large Language Models through Online Preference Optimisation
    by: Calandriello, Daniele, et al.
    Published: (2024)
  • Understanding the performance gap between online and offline alignment algorithms
    by: Tang, Yunhao, et al.
    Published: (2024)
  • Multi-turn Reinforcement Learning from Preference Human Feedback
    by: Shani, Lior, et al.
    Published: (2024)
  • West-of-N: Synthetic Preferences for Self-Improving Reward Models
    by: Pace, Alizée, et al.
    Published: (2024)

Search Options

  • Search History
  • Advanced Search

Find More

  • Browse the Catalog
  • Browse Alphabetically
  • Explore Channels
  • Course Reserves
  • New Items

Need Help?

  • Search Tips
  • Ask a Librarian
  • FAQs