AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback
Fuente:
arXiv
Saved in:
| Main Authors: | Dubois, Yann, Li, Xuechen, Taori, Rohan, Zhang, Tianyi, Gulrajani, Ishaan, Ba, Jimmy, Guestrin, Carlos, Liang, Percy, Hashimoto, Tatsunori B. |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
by: Dubois, Yann, et al.
Published: (2024)
by: Dubois, Yann, et al.
Published: (2024)
Evaluating Self-Supervised Learning via Risk Decomposition
by: Dubois, Yann, et al.
Published: (2023)
by: Dubois, Yann, et al.
Published: (2023)
Benchmarking Distributional Alignment of Large Language Models
by: Meister, Nicole, et al.
Published: (2024)
by: Meister, Nicole, et al.
Published: (2024)
The Extractive-Abstractive Spectrum: Uncovering Verifiability Trade-offs in LLM Generations
by: Worledge, Theodora, et al.
Published: (2024)
by: Worledge, Theodora, et al.
Published: (2024)
Model Equality Testing: Which Model Is This API Serving?
by: Gao, Irena, et al.
Published: (2024)
by: Gao, Irena, et al.
Published: (2024)
Identifying the Risks of LM Agents with an LM-Emulated Sandbox
by: Ruan, Yangjun, et al.
Published: (2023)
by: Ruan, Yangjun, et al.
Published: (2023)
Robust Distortion-free Watermarks for Language Models
by: Kuditipudi, Rohith, et al.
Published: (2023)
by: Kuditipudi, Rohith, et al.
Published: (2023)
Pre-training under infinite compute
by: Kim, Konwoo, et al.
Published: (2025)
by: Kim, Konwoo, et al.
Published: (2025)
Linguistic Calibration of Long-Form Generations
by: Band, Neil, et al.
Published: (2024)
by: Band, Neil, et al.
Published: (2024)
On the Learnability of Watermarks for Language Models
by: Gu, Chenchen, et al.
Published: (2023)
by: Gu, Chenchen, et al.
Published: (2023)
Learning to (Learn at Test Time): RNNs with Expressive Hidden States
by: Sun, Yu, et al.
Published: (2024)
by: Sun, Yu, et al.
Published: (2024)
Auditing Prompt Caching in Language Model APIs
by: Gu, Chenchen, et al.
Published: (2025)
by: Gu, Chenchen, et al.
Published: (2025)
Out-of-Domain Robustness via Targeted Augmentations
by: Gao, Irena, et al.
Published: (2023)
by: Gao, Irena, et al.
Published: (2023)
Data-efficient pre-training by scaling synthetic megadocs
by: Kim, Konwoo, et al.
Published: (2026)
by: Kim, Konwoo, et al.
Published: (2026)
The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas
by: Si, Chenglei, et al.
Published: (2025)
by: Si, Chenglei, et al.
Published: (2025)
Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
by: Si, Chenglei, et al.
Published: (2024)
by: Si, Chenglei, et al.
Published: (2024)
Language Models with Conformal Factuality Guarantees
by: Mohri, Christopher, et al.
Published: (2024)
by: Mohri, Christopher, et al.
Published: (2024)
Learning to (Learn at Test Time)
by: Sun, Yu, et al.
Published: (2023)
by: Sun, Yu, et al.
Published: (2023)
AutoBencher: Towards Declarative Benchmark Construction
by: Li, Xiang Lisa, et al.
Published: (2024)
by: Li, Xiang Lisa, et al.
Published: (2024)
Frontières et mobilité au quotidien
by: Dubois, Yann
Published: (2020)
by: Dubois, Yann
Published: (2020)
Eliciting Language Model Behaviors with Investigator Agents
by: Li, Xiang Lisa, et al.
Published: (2025)
by: Li, Xiang Lisa, et al.
Published: (2025)
Understanding Finetuning for Factual Knowledge Extraction
by: Ghosal, Gaurav, et al.
Published: (2024)
by: Ghosal, Gaurav, et al.
Published: (2024)
Improving Pretraining Data Using Perplexity Correlations
by: Thrush, Tristan, et al.
Published: (2024)
by: Thrush, Tristan, et al.
Published: (2024)
A Bitter Lesson for Data Filtering
by: Mohri, Christopher, et al.
Published: (2026)
by: Mohri, Christopher, et al.
Published: (2026)
Removing RLHF Protections in GPT-4 via Fine-Tuning
by: Zhan, Qiusi, et al.
Published: (2023)
by: Zhan, Qiusi, et al.
Published: (2023)
Doubly Optimal No-Regret Online Learning in Strongly Monotone Games with Bandit Feedback
by: Ba, Wenjia, et al.
Published: (2021)
by: Ba, Wenjia, et al.
Published: (2021)
EMA Policy Gradient: Taming Reinforcement Learning for LLMs with EMA Anchor and Top-k KL
by: Zhang, Lunjun, et al.
Published: (2026)
by: Zhang, Lunjun, et al.
Published: (2026)
Societal Impacts Research Requires Benchmarks for Creative Composition Tasks
by: Shen, Judy Hanwen, et al.
Published: (2025)
by: Shen, Judy Hanwen, et al.
Published: (2025)
Observational Scaling Laws and the Predictability of Language Model Performance
by: Ruan, Yangjun, et al.
Published: (2024)
by: Ruan, Yangjun, et al.
Published: (2024)
A Collaborative Framework for Quantum Optimisation and Quantum Neural Networks: Credit Feature Selection and Image Classification
by: Long, JiaNing, et al.
Published: (2025)
by: Long, JiaNing, et al.
Published: (2025)
Discovering Implicit Large Language Model Alignment Objectives
by: Chen, Edward, et al.
Published: (2026)
by: Chen, Edward, et al.
Published: (2026)
The Future of Open Human Feedback
by: Don-Yehiya, Shachar, et al.
Published: (2024)
by: Don-Yehiya, Shachar, et al.
Published: (2024)
Scaling Laws for the Value of Individual Data Points in Machine Learning
by: Covert, Ian, et al.
Published: (2024)
by: Covert, Ian, et al.
Published: (2024)
Locality Alignment Improves Vision-Language Models
by: Covert, Ian, et al.
Published: (2024)
by: Covert, Ian, et al.
Published: (2024)
Deep Learning Based Monthly Temperature Prediction for Jilin Province: A Multi Model Comparative Study 2000 2026
by: Deng, Xingyue, et al.
Published: (2026)
by: Deng, Xingyue, et al.
Published: (2026)
Research on Expressway Congestion Warning Technology Based on YOLOv11-DIoU and GRU-Attention
by: Yulin, Tong, et al.
Published: (2025)
by: Yulin, Tong, et al.
Published: (2025)
Quasi‐Static Closed‐Loop Wind‐Farm Control for Combined Power and Fatigue Optimization
by: Ishaan Sood, et al.
Published: (2025)
by: Ishaan Sood, et al.
Published: (2025)
Reasoning to Learn from Latent Thoughts
by: Ruan, Yangjun, et al.
Published: (2025)
by: Ruan, Yangjun, et al.
Published: (2025)
Long-term Safe Reinforcement Learning with Binary Feedback
by: Wachi, Akifumi, et al.
Published: (2024)
by: Wachi, Akifumi, et al.
Published: (2024)
s1: Simple test-time scaling
by: Muennighoff, Niklas, et al.
Published: (2025)
by: Muennighoff, Niklas, et al.
Published: (2025)
Similar Items
-
Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
by: Dubois, Yann, et al.
Published: (2024) -
Evaluating Self-Supervised Learning via Risk Decomposition
by: Dubois, Yann, et al.
Published: (2023) -
Benchmarking Distributional Alignment of Large Language Models
by: Meister, Nicole, et al.
Published: (2024) -
The Extractive-Abstractive Spectrum: Uncovering Verifiability Trade-offs in LLM Generations
by: Worledge, Theodora, et al.
Published: (2024) -
Model Equality Testing: Which Model Is This API Serving?
by: Gao, Irena, et al.
Published: (2024)