Aligning language models with human preferences
Fuente:
arXiv
Saved in:
| Main Author: | Korbak, Tomasz |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Compositional preference models for aligning LMs
by: Go, Dongyoung, et al.
Published: (2023)
by: Go, Dongyoung, et al.
Published: (2023)
Catalytic Role Of Noise And Necessity Of Inductive Biases In The Emergence Of Compositional Communication
by: Kuciński, Łukasz, et al.
Published: (2021)
by: Kuciński, Łukasz, et al.
Published: (2021)
InfAlign: Inference-aware language model alignment
by: Balashankar, Ananth, et al.
Published: (2024)
by: Balashankar, Ananth, et al.
Published: (2024)
Large language models can accurately predict searcher preferences
by: Thomas, Paul, et al.
Published: (2023)
by: Thomas, Paul, et al.
Published: (2023)
The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
by: Berglund, Lukas, et al.
Published: (2023)
by: Berglund, Lukas, et al.
Published: (2023)
Training Language Models with Language Feedback at Scale
by: Scheurer, Jérémy, et al.
Published: (2023)
by: Scheurer, Jérémy, et al.
Published: (2023)
Post-training makes large language models less human-like
by: Binz, Marcel, et al.
Published: (2026)
by: Binz, Marcel, et al.
Published: (2026)
A density estimation perspective on learning from pairwise human preferences
by: Dumoulin, Vincent, et al.
Published: (2023)
by: Dumoulin, Vincent, et al.
Published: (2023)
Ensemble Kalman filter for uncertainty in human language comprehension
by: Bhandari, Diksha, et al.
Published: (2025)
by: Bhandari, Diksha, et al.
Published: (2025)
Evaluating language models as risk scores
by: Cruz, André F., et al.
Published: (2024)
by: Cruz, André F., et al.
Published: (2024)
PathAlign: A vision-language model for whole slide images in histopathology
by: Ahmed, Faruk, et al.
Published: (2024)
by: Ahmed, Faruk, et al.
Published: (2024)
Amortizing intractable inference in large language models
by: Hu, Edward J., et al.
Published: (2023)
by: Hu, Edward J., et al.
Published: (2023)
Do language models plan ahead for future tokens?
by: Wu, Wilson, et al.
Published: (2024)
by: Wu, Wilson, et al.
Published: (2024)
Perturbed examples reveal invariances shared by language models
by: Rawal, Ruchit, et al.
Published: (2023)
by: Rawal, Ruchit, et al.
Published: (2023)
A mean teacher algorithm for unlearning of language models
by: Klochkov, Yegor
Published: (2025)
by: Klochkov, Yegor
Published: (2025)
Visualizing token importance for black-box language models
by: Rauba, Paulius, et al.
Published: (2025)
by: Rauba, Paulius, et al.
Published: (2025)
AlignSAE: Concept-Aligned Sparse Autoencoders
by: Yang, Minglai, et al.
Published: (2025)
by: Yang, Minglai, et al.
Published: (2025)
LIRE: listwise reward enhancement for preference alignment
by: Zhu, Mingye, et al.
Published: (2024)
by: Zhu, Mingye, et al.
Published: (2024)
Anatomical Heterogeneity in Transformer Language Models
by: Wietrzykowski, Tomasz
Published: (2026)
by: Wietrzykowski, Tomasz
Published: (2026)
The language of time: a language model perspective on time-series foundation models
by: Xie, Yi, et al.
Published: (2025)
by: Xie, Yi, et al.
Published: (2025)
Prompt reinforcing for long-term planning of large language models
by: Lin, Hsien-Chin, et al.
Published: (2025)
by: Lin, Hsien-Chin, et al.
Published: (2025)
Machine-generated text detection prevents language model collapse
by: Drayson, George, et al.
Published: (2025)
by: Drayson, George, et al.
Published: (2025)
The Consciousness Cluster: Emergent preferences of Models that Claim to be Conscious
by: Chua, James, et al.
Published: (2026)
by: Chua, James, et al.
Published: (2026)
Can large language models replace humans in the systematic review process? Evaluating GPT-4's efficacy in screening and extracting data from peer-reviewed and grey literature in multiple languages
by: Khraisha, Qusai, et al.
Published: (2023)
by: Khraisha, Qusai, et al.
Published: (2023)
Improving Code Generation by Training with Natural Language Feedback
by: Chen, Angelica, et al.
Published: (2023)
by: Chen, Angelica, et al.
Published: (2023)
Mixture-of-Depths: Dynamically allocating compute in transformer-based language models
by: Raposo, David, et al.
Published: (2024)
by: Raposo, David, et al.
Published: (2024)
Anchor function: a type of benchmark functions for studying language models
by: Zhang, Zhongwang, et al.
Published: (2024)
by: Zhang, Zhongwang, et al.
Published: (2024)
Simple linear attention language models balance the recall-throughput tradeoff
by: Arora, Simran, et al.
Published: (2024)
by: Arora, Simran, et al.
Published: (2024)
Just read twice: closing the recall gap for recurrent language models
by: Arora, Simran, et al.
Published: (2024)
by: Arora, Simran, et al.
Published: (2024)
Zero-shot generation of synthetic neurosurgical data with large language models
by: Barr, Austin A., et al.
Published: (2025)
by: Barr, Austin A., et al.
Published: (2025)
Repetitions are not all alike: distinct mechanisms sustain repetition in language models
by: Mahaut, Matéo, et al.
Published: (2025)
by: Mahaut, Matéo, et al.
Published: (2025)
Only relative ranks matter in weight-clustered large language models
by: Aizpurua, Borja, et al.
Published: (2026)
by: Aizpurua, Borja, et al.
Published: (2026)
How do language models learn facts? Dynamics, curricula and hallucinations
by: Zucchet, Nicolas, et al.
Published: (2025)
by: Zucchet, Nicolas, et al.
Published: (2025)
Representation in large language models
by: Yetman, Cameron
Published: (2025)
by: Yetman, Cameron
Published: (2025)
Inference time LLM alignment in single and multidomain preference spectrum
by: Shahriar, Sadat, et al.
Published: (2024)
by: Shahriar, Sadat, et al.
Published: (2024)
Question answering system of bridge design specification based on large language model
by: Zhang, Leye, et al.
Published: (2024)
by: Zhang, Leye, et al.
Published: (2024)
DataComp-LM: In search of the next generation of training sets for language models
by: Li, Jeffrey, et al.
Published: (2024)
by: Li, Jeffrey, et al.
Published: (2024)
The representation landscape of few-shot learning and fine-tuning in large language models
by: Doimo, Diego, et al.
Published: (2024)
by: Doimo, Diego, et al.
Published: (2024)
Linear representations in language models can change dramatically over a conversation
by: Lampinen, Andrew Kyle, et al.
Published: (2026)
by: Lampinen, Andrew Kyle, et al.
Published: (2026)
B-score: Detecting biases in large language models using response history
by: Vo, An, et al.
Published: (2025)
by: Vo, An, et al.
Published: (2025)
Similar Items
-
Compositional preference models for aligning LMs
by: Go, Dongyoung, et al.
Published: (2023) -
Catalytic Role Of Noise And Necessity Of Inductive Biases In The Emergence Of Compositional Communication
by: Kuciński, Łukasz, et al.
Published: (2021) -
InfAlign: Inference-aware language model alignment
by: Balashankar, Ananth, et al.
Published: (2024) -
Large language models can accurately predict searcher preferences
by: Thomas, Paul, et al.
Published: (2023) -
The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
by: Berglund, Lukas, et al.
Published: (2023)