Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Ahmadian, Arash, Cremer, Chris, Gallé, Matthias, Fadaee, Marzieh, Kreutzer, Julia, Pietquin, Olivier, Üstün, Ahmet, Hooker, Sara |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
RLHF Can Speak Many Languages: Unlocking Multilingual Preference Optimization for LLMs
por: Dang, John, et al.
Publicado: (2024)
por: Dang, John, et al.
Publicado: (2024)
The Multilingual Alignment Prism: Aligning Global and Local Preferences to Reduce Harm
por: Aakanksha, et al.
Publicado: (2024)
por: Aakanksha, et al.
Publicado: (2024)
SimMerge: Learning to Select Merge Operators from Similarity Signals
por: Bolton, Oliver, et al.
Publicado: (2026)
por: Bolton, Oliver, et al.
Publicado: (2026)
LLM See, LLM Do: Guiding Data Generation to Target Non-Differentiable Objectives
por: Shimabucoro, Luísa, et al.
Publicado: (2024)
por: Shimabucoro, Luísa, et al.
Publicado: (2024)
Mix Data or Merge Models? Optimizing for Diverse Multi-Task Learning
por: Aakanksha, et al.
Publicado: (2024)
por: Aakanksha, et al.
Publicado: (2024)
Verification Limits Code LLM Training
por: Gureja, Srishti, et al.
Publicado: (2025)
por: Gureja, Srishti, et al.
Publicado: (2025)
Treasure Hunt: Real-time Targeting of the Long Tail using Training-Time Markers
por: D'souza, Daniel, et al.
Publicado: (2025)
por: D'souza, Daniel, et al.
Publicado: (2025)
The Art of Asking: Multilingual Prompt Optimization for Synthetic Data
por: Mora, David, et al.
Publicado: (2025)
por: Mora, David, et al.
Publicado: (2025)
A Post-trainer's Guide to Multilingual Training Data: Uncovering Cross-lingual Transfer Dynamics
por: Shimabucoro, Luisa, et al.
Publicado: (2025)
por: Shimabucoro, Luisa, et al.
Publicado: (2025)
Diversify and Conquer: Diversity-Centric Data Selection with Iterative Refinement
por: Yu, Simon, et al.
Publicado: (2024)
por: Yu, Simon, et al.
Publicado: (2024)
The Multilingual Divide and Its Impact on Global AI Safety
por: Peppin, Aidan, et al.
Publicado: (2025)
por: Peppin, Aidan, et al.
Publicado: (2025)
Agents Explore but Agents Ignore: LLMs Lack Environmental Curiosity
por: Engländer, Leon, et al.
Publicado: (2026)
por: Engländer, Leon, et al.
Publicado: (2026)
To Code, or Not To Code? Exploring Impact of Code in Pre-training
por: Aryabumi, Viraat, et al.
Publicado: (2024)
por: Aryabumi, Viraat, et al.
Publicado: (2024)
Making, not Taking, the Best of N
por: Khairi, Ammar, et al.
Publicado: (2025)
por: Khairi, Ammar, et al.
Publicado: (2025)
How Well Do LLMs Imitate Human Writing Style?
por: Jemama, Rebira, et al.
Publicado: (2025)
por: Jemama, Rebira, et al.
Publicado: (2025)
Fine-Tuning LLMs with Fine-Grained Human Feedback on Text Spans
por: CH-Wang, Sky, et al.
Publicado: (2025)
por: CH-Wang, Sky, et al.
Publicado: (2025)
Déjà Vu: Multilingual LLM Evaluation through the Lens of Machine Translation Evaluation
por: Kreutzer, Julia, et al.
Publicado: (2025)
por: Kreutzer, Julia, et al.
Publicado: (2025)
If You Can't Use Them, Recycle Them: Optimizing Merging at Scale Mitigates Performance Tradeoffs
por: Khalifa, Muhammad, et al.
Publicado: (2024)
por: Khalifa, Muhammad, et al.
Publicado: (2024)
Human-Aligned Enhancement of Programming Answers with LLMs Guided by User Feedback
por: Bappon, Suborno Deb, et al.
Publicado: (2026)
por: Bappon, Suborno Deb, et al.
Publicado: (2026)
Automatic WordNet Construction Using Markov Chain Monte Carlo
por: Marzieh Fadaee
Publicado: (2013)
por: Marzieh Fadaee
Publicado: (2013)
One Tokenizer To Rule Them All: Emergent Language Plasticity via Multilingual Tokenizers
por: Abagyan, Diana, et al.
Publicado: (2025)
por: Abagyan, Diana, et al.
Publicado: (2025)
Contrastive Policy Gradient: Aligning LLMs on sequence-level scores in a supervised-friendly fashion
por: Flet-Berliac, Yannis, et al.
Publicado: (2024)
por: Flet-Berliac, Yannis, et al.
Publicado: (2024)
The State of Multilingual LLM Safety Research: From Measuring the Language Gap to Mitigating It
por: Yong, Zheng-Xin, et al.
Publicado: (2025)
por: Yong, Zheng-Xin, et al.
Publicado: (2025)
Please Make it Sound like Human: Encoder-Decoder vs. Decoder-Only Transformers for AI-to-Human Text Style Transfer
por: Paneru, Utsav
Publicado: (2026)
por: Paneru, Utsav
Publicado: (2026)
How Does Quantization Affect Multilingual LLMs?
por: Marchisio, Kelly, et al.
Publicado: (2024)
por: Marchisio, Kelly, et al.
Publicado: (2024)
A Framework for Fine-Tuning LLMs using Heterogeneous Feedback
por: Aponte, Ryan, et al.
Publicado: (2024)
por: Aponte, Ryan, et al.
Publicado: (2024)
The Disparate Impacts of Speculative Decoding
por: Sandler, Jameson, et al.
Publicado: (2025)
por: Sandler, Jameson, et al.
Publicado: (2025)
Nexus: Specialization meets Adaptability for Efficiently Training Mixture of Experts
por: Gritsch, Nikolas, et al.
Publicado: (2024)
por: Gritsch, Nikolas, et al.
Publicado: (2024)
When Life Gives You Samples: The Benefits of Scaling up Inference Compute for Multilingual LLMs
por: Khairi, Ammar, et al.
Publicado: (2025)
por: Khairi, Ammar, et al.
Publicado: (2025)
LLMs and the Human Condition
por: Wallis, Peter
Publicado: (2024)
por: Wallis, Peter
Publicado: (2024)
Reinforcement Learning for Latent-Space Thinking in LLMs
por: Özeren, Enes, et al.
Publicado: (2025)
por: Özeren, Enes, et al.
Publicado: (2025)
Self-Improving Robust Preference Optimization
por: Choi, Eugene, et al.
Publicado: (2024)
por: Choi, Eugene, et al.
Publicado: (2024)
MASH: Evading Black-Box AI-Generated Text Detectors via Style Humanization
por: Gu, Yongtong, et al.
Publicado: (2026)
por: Gu, Yongtong, et al.
Publicado: (2026)
Averaging log-likelihoods in direct alignment
por: Grinsztajn, Nathan, et al.
Publicado: (2024)
por: Grinsztajn, Nathan, et al.
Publicado: (2024)
Aya Vision: Advancing the Frontier of Multilingual Multimodality
por: Dash, Saurabh, et al.
Publicado: (2025)
por: Dash, Saurabh, et al.
Publicado: (2025)
Do Language Models Mirror Human Confidence? Exploring Psychological Insights to Address Overconfidence in LLMs
por: Xu, Chenjun, et al.
Publicado: (2025)
por: Xu, Chenjun, et al.
Publicado: (2025)
Toward Beginner-Friendly LLMs for Language Learning: Controlling Difficulty in Conversation
por: Jin, Meiqing, et al.
Publicado: (2025)
por: Jin, Meiqing, et al.
Publicado: (2025)
Revisiting Word Embeddings in the LLM Era
por: Mahajan, Yash, et al.
Publicado: (2024)
por: Mahajan, Yash, et al.
Publicado: (2024)
Diatom assemblage in surface sediments of the Laptev Sea
por: Cremer, Holger, et al.
Publicado: (1998)
por: Cremer, Holger, et al.
Publicado: (1998)
Personality, Role, and Expressive Style in Large Language Models: An Interactionist Analysis
por: Nagao, Moe, et al.
Publicado: (2026)
por: Nagao, Moe, et al.
Publicado: (2026)
Ejemplares similares
-
RLHF Can Speak Many Languages: Unlocking Multilingual Preference Optimization for LLMs
por: Dang, John, et al.
Publicado: (2024) -
The Multilingual Alignment Prism: Aligning Global and Local Preferences to Reduce Harm
por: Aakanksha, et al.
Publicado: (2024) -
SimMerge: Learning to Select Merge Operators from Similarity Signals
por: Bolton, Oliver, et al.
Publicado: (2026) -
LLM See, LLM Do: Guiding Data Generation to Target Non-Differentiable Objectives
por: Shimabucoro, Luísa, et al.
Publicado: (2024) -
Mix Data or Merge Models? Optimizing for Diverse Multi-Task Learning
por: Aakanksha, et al.
Publicado: (2024)