Conditional Language Policy: A General Framework for Steerable Multi-Objective Finetuning
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Kaiwen, Kidambi, Rahul, Sullivan, Ryan, Agarwal, Alekh, Dann, Christoph, Michi, Andrea, Gelmi, Marco, Li, Yunxuan, Gupta, Raghav, Dubey, Avinava, Ramé, Alexandre, Ferret, Johan, Cideron, Geoffrey, Hou, Le, Yu, Hongkun, Ahmed, Amr, Mehta, Aranyak, Hussenot, Léonard, Bachem, Olivier, Leurent, Edouard |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
WARM: On the Benefits of Weight Averaged Reward Models
by: Ramé, Alexandre, et al.
Published: (2024)
by: Ramé, Alexandre, et al.
Published: (2024)
Auctions with LLM Summaries
by: Dubey, Kumar Avinava, et al.
Published: (2024)
by: Dubey, Kumar Avinava, et al.
Published: (2024)
Inference-time Unlearning Using Conformal Prediction
by: Chowdhury, Somnath Basu Roy, et al.
Published: (2026)
by: Chowdhury, Somnath Basu Roy, et al.
Published: (2026)
Diversity-Rewarded CFG Distillation
by: Cideron, Geoffrey, et al.
Published: (2024)
by: Cideron, Geoffrey, et al.
Published: (2024)
BOND: Aligning LLMs with Best-of-N Distillation
by: Sessa, Pier Giuseppe, et al.
Published: (2024)
by: Sessa, Pier Giuseppe, et al.
Published: (2024)
WARP: On the Benefits of Weight Averaged Rewarded Policies
by: Ramé, Alexandre, et al.
Published: (2024)
by: Ramé, Alexandre, et al.
Published: (2024)
A Minimaximalist Approach to Reinforcement Learning from Human Feedback
by: Swamy, Gokul, et al.
Published: (2024)
by: Swamy, Gokul, et al.
Published: (2024)
Fundamental Limits of Perfect Concept Erasure
by: Chowdhury, Somnath Basu Roy, et al.
Published: (2025)
by: Chowdhury, Somnath Basu Roy, et al.
Published: (2025)
Enhancing Group Fairness in Online Settings Using Oblique Decision Forests
by: Chowdhury, Somnath Basu Roy, et al.
Published: (2023)
by: Chowdhury, Somnath Basu Roy, et al.
Published: (2023)
Design Considerations in Offline Preference-based RL
by: Agarwal, Alekh, et al.
Published: (2025)
by: Agarwal, Alekh, et al.
Published: (2025)
MusicRL: Aligning Music Generation to Human Preferences
by: Cideron, Geoffrey, et al.
Published: (2024)
by: Cideron, Geoffrey, et al.
Published: (2024)
Mitigating Preference Hacking in Policy Optimization with Pessimism
by: Gupta, Dhawal, et al.
Published: (2025)
by: Gupta, Dhawal, et al.
Published: (2025)
Estimación de la carga de nitratos en una cuenca rural y su relación con la variabilidad climática
by: Mónica Gelmi
Published: (2012)
by: Mónica Gelmi
Published: (2012)
Linear Transformer Topological Masking with Graph Random Features
by: Reid, Isaac, et al.
Published: (2024)
by: Reid, Isaac, et al.
Published: (2024)
Incentive-Aligned Multi-Source LLM Summaries
by: Jiang, Yanchen, et al.
Published: (2025)
by: Jiang, Yanchen, et al.
Published: (2025)
Efficiency of Non-Truthful Auctions in Auto-bidding with Budget Constraints
by: Liaw, Christopher, et al.
Published: (2023)
by: Liaw, Christopher, et al.
Published: (2023)
Ads in Conversations
by: Banchio, Martino, et al.
Published: (2024)
by: Banchio, Martino, et al.
Published: (2024)
Incentive Compatibility in the Auto-bidding World
by: Alimohammadi, Yeganeh, et al.
Published: (2023)
by: Alimohammadi, Yeganeh, et al.
Published: (2023)
Optimal Time Complexity Algorithms for Computing General Random Walk Graph Kernels on Sparse Graphs
by: Choromanski, Krzysztof, et al.
Published: (2024)
by: Choromanski, Krzysztof, et al.
Published: (2024)
Computationally-efficient Graph Modeling with Refined Graph Random Features
by: Choromanski, Krzysztof, et al.
Published: (2025)
by: Choromanski, Krzysztof, et al.
Published: (2025)
On Teacher Hacking in Language Model Distillation
by: Tiapkin, Daniil, et al.
Published: (2025)
by: Tiapkin, Daniil, et al.
Published: (2025)
Incremental Extractive Opinion Summarization Using Cover Trees
by: Chowdhury, Somnath Basu Roy, et al.
Published: (2024)
by: Chowdhury, Somnath Basu Roy, et al.
Published: (2024)
Scalable Neural Network Kernels
by: Sehanobish, Arijit, et al.
Published: (2023)
by: Sehanobish, Arijit, et al.
Published: (2023)
SWING: Unlocking Implicit Graph Representations for Graph Random Features
by: Manenti, Alessandro, et al.
Published: (2026)
by: Manenti, Alessandro, et al.
Published: (2026)
Scaling Inference-Time Computation via Opponent Simulation: Enabling Online Strategic Adaptation in Repeated Negotiation
by: Liu, Xiangyu, et al.
Published: (2026)
by: Liu, Xiangyu, et al.
Published: (2026)
The Power of Two-sided Recruitment in Two-sided Markets
by: Cai, Yang, et al.
Published: (2023)
by: Cai, Yang, et al.
Published: (2023)
A New Lower Bound for the Random Offerer Mechanism in Bilateral Trade using AI-Guided Evolutionary Search
by: Cai, Yang, et al.
Published: (2026)
by: Cai, Yang, et al.
Published: (2026)
The Memory Engine: Self-Organized Coherence from Internal Feedback
by: Sarkar, Aranyak
Published: (2025)
by: Sarkar, Aranyak
Published: (2025)
A Non-Markovian Route to Coherence in Heterogeneous Diffusive Systems
by: Sarkar, Aranyak
Published: (2025)
by: Sarkar, Aranyak
Published: (2025)
Crise sociale, question nationale et violence urbaine. Retour sur la mystérieuse Kale Borroka en Espagne.
by: Jérôme Ferret
Published: (2012)
by: Jérôme Ferret
Published: (2012)
EUGens: Efficient, Unified, and General Dense Layers
by: Kim, Sang Min, et al.
Published: (2024)
by: Kim, Sang Min, et al.
Published: (2024)
EUGens: Efficient, Unified, and General Dense Layers
by: Kim, Sang Min, et al.
Published: (2026)
by: Kim, Sang Min, et al.
Published: (2026)
Adaptive and Explainable AI Agents for Anomaly Detection in Critical IoT Infrastructure using LLM-Enhanced Contextual Reasoning
by: Sharma, Raghav, et al.
Published: (2025)
by: Sharma, Raghav, et al.
Published: (2025)
Small Language Models for Agentic Systems: A Survey of Architectures, Capabilities, and Deployment Trade offs
by: Sharma, Raghav, et al.
Published: (2025)
by: Sharma, Raghav, et al.
Published: (2025)
Preserving Expert-Level Privacy in Offline Reinforcement Learning
by: Sharma, Navodita, et al.
Published: (2024)
by: Sharma, Navodita, et al.
Published: (2024)
More Benefits of Being Distributional: Second-Order Bounds for Reinforcement Learning
by: Wang, Kaiwen, et al.
Published: (2024)
by: Wang, Kaiwen, et al.
Published: (2024)
RelFlexformer: Efficient Attention 3D-Transformers for Integrable Relative Positional Encodings
by: Kim, Byeongchan, et al.
Published: (2026)
by: Kim, Byeongchan, et al.
Published: (2026)
Hyperspectral Imaging for Detection and Classification of Plant Primary and Secondary Metabolites: A Review
by: Muskan Raghav, et al.
Published: (2025)
by: Muskan Raghav, et al.
Published: (2025)
The interplay of geopolitics and agricultural commodity prices
by: Raghav Goyal, et al.
Published: (2024)
by: Raghav Goyal, et al.
Published: (2024)
EVALUACIÓN DE LA ACCIÓN DE DIFERENTES FITORREGULADORES SOBRE LAS POBLACIONES DE STENEOTARSONEMUS SPINKI SMILEY EN DOS VARIEDADES COMERCIALES DE ARROZ
by: Eleazar Botta Ferret
Published: (2008)
by: Eleazar Botta Ferret
Published: (2008)
Similar Items
-
WARM: On the Benefits of Weight Averaged Reward Models
by: Ramé, Alexandre, et al.
Published: (2024) -
Auctions with LLM Summaries
by: Dubey, Kumar Avinava, et al.
Published: (2024) -
Inference-time Unlearning Using Conformal Prediction
by: Chowdhury, Somnath Basu Roy, et al.
Published: (2026) -
Diversity-Rewarded CFG Distillation
by: Cideron, Geoffrey, et al.
Published: (2024) -
BOND: Aligning LLMs with Best-of-N Distillation
by: Sessa, Pier Giuseppe, et al.
Published: (2024)