Preference Learning for AI Alignment: a Causal Perspective
Fuente:
arXiv
Salvato in:
| Autori principali: | Kobalczyk, Katarzyna, van der Schaar, Mihaela |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Discovery of Hidden Miscalibration Regimes
di: Kobalczyk, Katarzyna, et al.
Pubblicazione: (2026)
di: Kobalczyk, Katarzyna, et al.
Pubblicazione: (2026)
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes
di: Kobalczyk, Katarzyna, et al.
Pubblicazione: (2024)
di: Kobalczyk, Katarzyna, et al.
Pubblicazione: (2024)
Eliciting Numerical Predictive Distributions of LLMs Without Autoregression
di: Piskorz, Julianna, et al.
Pubblicazione: (2026)
di: Piskorz, Julianna, et al.
Pubblicazione: (2026)
Active Task Disambiguation with LLMs
di: Kobalczyk, Katarzyna, et al.
Pubblicazione: (2025)
di: Kobalczyk, Katarzyna, et al.
Pubblicazione: (2025)
Towards Automated Knowledge Integration From Human-Interpretable Representations
di: Kobalczyk, Katarzyna, et al.
Pubblicazione: (2024)
di: Kobalczyk, Katarzyna, et al.
Pubblicazione: (2024)
Inverse-RLignment: Large Language Model Alignment from Demonstrations through Inverse Reinforcement Learning
di: Sun, Hao, et al.
Pubblicazione: (2024)
di: Sun, Hao, et al.
Pubblicazione: (2024)
Causal Deep Learning
di: Berrevoets, Jeroen, et al.
Pubblicazione: (2023)
di: Berrevoets, Jeroen, et al.
Pubblicazione: (2023)
Not All Explanations for Deep Learning Phenomena Are Equally Valuable
di: Jeffares, Alan, et al.
Pubblicazione: (2025)
di: Jeffares, Alan, et al.
Pubblicazione: (2025)
Active Timepoint Selection for Learning Measure-Valued Trajectories
di: Huynh, Nicolas, et al.
Pubblicazione: (2026)
di: Huynh, Nicolas, et al.
Pubblicazione: (2026)
The Synergy of LLMs & RL Unlocks Offline Learning of Generalizable Language-Conditioned Policies with Low-fidelity Data
di: Pouplin, Thomas, et al.
Pubblicazione: (2024)
di: Pouplin, Thomas, et al.
Pubblicazione: (2024)
Hyperparameter Trajectory Inference with Conditional Lagrangian Optimal Transport
di: Amad, Harry, et al.
Pubblicazione: (2026)
di: Amad, Harry, et al.
Pubblicazione: (2026)
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities
di: Sun, Hao, et al.
Pubblicazione: (2025)
di: Sun, Hao, et al.
Pubblicazione: (2025)
Deep Learning Through A Telescoping Lens: A Simple Model Provides Empirical Insights On Grokking, Gradient Boosting & Beyond
di: Jeffares, Alan, et al.
Pubblicazione: (2024)
di: Jeffares, Alan, et al.
Pubblicazione: (2024)
Language Bottleneck Models for Qualitative Knowledge State Modeling
di: Berthon, Antonin, et al.
Pubblicazione: (2025)
di: Berthon, Antonin, et al.
Pubblicazione: (2025)
Simulating Viva Voce Examinations to Evaluate Clinical Reasoning in Large Language Models
di: Chiu, Christopher, et al.
Pubblicazione: (2025)
di: Chiu, Christopher, et al.
Pubblicazione: (2025)
Self-Healing Machine Learning: A Framework for Autonomous Adaptation in Real-World Environments
di: Rauba, Paulius, et al.
Pubblicazione: (2024)
di: Rauba, Paulius, et al.
Pubblicazione: (2024)
DC-Check: A Data-Centric AI checklist to guide the development of reliable machine learning systems
di: Seedat, Nabeel, et al.
Pubblicazione: (2022)
di: Seedat, Nabeel, et al.
Pubblicazione: (2022)
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs
di: Sun, Hao, et al.
Pubblicazione: (2025)
di: Sun, Hao, et al.
Pubblicazione: (2025)
Machine Learning with Requirements: a Manifesto
di: Giunchiglia, Eleonora, et al.
Pubblicazione: (2023)
di: Giunchiglia, Eleonora, et al.
Pubblicazione: (2023)
When is Off-Policy Evaluation (Reward Modeling) Useful in Contextual Bandits? A Data-Centric Perspective
di: Sun, Hao, et al.
Pubblicazione: (2023)
di: Sun, Hao, et al.
Pubblicazione: (2023)
Interpretable Reward Modeling with Active Concept Bottlenecks
di: Laguna, Sonia, et al.
Pubblicazione: (2025)
di: Laguna, Sonia, et al.
Pubblicazione: (2025)
Query-Dependent Prompt Evaluation and Optimization with Offline Inverse RL
di: Sun, Hao, et al.
Pubblicazione: (2023)
di: Sun, Hao, et al.
Pubblicazione: (2023)
Curated LLM: Synergy of LLMs and Data Curation for tabular augmentation in low-data regimes
di: Seedat, Nabeel, et al.
Pubblicazione: (2023)
di: Seedat, Nabeel, et al.
Pubblicazione: (2023)
You can't handle the (dirty) truth: Data-centric insights improve pseudo-labeling
di: Seedat, Nabeel, et al.
Pubblicazione: (2024)
di: Seedat, Nabeel, et al.
Pubblicazione: (2024)
Large Language Models to Enhance Bayesian Optimization
di: Liu, Tennison, et al.
Pubblicazione: (2024)
di: Liu, Tennison, et al.
Pubblicazione: (2024)
Time Series Diffusion in the Frequency Domain
di: Crabbé, Jonathan, et al.
Pubblicazione: (2024)
di: Crabbé, Jonathan, et al.
Pubblicazione: (2024)
Meta-Learners for Partially-Identified Treatment Effects Across Multiple Environments
di: Schweisthal, Jonas, et al.
Pubblicazione: (2024)
di: Schweisthal, Jonas, et al.
Pubblicazione: (2024)
AutoPrognosis 2.0: Democratizing Diagnostic and Prognostic Modeling in Healthcare with Automated Machine Learning
di: Imrie, Fergus, et al.
Pubblicazione: (2022)
di: Imrie, Fergus, et al.
Pubblicazione: (2022)
Defining Expertise: Applications to Treatment Effect Estimation
di: Hüyük, Alihan, et al.
Pubblicazione: (2024)
di: Hüyük, Alihan, et al.
Pubblicazione: (2024)
CliMB: An AI-enabled Partner for Clinical Predictive Modeling
di: Saveliev, Evgeny, et al.
Pubblicazione: (2024)
di: Saveliev, Evgeny, et al.
Pubblicazione: (2024)
Hypothesis Hunting with Evolving Networks of Autonomous Scientific Agents
di: Liu, Tennison, et al.
Pubblicazione: (2025)
di: Liu, Tennison, et al.
Pubblicazione: (2025)
On the consistent reasoning paradox of intelligence and optimal trust in AI: The power of 'I don't know'
di: Bastounis, Alexander, et al.
Pubblicazione: (2024)
di: Bastounis, Alexander, et al.
Pubblicazione: (2024)
Interpretable DNA Sequence Classification via Dynamic Feature Generation in Decision Trees
di: Huynh, Nicolas, et al.
Pubblicazione: (2026)
di: Huynh, Nicolas, et al.
Pubblicazione: (2026)
DAGnosis: Localized Identification of Data Inconsistencies using Structures
di: Huynh, Nicolas, et al.
Pubblicazione: (2024)
di: Huynh, Nicolas, et al.
Pubblicazione: (2024)
Influence-Guided Symbolic Regression: Scientific Discovery via LLM-Driven Equation Search with Granular Feedback
di: Saveliev, Evgeny S., et al.
Pubblicazione: (2026)
di: Saveliev, Evgeny S., et al.
Pubblicazione: (2026)
Unveiling the Power of Sparse Neural Networks for Feature Selection
di: Atashgahi, Zahra, et al.
Pubblicazione: (2024)
di: Atashgahi, Zahra, et al.
Pubblicazione: (2024)
CauSim: Scaling Causal Reasoning with Increasingly Complex Causal Simulators
di: Astorga, Nicolás, et al.
Pubblicazione: (2026)
di: Astorga, Nicolás, et al.
Pubblicazione: (2026)
Truly Self-Improving Agents Require Intrinsic Metacognitive Learning
di: Liu, Tennison, et al.
Pubblicazione: (2025)
di: Liu, Tennison, et al.
Pubblicazione: (2025)
Cascaded Language Models for Cost-effective Human-AI Decision-Making
di: Fanconi, Claudio, et al.
Pubblicazione: (2025)
di: Fanconi, Claudio, et al.
Pubblicazione: (2025)
GameTalk: Training LLMs for Strategic Conversation
di: Vendrell, Victor Conchello, et al.
Pubblicazione: (2026)
di: Vendrell, Victor Conchello, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Discovery of Hidden Miscalibration Regimes
di: Kobalczyk, Katarzyna, et al.
Pubblicazione: (2026) -
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes
di: Kobalczyk, Katarzyna, et al.
Pubblicazione: (2024) -
Eliciting Numerical Predictive Distributions of LLMs Without Autoregression
di: Piskorz, Julianna, et al.
Pubblicazione: (2026) -
Active Task Disambiguation with LLMs
di: Kobalczyk, Katarzyna, et al.
Pubblicazione: (2025) -
Towards Automated Knowledge Integration From Human-Interpretable Representations
di: Kobalczyk, Katarzyna, et al.
Pubblicazione: (2024)