Interpretable Reward Modeling with Active Concept Bottlenecks
Fuente:
arXiv
Saved in:
| Main Authors: | Laguna, Sonia, Kobalczyk, Katarzyna, Vogt, Julia E., Van der Schaar, Mihaela |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Automated Knowledge Integration From Human-Interpretable Representations
by: Kobalczyk, Katarzyna, et al.
Published: (2024)
by: Kobalczyk, Katarzyna, et al.
Published: (2024)
Preference Learning for AI Alignment: a Causal Perspective
by: Kobalczyk, Katarzyna, et al.
Published: (2025)
by: Kobalczyk, Katarzyna, et al.
Published: (2025)
Discovery of Hidden Miscalibration Regimes
by: Kobalczyk, Katarzyna, et al.
Published: (2026)
by: Kobalczyk, Katarzyna, et al.
Published: (2026)
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes
by: Kobalczyk, Katarzyna, et al.
Published: (2024)
by: Kobalczyk, Katarzyna, et al.
Published: (2024)
Active Task Disambiguation with LLMs
by: Kobalczyk, Katarzyna, et al.
Published: (2025)
by: Kobalczyk, Katarzyna, et al.
Published: (2025)
Eliciting Numerical Predictive Distributions of LLMs Without Autoregression
by: Piskorz, Julianna, et al.
Published: (2026)
by: Piskorz, Julianna, et al.
Published: (2026)
Stochastic Concept Bottleneck Models
by: Vandenhirtz, Moritz, et al.
Published: (2024)
by: Vandenhirtz, Moritz, et al.
Published: (2024)
Beyond Concept Bottleneck Models: How to Make Black Boxes Intervenable?
by: Laguna, Sonia, et al.
Published: (2024)
by: Laguna, Sonia, et al.
Published: (2024)
Post-hoc Stochastic Concept Bottleneck Models
by: Hoffmann, Wiktor Jan, et al.
Published: (2025)
by: Hoffmann, Wiktor Jan, et al.
Published: (2025)
Language Bottleneck Models for Qualitative Knowledge State Modeling
by: Berthon, Antonin, et al.
Published: (2025)
by: Berthon, Antonin, et al.
Published: (2025)
Exploiting Interpretable Capabilities with Concept-Enhanced Diffusion and Prototype Networks
by: Carballo-Castro, Alba, et al.
Published: (2024)
by: Carballo-Castro, Alba, et al.
Published: (2024)
Soft Mixture Denoising: Beyond the Expressive Bottleneck of Diffusion Models
by: Li, Yangming, et al.
Published: (2023)
by: Li, Yangming, et al.
Published: (2023)
Active Timepoint Selection for Learning Measure-Valued Trajectories
by: Huynh, Nicolas, et al.
Published: (2026)
by: Huynh, Nicolas, et al.
Published: (2026)
Measuring Leakage in Concept-Based Methods: An Information Theoretic Approach
by: Makonnen, Mikael, et al.
Published: (2025)
by: Makonnen, Mikael, et al.
Published: (2025)
Beyond the ATE: Interpretable Modelling of Treatment Effects over Dose and Time
by: Piskorz, Julianna, et al.
Published: (2025)
by: Piskorz, Julianna, et al.
Published: (2025)
Deep Hierarchical Learning with Nested Subspace Networks for Large Language Models
by: Rauba, Paulius, et al.
Published: (2025)
by: Rauba, Paulius, et al.
Published: (2025)
Matchmaker: Self-Improving Large Language Model Programs for Schema Matching
by: Seedat, Nabeel, et al.
Published: (2024)
by: Seedat, Nabeel, et al.
Published: (2024)
Dense Reward for Free in Reinforcement Learning from Human Feedback
by: Chan, Alex J., et al.
Published: (2024)
by: Chan, Alex J., et al.
Published: (2024)
Why Tabular Foundation Models Should Be a Research Priority
by: van Breugel, Boris, et al.
Published: (2024)
by: van Breugel, Boris, et al.
Published: (2024)
Inverse-RLignment: Large Language Model Alignment from Demonstrations through Inverse Reinforcement Learning
by: Sun, Hao, et al.
Published: (2024)
by: Sun, Hao, et al.
Published: (2024)
Reference-Guided Machine Unlearning
by: Mirlach, Jonas, et al.
Published: (2026)
by: Mirlach, Jonas, et al.
Published: (2026)
No Equations Needed: Learning System Dynamics Without Relying on Closed-Form ODEs
by: Kacprzyk, Krzysztof, et al.
Published: (2025)
by: Kacprzyk, Krzysztof, et al.
Published: (2025)
Shape Arithmetic Expressions: Advancing Scientific Discovery Beyond Closed-Form Equations
by: Kacprzyk, Krzysztof, et al.
Published: (2024)
by: Kacprzyk, Krzysztof, et al.
Published: (2024)
On Error Propagation of Diffusion Models
by: Li, Yangming, et al.
Published: (2023)
by: Li, Yangming, et al.
Published: (2023)
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs
by: Sun, Hao, et al.
Published: (2025)
by: Sun, Hao, et al.
Published: (2025)
Tiny Autoregressive Recursive Models
by: Rauba, Paulius, et al.
Published: (2026)
by: Rauba, Paulius, et al.
Published: (2026)
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities
by: Sun, Hao, et al.
Published: (2025)
by: Sun, Hao, et al.
Published: (2025)
Stochastic Encodings for Active Feature Acquisition
by: Norcliffe, Alexander, et al.
Published: (2025)
by: Norcliffe, Alexander, et al.
Published: (2025)
The Synergy of LLMs & RL Unlocks Offline Learning of Generalizable Language-Conditioned Policies with Low-fidelity Data
by: Pouplin, Thomas, et al.
Published: (2024)
by: Pouplin, Thomas, et al.
Published: (2024)
Not All Explanations for Deep Learning Phenomena Are Equally Valuable
by: Jeffares, Alan, et al.
Published: (2025)
by: Jeffares, Alan, et al.
Published: (2025)
Hyperparameter Trajectory Inference with Conditional Lagrangian Optimal Transport
by: Amad, Harry, et al.
Published: (2026)
by: Amad, Harry, et al.
Published: (2026)
When is Off-Policy Evaluation (Reward Modeling) Useful in Contextual Bandits? A Data-Centric Perspective
by: Sun, Hao, et al.
Published: (2023)
by: Sun, Hao, et al.
Published: (2023)
Risk-Sensitive Diffusion: Robustly Optimizing Diffusion Models with Noisy Samples
by: Li, Yangming, et al.
Published: (2024)
by: Li, Yangming, et al.
Published: (2024)
Interpretable Prognostics with Concept Bottleneck Models
by: Forest, Florent, et al.
Published: (2024)
by: Forest, Florent, et al.
Published: (2024)
Simulating Viva Voce Examinations to Evaluate Clinical Reasoning in Large Language Models
by: Chiu, Christopher, et al.
Published: (2025)
by: Chiu, Christopher, et al.
Published: (2025)
Deep Learning Through A Telescoping Lens: A Simple Model Provides Empirical Insights On Grokking, Gradient Boosting & Beyond
by: Jeffares, Alan, et al.
Published: (2024)
by: Jeffares, Alan, et al.
Published: (2024)
Fast Generalization after Interpolation via Critically Damped Momentum Optimization
by: Muscarnera, Luca, et al.
Published: (2026)
by: Muscarnera, Luca, et al.
Published: (2026)
Decision Tree Induction Through LLMs via Semantically-Aware Evolution
by: Liu, Tennison, et al.
Published: (2025)
by: Liu, Tennison, et al.
Published: (2025)
What's the next frontier for Data-centric AI? Data Savvy Agents
by: Seedat, Nabeel, et al.
Published: (2025)
by: Seedat, Nabeel, et al.
Published: (2025)
A Study of Posterior Stability for Time-Series Latent Diffusion
by: Li, Yangming, et al.
Published: (2024)
by: Li, Yangming, et al.
Published: (2024)
Similar Items
-
Towards Automated Knowledge Integration From Human-Interpretable Representations
by: Kobalczyk, Katarzyna, et al.
Published: (2024) -
Preference Learning for AI Alignment: a Causal Perspective
by: Kobalczyk, Katarzyna, et al.
Published: (2025) -
Discovery of Hidden Miscalibration Regimes
by: Kobalczyk, Katarzyna, et al.
Published: (2026) -
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes
by: Kobalczyk, Katarzyna, et al.
Published: (2024) -
Active Task Disambiguation with LLMs
by: Kobalczyk, Katarzyna, et al.
Published: (2025)