Expected Reward Prediction, with Applications to Model Routing
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Hasanaliyev, Kenan, Alberti, Silas, Hamer, Jenny, Rajagopal, Dheeraj, Robinson, Kevin, Snoek, Jasper, Veitch, Victor, D'Amour, Alexander Nicholas |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Data Unlearning in Diffusion Models
par: Alberti, Silas, et autres
Publié: (2025)
par: Alberti, Silas, et autres
Publié: (2025)
Transforming and Combining Rewards for Aligning Large Language Models
par: Wang, Zihao, et autres
Publié: (2024)
par: Wang, Zihao, et autres
Publié: (2024)
Bias in Language Models: Beyond Trick Tests and Toward RUTEd Evaluation
par: Lum, Kristian, et autres
Publié: (2024)
par: Lum, Kristian, et autres
Publié: (2024)
The Illusion of Intervention: Your LLM-Simulated Experiment is an Observational Study
par: Lin, Victoria, et autres
Publié: (2026)
par: Lin, Victoria, et autres
Publié: (2026)
Copula-based Sensitivity Analysis for Multi-Treatment Causal Inference with Unobserved Confounding
par: Zheng, Jiajing, et autres
Publié: (2021)
par: Zheng, Jiajing, et autres
Publié: (2021)
Theoretical guarantees on the best-of-n alignment policy
par: Beirami, Ahmad, et autres
Publié: (2024)
par: Beirami, Ahmad, et autres
Publié: (2024)
RATE: Causal Explainability of Reward Models with Imperfect Counterfactuals
par: Reber, David, et autres
Publié: (2024)
par: Reber, David, et autres
Publié: (2024)
How Far Can We Extract Diverse Perspectives from Large Language Models?
par: Hayati, Shirley Anugrah, et autres
Publié: (2023)
par: Hayati, Shirley Anugrah, et autres
Publié: (2023)
Confidence Calibration and Rationalization for LLMs via Multi-Agent Deliberation
par: Yang, Ruixin, et autres
Publié: (2024)
par: Yang, Ruixin, et autres
Publié: (2024)
Scalable Influence and Fact Tracing for Large Language Model Pretraining
par: Chang, Tyler A., et autres
Publié: (2024)
par: Chang, Tyler A., et autres
Publié: (2024)
Choosing a Proxy Metric from Past Experiments
par: Tripuraneni, Nilesh, et autres
Publié: (2023)
par: Tripuraneni, Nilesh, et autres
Publié: (2023)
NeoBabel: A Multilingual Open Tower for Visual Generation
par: Derakhshani, Mohammad Mahdi, et autres
Publié: (2025)
par: Derakhshani, Mohammad Mahdi, et autres
Publié: (2025)
Steering off Course: Reliability Challenges in Steering Language Models
par: Da Silva, Patrick Queiroz, et autres
Publié: (2025)
par: Da Silva, Patrick Queiroz, et autres
Publié: (2025)
J-P: MDP. FP. PP.: Characterizing Total Expected Rewards in Markov Decision Processes as Least Fixed Points with an Application to Operational Semantics of Probabilistic Programs (Technical Report)
par: Batz, Kevin, et autres
Publié: (2024)
par: Batz, Kevin, et autres
Publié: (2024)
A Unified Approach to Routing and Cascading for LLMs
par: Dekoninck, Jasper, et autres
Publié: (2024)
par: Dekoninck, Jasper, et autres
Publié: (2024)
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs
par: von Recum, Alexander, et autres
Publié: (2024)
par: von Recum, Alexander, et autres
Publié: (2024)
BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n Sampling
par: Gui, Lin, et autres
Publié: (2024)
par: Gui, Lin, et autres
Publié: (2024)
Predictive Churn with the Set of Good Models
par: Watson-Daniels, Jamelle, et autres
Publié: (2024)
par: Watson-Daniels, Jamelle, et autres
Publié: (2024)
Deconfounding Scores and Representation Learning for Causal Effect Estimation with Weak Overlap
par: Clivio, Oscar, et autres
Publié: (2026)
par: Clivio, Oscar, et autres
Publié: (2026)
CLIP the Bias: How Useful is Balancing Data in Multimodal Learning?
par: Alabdulmohsin, Ibrahim, et autres
Publié: (2024)
par: Alabdulmohsin, Ibrahim, et autres
Publié: (2024)
Making physical activity fun and accessible to adults with intellectual disabilities: A pilot study of a gamification intervention
par: Stéphanie Turgeon, et autres
Publié: (2024)
par: Stéphanie Turgeon, et autres
Publié: (2024)
The Linear Representation Hypothesis and the Geometry of Large Language Models
par: Park, Kiho, et autres
Publié: (2023)
par: Park, Kiho, et autres
Publié: (2023)
Concept Algebra for (Score-Based) Text-Controlled Generative Models
par: Wang, Zihao, et autres
Publié: (2023)
par: Wang, Zihao, et autres
Publié: (2023)
Rubric-Guided Process Reward for Stepwise Model Routing
par: Ye, Shenghao, et autres
Publié: (2026)
par: Ye, Shenghao, et autres
Publié: (2026)
TIER: Trajectory-Invariant Execution Rewards for Multi-Step Tool Composition
par: Kulkarni, Anay, et autres
Publié: (2026)
par: Kulkarni, Anay, et autres
Publié: (2026)
Reward-Based Online LLM Routing via NeuralUCB
par: Tsai, Ming-Hua, et autres
Publié: (2026)
par: Tsai, Ming-Hua, et autres
Publié: (2026)
Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
par: Eisenstein, Jacob, et autres
Publié: (2023)
par: Eisenstein, Jacob, et autres
Publié: (2023)
Incorporating Hierarchical Semantics in Sparse Autoencoder Architectures
par: Muchane, Mark, et autres
Publié: (2025)
par: Muchane, Mark, et autres
Publié: (2025)
An Iris for Expected Cost Analysis
par: Lohse, Janine, et autres
Publié: (2024)
par: Lohse, Janine, et autres
Publié: (2024)
On the Origins of Linear Representations in Large Language Models
par: Jiang, Yibo, et autres
Publié: (2024)
par: Jiang, Yibo, et autres
Publié: (2024)
Reward Prediction with Factorized World States
par: Shen, Yijun, et autres
Publié: (2026)
par: Shen, Yijun, et autres
Publié: (2026)
The Geometry of Categorical and Hierarchical Concepts in Large Language Models
par: Park, Kiho, et autres
Publié: (2024)
par: Park, Kiho, et autres
Publié: (2024)
The Information Geometry of Softmax: Probing and Steering
par: Park, Kiho, et autres
Publié: (2026)
par: Park, Kiho, et autres
Publié: (2026)
Simple linear attention language models balance the recall-throughput tradeoff
par: Arora, Simran, et autres
Publié: (2024)
par: Arora, Simran, et autres
Publié: (2024)
PIGEON: Predicting Image Geolocations
par: Haas, Lukas, et autres
Publié: (2023)
par: Haas, Lukas, et autres
Publié: (2023)
Smaller Language Models are capable of selecting Instruction-Tuning Training Data for Larger Language Models
par: Mekala, Dheeraj, et autres
Publié: (2024)
par: Mekala, Dheeraj, et autres
Publié: (2024)
DIY-MKG: An LLM-Based Polyglot Language Learning System
par: Tang, Kenan, et autres
Publié: (2025)
par: Tang, Kenan, et autres
Publié: (2025)
Expect the Unexpected? Testing the Surprisal of Salient Entities
par: Lin, Jessica, et autres
Publié: (2026)
par: Lin, Jessica, et autres
Publié: (2026)
Automated Expected Cost Analysis for Quantum Programs
par: Moser, Georg, et autres
Publié: (2026)
par: Moser, Georg, et autres
Publié: (2026)
A Novel Constructive Routing Algorithm for Fleet Size and Mix Vehicle Routing Problem
par: Kenan KARAGUL
Publié: (2014)
par: Kenan KARAGUL
Publié: (2014)
Documents similaires
-
Data Unlearning in Diffusion Models
par: Alberti, Silas, et autres
Publié: (2025) -
Transforming and Combining Rewards for Aligning Large Language Models
par: Wang, Zihao, et autres
Publié: (2024) -
Bias in Language Models: Beyond Trick Tests and Toward RUTEd Evaluation
par: Lum, Kristian, et autres
Publié: (2024) -
The Illusion of Intervention: Your LLM-Simulated Experiment is an Observational Study
par: Lin, Victoria, et autres
Publié: (2026) -
Copula-based Sensitivity Analysis for Multi-Treatment Causal Inference with Unobserved Confounding
par: Zheng, Jiajing, et autres
Publié: (2021)