Guiding LLM Decision-Making with Fairness Reward Models
Fuente:
arXiv
Saved in:
| Main Authors: | Hall, Zara, Subbiah, Melanie, Zollo, Thomas P, McKeown, Kathleen, Zemel, Richard |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Confidence Calibration in Vision-Language-Action Models
by: Zollo, Thomas P, et al.
Published: (2025)
by: Zollo, Thomas P, et al.
Published: (2025)
Test-Time Warmup for Multimodal Large Language Models
by: Rajaneesh, Nikita, et al.
Published: (2025)
by: Rajaneesh, Nikita, et al.
Published: (2025)
Unsupervised Confidence Calibration for Reasoning LLMs from a Single Generation
by: Zollo, Thomas, et al.
Published: (2026)
by: Zollo, Thomas, et al.
Published: (2026)
Tell Me What To Learn: Generalizing Neural Memory to be Controllable in Natural Language
by: Bennett, Max S., et al.
Published: (2026)
by: Bennett, Max S., et al.
Published: (2026)
DYSTIL: Dynamic Strategy Induction with Large Language Models for Reinforcement Learning
by: Wang, Borui, et al.
Published: (2025)
by: Wang, Borui, et al.
Published: (2025)
LiveNewsBench: Evaluating LLM Web Search Capabilities with Freshly Curated News
by: Zhang, Yunfan, et al.
Published: (2026)
by: Zhang, Yunfan, et al.
Published: (2026)
Exploring Chain-of-Thought Reasoning for Steerable Pluralistic Alignment
by: Zhang, Yunfan, et al.
Published: (2025)
by: Zhang, Yunfan, et al.
Published: (2025)
Adaptive Elicitation of Latent Information Using Natural Language
by: Wang, Jimmy, et al.
Published: (2025)
by: Wang, Jimmy, et al.
Published: (2025)
QuEst: Enhancing Estimates of Quantile-Based Distributional Measures Using Model Predictions
by: Deng, Zhun, et al.
Published: (2025)
by: Deng, Zhun, et al.
Published: (2025)
Distribution-Free Statistical Dispersion Control for Societal Applications
by: Deng, Zhun, et al.
Published: (2023)
by: Deng, Zhun, et al.
Published: (2023)
Counterfactual Simulatability of LLM Explanations for Generation Tasks
by: Limpijankit, Marvin, et al.
Published: (2025)
by: Limpijankit, Marvin, et al.
Published: (2025)
Reading Subtext: Evaluating Large Language Models on Short Story Summarization with Writers
by: Subbiah, Melanie, et al.
Published: (2024)
by: Subbiah, Melanie, et al.
Published: (2024)
Improving Predictor Reliability with Selective Recalibration
by: Zollo, Thomas P., et al.
Published: (2024)
by: Zollo, Thomas P., et al.
Published: (2024)
Whom to Query for What: Adaptive Group Elicitation via Multi-Turn LLM Interactions
by: Ding, Ruomeng, et al.
Published: (2026)
by: Ding, Ruomeng, et al.
Published: (2026)
Towards Effective Discrimination Testing for Generative AI
by: Zollo, Thomas P., et al.
Published: (2024)
by: Zollo, Thomas P., et al.
Published: (2024)
Prompt Risk Control: A Rigorous Framework for Responsible Deployment of Large Language Models
by: Zollo, Thomas P., et al.
Published: (2023)
by: Zollo, Thomas P., et al.
Published: (2023)
Parallel Structures in Pre-training Data Yield In-Context Learning
by: Chen, Yanda, et al.
Published: (2024)
by: Chen, Yanda, et al.
Published: (2024)
On the Relation between Sensitivity and Accuracy in In-context Learning
by: Chen, Yanda, et al.
Published: (2022)
by: Chen, Yanda, et al.
Published: (2022)
Computational Representations of Character Significance in Novels
by: Mian, Haaris, et al.
Published: (2026)
by: Mian, Haaris, et al.
Published: (2026)
No Compute Left Behind: Rethinking Reasoning and Sampling with Masked Diffusion Models
by: Horvitz, Zachary, et al.
Published: (2025)
by: Horvitz, Zachary, et al.
Published: (2025)
Whose Story Gets Told? Positionality and Bias in LLM Summaries of Life Narratives
by: Subbiah, Melanie, et al.
Published: (2026)
by: Subbiah, Melanie, et al.
Published: (2026)
Estimating Tail Risks in Language Model Output Distributions
by: Angell, Rico, et al.
Published: (2026)
by: Angell, Rico, et al.
Published: (2026)
Social Orientation: A New Feature for Dialogue Analysis
by: Morrill, Todd, et al.
Published: (2024)
by: Morrill, Todd, et al.
Published: (2024)
AdvSumm: Adversarial Training for Bias Mitigation in Text Summarization
by: Gupta, Mukur, et al.
Published: (2025)
by: Gupta, Mukur, et al.
Published: (2025)
Getting Serious about Humor: Crafting Humor Datasets with Unfunny Large Language Models
by: Horvitz, Zachary, et al.
Published: (2024)
by: Horvitz, Zachary, et al.
Published: (2024)
Online Algorithmic Recourse by Collective Action
by: Creager, Elliot, et al.
Published: (2023)
by: Creager, Elliot, et al.
Published: (2023)
A General Framework for Inference-time Scaling and Steering of Diffusion Models
by: Singhal, Raghav, et al.
Published: (2025)
by: Singhal, Raghav, et al.
Published: (2025)
Artificial Impressions: Evaluating Large Language Model Behavior Through the Lens of Trait Impressions
by: Deas, Nicholas, et al.
Published: (2025)
by: Deas, Nicholas, et al.
Published: (2025)
StyleDistance: Stronger Content-Independent Style Embeddings with Synthetic Parallel Examples
by: Patel, Ajay, et al.
Published: (2024)
by: Patel, Ajay, et al.
Published: (2024)
Is the Top Still Spinning? Evaluating Subjectivity in Narrative Understanding
by: Subbiah, Melanie, et al.
Published: (2025)
by: Subbiah, Melanie, et al.
Published: (2025)
Remembering to Be Fair: Non-Markovian Fairness in Sequential Decision Making
by: Alamdari, Parand A., et al.
Published: (2023)
by: Alamdari, Parand A., et al.
Published: (2023)
ModelingAgent: Bridging LLMs and Mathematical Modeling for Real-World Challenges
by: Qian, Cheng, et al.
Published: (2025)
by: Qian, Cheng, et al.
Published: (2025)
Solving Zebra Puzzles Using Constraint-Guided Multi-Agent Systems
by: Berman, Shmuel, et al.
Published: (2024)
by: Berman, Shmuel, et al.
Published: (2024)
Reward Design for Justifiable Sequential Decision-Making
by: Sukovic, Aleksa, et al.
Published: (2024)
by: Sukovic, Aleksa, et al.
Published: (2024)
Summarization of Opinionated Political Documents with Varied Perspectives
by: Deas, Nicholas, et al.
Published: (2024)
by: Deas, Nicholas, et al.
Published: (2024)
STORYSUMM: Evaluating Faithfulness in Story Summarization
by: Subbiah, Melanie, et al.
Published: (2024)
by: Subbiah, Melanie, et al.
Published: (2024)
Few-Shot Design Optimization by Exploiting Auxiliary Information
by: Mani, Arjun, et al.
Published: (2026)
by: Mani, Arjun, et al.
Published: (2026)
Long-Term Fair Decision Making through Deep Generative Models
by: Hu, Yaowei, et al.
Published: (2024)
by: Hu, Yaowei, et al.
Published: (2024)
A First Full Physics Benchmark for Highly Granular Calorimeter Surrogates
by: Buss, Thorsten, et al.
Published: (2025)
by: Buss, Thorsten, et al.
Published: (2025)
Marginal Fairness: Fair Decision-Making under Risk Measures
by: Huang, Fei, et al.
Published: (2025)
by: Huang, Fei, et al.
Published: (2025)
Similar Items
-
Confidence Calibration in Vision-Language-Action Models
by: Zollo, Thomas P, et al.
Published: (2025) -
Test-Time Warmup for Multimodal Large Language Models
by: Rajaneesh, Nikita, et al.
Published: (2025) -
Unsupervised Confidence Calibration for Reasoning LLMs from a Single Generation
by: Zollo, Thomas, et al.
Published: (2026) -
Tell Me What To Learn: Generalizing Neural Memory to be Controllable in Natural Language
by: Bennett, Max S., et al.
Published: (2026) -
DYSTIL: Dynamic Strategy Induction with Large Language Models for Reinforcement Learning
by: Wang, Borui, et al.
Published: (2025)