Can Revealed Preferences Clarify LLM Alignment and Steering?
Fuente:
arXiv
Saved in:
| Main Authors: | Yamin, Khurram, Tang, Jingjing, Horvitz, Eric, Wilder, Bryan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can LLMs Reconcile Knowledge Conflicts in Counterfactual Reasoning
by: Yamin, Khurram, et al.
Published: (2025)
by: Yamin, Khurram, et al.
Published: (2025)
Dependent Randomized Rounding for Budget Constrained Experimental Design
by: Yamin, Khurram, et al.
Published: (2025)
by: Yamin, Khurram, et al.
Published: (2025)
Accounting for Missing Covariates in Heterogeneous Treatment Estimation
by: Yamin, Khurram, et al.
Published: (2024)
by: Yamin, Khurram, et al.
Published: (2024)
When Agents Say One Thing and Do Another: Validating Elicited Beliefs from LLMs
by: Yamin, Khurram, et al.
Published: (2026)
by: Yamin, Khurram, et al.
Published: (2026)
Failure Modes of LLMs for Causal Reasoning on Narratives
by: Yamin, Khurram, et al.
Published: (2024)
by: Yamin, Khurram, et al.
Published: (2024)
Predicting Language Models' Success at Zero-Shot Probabilistic Prediction
by: Ren, Kevin, et al.
Published: (2025)
by: Ren, Kevin, et al.
Published: (2025)
Utility-Directed Conformal Prediction: A Decision-Aware Framework for Actionable Uncertainty Quantification
by: Cortes-Gomez, Santiago, et al.
Published: (2024)
by: Cortes-Gomez, Santiago, et al.
Published: (2024)
Improving Instruction-Following in Language Models through Activation Steering
by: Stolfo, Alessandro, et al.
Published: (2024)
by: Stolfo, Alessandro, et al.
Published: (2024)
Comparing Targeting Strategies for Maximizing Social Welfare with Limited Resources
by: Sharma, Vibhhu, et al.
Published: (2024)
by: Sharma, Vibhhu, et al.
Published: (2024)
Where to Steer: Input-Dependent Layer Selection for Steering Improves LLM Alignment
by: Gadgil, Soham, et al.
Published: (2026)
by: Gadgil, Soham, et al.
Published: (2026)
Learning treatment effects while treating those in need
by: Wilder, Bryan, et al.
Published: (2024)
by: Wilder, Bryan, et al.
Published: (2024)
Steering Language Model Refusal with Sparse Autoencoders
by: O'Brien, Kyle, et al.
Published: (2024)
by: O'Brien, Kyle, et al.
Published: (2024)
Improving constraint-based discovery with robust propagation and reliable LLM priors
by: Lyu, Ruiqi, et al.
Published: (2025)
by: Lyu, Ruiqi, et al.
Published: (2025)
Fostering the Ecosystem of AI for Social Impact Requires Expanding and Strengthening Evaluation Standards
by: Wilder, Bryan, et al.
Published: (2025)
by: Wilder, Bryan, et al.
Published: (2025)
DSPA: Dynamic SAE Steering for Data-Efficient Preference Alignment
by: Wedgwood, James, et al.
Published: (2026)
by: Wedgwood, James, et al.
Published: (2026)
BarrierSteer: LLM Safety via Learning Barrier Steering
by: Tran, Thanh Q., et al.
Published: (2026)
by: Tran, Thanh Q., et al.
Published: (2026)
Adversarial Preference Learning for Robust LLM Alignment
by: Wang, Yuanfu, et al.
Published: (2025)
by: Wang, Yuanfu, et al.
Published: (2025)
Explaining Concept Shift with Interpretable Feature Attribution
by: Lyu, Ruiqi, et al.
Published: (2025)
by: Lyu, Ruiqi, et al.
Published: (2025)
OEUVRE: OnlinE Unbiased Variance-Reduced loss Estimation
by: Pardeshi, Kanad, et al.
Published: (2025)
by: Pardeshi, Kanad, et al.
Published: (2025)
Decision-Focused Evaluation of Worst-Case Distribution Shift
by: Ren, Kevin, et al.
Published: (2024)
by: Ren, Kevin, et al.
Published: (2024)
Combining digital data streams and epidemic networks for real time outbreak detection
by: Lyu, Ruiqi, et al.
Published: (2025)
by: Lyu, Ruiqi, et al.
Published: (2025)
Distributionally Robust Feature Selection
by: Swaroop, Maitreyi, et al.
Published: (2025)
by: Swaroop, Maitreyi, et al.
Published: (2025)
Alignment Revisited: Are Large Language Models Consistent in Stated and Revealed Preferences?
by: Gu, Zhuojun, et al.
Published: (2025)
by: Gu, Zhuojun, et al.
Published: (2025)
Data Selection for LLM Alignment Using Fine-Grained Preferences
by: Zhang, Jia, et al.
Published: (2025)
by: Zhang, Jia, et al.
Published: (2025)
Alignment with Preference Optimization Is All You Need for LLM Safety
by: Alami, Reda, et al.
Published: (2024)
by: Alami, Reda, et al.
Published: (2024)
Nearly Optimal Active Preference Learning and Its Application to LLM Alignment
by: Zhao, Yao, et al.
Published: (2026)
by: Zhao, Yao, et al.
Published: (2026)
Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior
by: Wurgaft, Daniel, et al.
Published: (2026)
by: Wurgaft, Daniel, et al.
Published: (2026)
D-STEER - Preference Alignment Techniques Learn to Behave, not to Believe -- Beneath the Surface, DPO as Steering Vector Perturbation in Activation Space
by: Raina, Samarth, et al.
Published: (2025)
by: Raina, Samarth, et al.
Published: (2025)
Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
by: Kim, Dongyoung, et al.
Published: (2024)
by: Kim, Dongyoung, et al.
Published: (2024)
HelpSteer2-Preference: Complementing Ratings with Preferences
by: Wang, Zhilin, et al.
Published: (2024)
by: Wang, Zhilin, et al.
Published: (2024)
AMPS: Adaptive Modality Preference Steering via Functional Entropy
by: Huang, Zihan, et al.
Published: (2026)
by: Huang, Zihan, et al.
Published: (2026)
Healthcare LLM Benchmarks Are Only as Good as Their Explicit Assumptions
by: Raman, Naveen, et al.
Published: (2026)
by: Raman, Naveen, et al.
Published: (2026)
Preference Alignment with Flow Matching
by: Kim, Minu, et al.
Published: (2024)
by: Kim, Minu, et al.
Published: (2024)
ActiveDPO: Active Direct Preference Optimization for Sample-Efficient Alignment
by: Lin, Xiaoqiang, et al.
Published: (2025)
by: Lin, Xiaoqiang, et al.
Published: (2025)
A General Framework for Inference-time Scaling and Steering of Diffusion Models
by: Singhal, Raghav, et al.
Published: (2025)
by: Singhal, Raghav, et al.
Published: (2025)
Robust LLM Alignment via Distributionally Robust Direct Preference Optimization
by: Xu, Zaiyan, et al.
Published: (2025)
by: Xu, Zaiyan, et al.
Published: (2025)
Leaving the Nest: Going Beyond Local Loss Functions for Predict-Then-Optimize
by: Shah, Sanket, et al.
Published: (2023)
by: Shah, Sanket, et al.
Published: (2023)
Data-Centric Human Preference with Rationales for Direct Preference Alignment
by: Just, Hoang Anh, et al.
Published: (2024)
by: Just, Hoang Anh, et al.
Published: (2024)
Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering
by: Bigelow, Eric, et al.
Published: (2025)
by: Bigelow, Eric, et al.
Published: (2025)
Reading Between the Lines: Modeling User Behavior and Costs in AI-Assisted Programming
by: Mozannar, Hussein, et al.
Published: (2022)
by: Mozannar, Hussein, et al.
Published: (2022)
Similar Items
-
Can LLMs Reconcile Knowledge Conflicts in Counterfactual Reasoning
by: Yamin, Khurram, et al.
Published: (2025) -
Dependent Randomized Rounding for Budget Constrained Experimental Design
by: Yamin, Khurram, et al.
Published: (2025) -
Accounting for Missing Covariates in Heterogeneous Treatment Estimation
by: Yamin, Khurram, et al.
Published: (2024) -
When Agents Say One Thing and Do Another: Validating Elicited Beliefs from LLMs
by: Yamin, Khurram, et al.
Published: (2026) -
Failure Modes of LLMs for Causal Reasoning on Narratives
by: Yamin, Khurram, et al.
Published: (2024)