Saved in:
| Main Authors: | Yamin, Khurram, Tang, Jingjing, Cortes-Gomez, Santiago, Sharma, Amit, Horvitz, Eric, Wilder, Bryan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2602.06286 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can Revealed Preferences Clarify LLM Alignment and Steering?
by: Yamin, Khurram, et al.
Published: (2026)
by: Yamin, Khurram, et al.
Published: (2026)
Can LLMs Reconcile Knowledge Conflicts in Counterfactual Reasoning
by: Yamin, Khurram, et al.
Published: (2025)
by: Yamin, Khurram, et al.
Published: (2025)
Accounting for Missing Covariates in Heterogeneous Treatment Estimation
by: Yamin, Khurram, et al.
Published: (2024)
by: Yamin, Khurram, et al.
Published: (2024)
Dependent Randomized Rounding for Budget Constrained Experimental Design
by: Yamin, Khurram, et al.
Published: (2025)
by: Yamin, Khurram, et al.
Published: (2025)
Utility-Directed Conformal Prediction: A Decision-Aware Framework for Actionable Uncertainty Quantification
by: Cortes-Gomez, Santiago, et al.
Published: (2024)
by: Cortes-Gomez, Santiago, et al.
Published: (2024)
Large Language Models Often Say One Thing and Do Another
by: Xu, Ruoxi, et al.
Published: (2025)
by: Xu, Ruoxi, et al.
Published: (2025)
Failure Modes of LLMs for Causal Reasoning on Narratives
by: Yamin, Khurram, et al.
Published: (2024)
by: Yamin, Khurram, et al.
Published: (2024)
Predicting Language Models' Success at Zero-Shot Probabilistic Prediction
by: Ren, Kevin, et al.
Published: (2025)
by: Ren, Kevin, et al.
Published: (2025)
The Limits of AI-Driven Allocation: Optimal Screening under Aleatoric Uncertainty
by: Cortes-Gomez, Santiago, et al.
Published: (2026)
by: Cortes-Gomez, Santiago, et al.
Published: (2026)
Auditing Fairness by Betting
by: Chugg, Ben, et al.
Published: (2023)
by: Chugg, Ben, et al.
Published: (2023)
Say One Thing, Do Another? Diagnosing Reasoning-Execution Gaps in VLM-Powered Mobile-Use Agents
by: Dong, Lingzhong, et al.
Published: (2025)
by: Dong, Lingzhong, et al.
Published: (2025)
Healthcare LLM Benchmarks Are Only as Good as Their Explicit Assumptions
by: Raman, Naveen, et al.
Published: (2026)
by: Raman, Naveen, et al.
Published: (2026)
"Is This It?": Towards Ecologically Valid Benchmarks for Situated Collaboration
by: Bohus, Dan, et al.
Published: (2024)
by: Bohus, Dan, et al.
Published: (2024)
Say Anything but This: When Tokenizer Betrays Reasoning in LLMs
by: Ayoobi, Navid, et al.
Published: (2026)
by: Ayoobi, Navid, et al.
Published: (2026)
Do as I Say, Not as I Do: Instruction-Induction Conflict in LLMs
by: Camassa, Carolina, et al.
Published: (2026)
by: Camassa, Carolina, et al.
Published: (2026)
SayNext-Bench: Why Do LLMs Struggle with Next-Utterance Anticipation?
by: Yang, Yueyi, et al.
Published: (2026)
by: Yang, Yueyi, et al.
Published: (2026)
One Thing Follows Another
by: Rosenthal, Sarah
Published: (2025)
by: Rosenthal, Sarah
Published: (2025)
Data-driven Design of Randomized Control Trials with Guaranteed Treatment Effects
by: Cortes-Gomez, Santiago, et al.
Published: (2024)
by: Cortes-Gomez, Santiago, et al.
Published: (2024)
Valid Inference with Imperfect Synthetic Data
by: Byun, Yewon, et al.
Published: (2025)
by: Byun, Yewon, et al.
Published: (2025)
Accurate Measures of Vaccination and Concerns of Vaccine Holdouts from Web Search Logs
by: Chang, Serina, et al.
Published: (2023)
by: Chang, Serina, et al.
Published: (2023)
RealFin: How Well Do LLMs Reason About Finance When Users Leave Things Unsaid?
by: Dai, Yuyang, et al.
Published: (2026)
by: Dai, Yuyang, et al.
Published: (2026)
IV Co-Scientist: Multi-Agent LLM Framework for Causal Instrumental Variable Discovery
by: Sheth, Ivaxi, et al.
Published: (2026)
by: Sheth, Ivaxi, et al.
Published: (2026)
Eliciting Better Multilingual Structured Reasoning from LLMs through Code
by: Li, Bryan, et al.
Published: (2024)
by: Li, Bryan, et al.
Published: (2024)
SpatialEpiBench: Benchmarking Spatial Information and Epidemic Priors in Forecasting
by: Lyu, Ruiqi, et al.
Published: (2026)
by: Lyu, Ruiqi, et al.
Published: (2026)
Doing What They Say, Not What They Reason: Locating the Faithfulness Gap in LLM Agents
by: Wang, Yufeng
Published: (2026)
by: Wang, Yufeng
Published: (2026)
When Models Know More Than They Say: Probing Analogical Reasoning in LLMs
by: McGovern, Hope, et al.
Published: (2026)
by: McGovern, Hope, et al.
Published: (2026)
Tandem Training for Language Models
by: West, Robert, et al.
Published: (2025)
by: West, Robert, et al.
Published: (2025)
MEMENTO: Teaching LLMs to Manage Their Own Context
by: Kontonis, Vasilis, et al.
Published: (2026)
by: Kontonis, Vasilis, et al.
Published: (2026)
When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use Agents
by: Jones, Jaylen, et al.
Published: (2026)
by: Jones, Jaylen, et al.
Published: (2026)
Federated Epidemic Surveillance
by: Lyu, Ruiqi, et al.
Published: (2023)
by: Lyu, Ruiqi, et al.
Published: (2023)
FlipLLM: Efficient Bit-Flip Attacks on Multimodal LLMs using Reinforcement Learning
by: Khalil, Khurram, et al.
Published: (2025)
by: Khalil, Khurram, et al.
Published: (2025)
Comparing Targeting Strategies for Maximizing Social Welfare with Limited Resources
by: Sharma, Vibhhu, et al.
Published: (2024)
by: Sharma, Vibhhu, et al.
Published: (2024)
When Agents See Humans as the Outgroup: Belief-Dependent Bias in LLM-Powered Agents
by: Wang, Zongwei, et al.
Published: (2026)
by: Wang, Zongwei, et al.
Published: (2026)
When is an Embedding Model More Promising than Another?
by: Darrin, Maxime, et al.
Published: (2024)
by: Darrin, Maxime, et al.
Published: (2024)
Computationally Assisted Quality Control for Public Health Data Streams
by: Joshi, Ananya, et al.
Published: (2023)
by: Joshi, Ananya, et al.
Published: (2023)
Collaborative Belief Reasoning with LLMs for Efficient Multi-Agent Collaboration
by: Wang, Zhimin, et al.
Published: (2025)
by: Wang, Zhimin, et al.
Published: (2025)
Challenges in Human-Agent Communication
by: Bansal, Gagan, et al.
Published: (2024)
by: Bansal, Gagan, et al.
Published: (2024)
LLMs on a Budget? Say HOLA
by: Siddiqui, Zohaib Hasan, et al.
Published: (2025)
by: Siddiqui, Zohaib Hasan, et al.
Published: (2025)
Leaving the Nest: Going Beyond Local Loss Functions for Predict-Then-Optimize
by: Shah, Sanket, et al.
Published: (2023)
by: Shah, Sanket, et al.
Published: (2023)
When Correct Beliefs Collapse: Epistemic Resilience of LLMs under Clinical Pressure
by: Xiao, Boyu, et al.
Published: (2026)
by: Xiao, Boyu, et al.
Published: (2026)
Similar Items
-
Can Revealed Preferences Clarify LLM Alignment and Steering?
by: Yamin, Khurram, et al.
Published: (2026) -
Can LLMs Reconcile Knowledge Conflicts in Counterfactual Reasoning
by: Yamin, Khurram, et al.
Published: (2025) -
Accounting for Missing Covariates in Heterogeneous Treatment Estimation
by: Yamin, Khurram, et al.
Published: (2024) -
Dependent Randomized Rounding for Budget Constrained Experimental Design
by: Yamin, Khurram, et al.
Published: (2025) -
Utility-Directed Conformal Prediction: A Decision-Aware Framework for Actionable Uncertainty Quantification
by: Cortes-Gomez, Santiago, et al.
Published: (2024)