Omitted Variable Bias in Language Models Under Distribution Shift
Fuente:
arXiv
Saved in:
| Main Authors: | Lin, Victoria, Morency, Louis-Philippe, Ben-Michael, Eli |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Optimizing Language Models for Human Preferences is a Causal Inference Problem
by: Lin, Victoria, et al.
Published: (2024)
by: Lin, Victoria, et al.
Published: (2024)
Isolated Causal Effects of Natural Language
by: Lin, Victoria, et al.
Published: (2024)
by: Lin, Victoria, et al.
Published: (2024)
Long Story Short: Omitted Variable Bias in Causal Machine Learning
by: Chernozhukov, Victor, et al.
Published: (2021)
by: Chernozhukov, Victor, et al.
Published: (2021)
Zipf Distributions from Two-Stage Symbolic Processes: Stability Under Stochastic Lexical Filtering
by: Berman, Vladimir
Published: (2025)
by: Berman, Vladimir
Published: (2025)
Robust Detection of Watermarks for Large Language Models Under Human Edits
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
LIDS: LLM Summary Inference Under the Layered Lens
by: Park, Dylan, et al.
Published: (2026)
by: Park, Dylan, et al.
Published: (2026)
The Illusion of Intervention: Your LLM-Simulated Experiment is an Observational Study
by: Lin, Victoria, et al.
Published: (2026)
by: Lin, Victoria, et al.
Published: (2026)
Social Caption: Evaluating Social Understanding in Multimodal Models
by: Thumu, Bhaavanaa, et al.
Published: (2026)
by: Thumu, Bhaavanaa, et al.
Published: (2026)
CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models
by: Tu, Ruibo, et al.
Published: (2024)
by: Tu, Ruibo, et al.
Published: (2024)
Causal Graph Discovery with Retrieval-Augmented Generation based Large Language Models
by: Zhang, Yuzhe, et al.
Published: (2024)
by: Zhang, Yuzhe, et al.
Published: (2024)
TWIN-GPT: Digital Twins for Clinical Trials via Large Language Model
by: Wang, Yue, et al.
Published: (2024)
by: Wang, Yue, et al.
Published: (2024)
Social Genome: Grounded Social Reasoning Abilities of Multimodal Models
by: Mathur, Leena, et al.
Published: (2025)
by: Mathur, Leena, et al.
Published: (2025)
Using Imperfect Surrogates for Downstream Inference: Design-based Supervised Learning for Social Science Applications of Large Language Models
by: Egami, Naoki, et al.
Published: (2023)
by: Egami, Naoki, et al.
Published: (2023)
Multi-Source Conformal Inference Under Distribution Shift
by: Liu, Yi, et al.
Published: (2024)
by: Liu, Yi, et al.
Published: (2024)
End-To-End Causal Effect Estimation from Unstructured Natural Language Data
by: Dhawan, Nikita, et al.
Published: (2024)
by: Dhawan, Nikita, et al.
Published: (2024)
Advancing Social Intelligence in AI Agents: Technical Challenges and Open Questions
by: Mathur, Leena, et al.
Published: (2024)
by: Mathur, Leena, et al.
Published: (2024)
A Design-based Solution for Causal Inference with Text: Can a Language Model Be Too Large?
by: Tierney, Graham, et al.
Published: (2025)
by: Tierney, Graham, et al.
Published: (2025)
MPO: An Efficient Post-Processing Framework for Mixing Diverse Preference Alignment
by: Wang, Tianze, et al.
Published: (2025)
by: Wang, Tianze, et al.
Published: (2025)
Causal Imitation Learning Under Measurement Error and Distribution Shift
by: Bo, Shi, et al.
Published: (2026)
by: Bo, Shi, et al.
Published: (2026)
IoT-LM: Large Multisensory Language Models for the Internet of Things
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
Annotation Sensitivity: Training Data Collection Methods Affect Model Performance
by: Kern, Christoph, et al.
Published: (2023)
by: Kern, Christoph, et al.
Published: (2023)
Evaluating Interventional Reasoning Capabilities of Large Language Models
by: Kasetty, Tejas, et al.
Published: (2024)
by: Kasetty, Tejas, et al.
Published: (2024)
Estimate Level Adjustment For Inference With Proxies Under Random Distribution Shifts
by: Wilkins-Reeves, Steven, et al.
Published: (2026)
by: Wilkins-Reeves, Steven, et al.
Published: (2026)
Debiasing Watermarks for Large Language Models via Maximal Coupling
by: Xie, Yangxinyu, et al.
Published: (2024)
by: Xie, Yangxinyu, et al.
Published: (2024)
CLEAR: Can Language Models Really Understand Causal Graphs?
by: Chen, Sirui, et al.
Published: (2024)
by: Chen, Sirui, et al.
Published: (2024)
Enhancing Causal Reasoning in Large Language Models: A Causal Attribution Model for Precision Fine-Tuning
by: Cai, Hengrui, et al.
Published: (2023)
by: Cai, Hengrui, et al.
Published: (2023)
Random Text, Zipf's Law, Critical Length,and Implications for Large Language Models
by: Berman, Vladimir
Published: (2025)
by: Berman, Vladimir
Published: (2025)
Quantifying Omitted Variable Bias in Nonlinear Instrumental Variable Estimators
by: Yen, Yu-Min
Published: (2026)
by: Yen, Yu-Min
Published: (2026)
Language Models as Causal Effect Generators
by: Bynum, Lucius E. J., et al.
Published: (2024)
by: Bynum, Lucius E. J., et al.
Published: (2024)
Assessing Omitted Variable Bias when the Controls are Endogenous
by: Diegert, Paul, et al.
Published: (2022)
by: Diegert, Paul, et al.
Published: (2022)
AutoEval Done Right: Using Synthetic Data for Model Evaluation
by: Boyeau, Pierre, et al.
Published: (2024)
by: Boyeau, Pierre, et al.
Published: (2024)
Bayesian Federated Cause-of-Death Classification and Quantification Under Distribution Shift
by: Zhu, Yu, et al.
Published: (2025)
by: Zhu, Yu, et al.
Published: (2025)
LLMs as Implicit Imputers: Uncertainty Should Scale with Missing Information
by: van Buuren, Stef
Published: (2026)
by: van Buuren, Stef
Published: (2026)
Spiking the training data to correct for test set contamination
by: Wei, Johnny Tian-Zheng, et al.
Published: (2026)
by: Wei, Johnny Tian-Zheng, et al.
Published: (2026)
Learning Dynamic Representations and Policies from Multimodal Clinical Time-Series with Informative Missingness
by: Liang, Zihan, et al.
Published: (2026)
by: Liang, Zihan, et al.
Published: (2026)
Causal Representation Learning from Multimodal Clinical Records under Non-Random Modality Missingness
by: Liang, Zihan, et al.
Published: (2025)
by: Liang, Zihan, et al.
Published: (2025)
MuPlon: Multi-Path Causal Optimization for Claim Verification through Controlling Confounding
by: Guo, Hanghui, et al.
Published: (2025)
by: Guo, Hanghui, et al.
Published: (2025)
How to Evaluate Entity Resolution Systems: An Entity-Centric Framework with Application to Inventor Name Disambiguation
by: Binette, Olivier, et al.
Published: (2024)
by: Binette, Olivier, et al.
Published: (2024)
Discovering influential text using convolutional neural networks
by: Ayers, Megan, et al.
Published: (2024)
by: Ayers, Megan, et al.
Published: (2024)
Optimal Estimation of Watermark Proportions in Hybrid AI-Human Texts
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
Similar Items
-
Optimizing Language Models for Human Preferences is a Causal Inference Problem
by: Lin, Victoria, et al.
Published: (2024) -
Isolated Causal Effects of Natural Language
by: Lin, Victoria, et al.
Published: (2024) -
Long Story Short: Omitted Variable Bias in Causal Machine Learning
by: Chernozhukov, Victor, et al.
Published: (2021) -
Zipf Distributions from Two-Stage Symbolic Processes: Stability Under Stochastic Lexical Filtering
by: Berman, Vladimir
Published: (2025) -
Robust Detection of Watermarks for Large Language Models Under Human Edits
by: Li, Xiang, et al.
Published: (2024)