Saved in:
Bibliographic Details
Main Authors: Skinner, Caleb, Guo, Yihan, Li, Meng
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2605.23102
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911707397357568
author Skinner, Caleb
Guo, Yihan
Li, Meng
author_facet Skinner, Caleb
Guo, Yihan
Li, Meng
contents Large language models (LLMs) offer a scalable mechanism to elicit domain-informed prior information for high-dimensional variable selection. However, existing methods such as LLM-Lasso are sensitive to weight quality, with performance degrading substantially when LLM-generated weights are inaccurate. To address this challenge, we first introduce a framework for quantifying the quality of LLM-generated weights, enabling rigorous evaluation of LLM-informed methods across varying weight regimes. We then propose the LLM Sparsity Prior (LSP), which integrates LLM-generated weights into the prior inclusion probabilities of Spike-and-Slab and Spike-and-Slab Lasso models via two interpretable hyperparameters governing global sparsity and weight concentration. Hierarchical hyperpriors on these parameters allow the model to dynamically discount uninformative or misleading weights, improving robustness without sacrificing gains when weights are accurate. Finally, we develop principled prompt engineering strategies and validate the method on a private medical dataset studying Acute Kidney Injury. LSP improves prediction accuracy and identifies clinically relevant features missed by the baselines, with robustness to prompt variation and particular effectiveness in low-data regimes.
format Preprint
id arxiv_https___arxiv_org_abs_2605_23102
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LLM Sparsity Prior for Robust Feature Selection
Skinner, Caleb
Guo, Yihan
Li, Meng
Machine Learning
Methodology
Large language models (LLMs) offer a scalable mechanism to elicit domain-informed prior information for high-dimensional variable selection. However, existing methods such as LLM-Lasso are sensitive to weight quality, with performance degrading substantially when LLM-generated weights are inaccurate. To address this challenge, we first introduce a framework for quantifying the quality of LLM-generated weights, enabling rigorous evaluation of LLM-informed methods across varying weight regimes. We then propose the LLM Sparsity Prior (LSP), which integrates LLM-generated weights into the prior inclusion probabilities of Spike-and-Slab and Spike-and-Slab Lasso models via two interpretable hyperparameters governing global sparsity and weight concentration. Hierarchical hyperpriors on these parameters allow the model to dynamically discount uninformative or misleading weights, improving robustness without sacrificing gains when weights are accurate. Finally, we develop principled prompt engineering strategies and validate the method on a private medical dataset studying Acute Kidney Injury. LSP improves prediction accuracy and identifies clinically relevant features missed by the baselines, with robustness to prompt variation and particular effectiveness in low-data regimes.
title LLM Sparsity Prior for Robust Feature Selection
topic Machine Learning
Methodology
url https://arxiv.org/abs/2605.23102