Identifying Evidence-Based Nudges in Biomedical Literature with Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chauhan, Jaydeep, Seidman, Mark, Parvari, Pezhman Raeisian, Zheng, Zhi, Ben-Miled, Zina, Barboi, Cristina, Gonzalez, Andrew, Boustani, Malaz
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911439610970112
author Chauhan, Jaydeep
Seidman, Mark
Parvari, Pezhman Raeisian
Zheng, Zhi
Ben-Miled, Zina
Barboi, Cristina
Gonzalez, Andrew
Boustani, Malaz
author_facet Chauhan, Jaydeep
Seidman, Mark
Parvari, Pezhman Raeisian
Zheng, Zhi
Ben-Miled, Zina
Barboi, Cristina
Gonzalez, Andrew
Boustani, Malaz
contents We present a scalable, AI-powered system that identifies and extracts evidence-based behavioral nudges from unstructured biomedical literature. Nudges are subtle, non-coercive interventions that influence behavior without limiting choice, showing strong impact on health outcomes like medication adherence. However, identifying these interventions from PubMed's 8 million+ articles is a bottleneck. Our system uses a novel multi-stage pipeline: first, hybrid filtering (keywords, TF-IDF, cosine similarity, and a "nudge-term bonus") reduces the corpus to about 81,000 candidates. Second, we use OpenScholar (quantized LLaMA 3.1 8B) to classify papers and extract structured fields like nudge type and target behavior in a single pass, validated against a JSON schema. We evaluated four configurations on a labeled test set (N=197). The best setup (Title/Abstract/Intro) achieved a 67.0% F1 score and 72.0% recall, ideal for discovery. A high-precision variant using self-consistency (7 randomized passes) achieved 100% precision with 12% recall, demonstrating a tunable trade-off for high-trust use cases. This system is being integrated into Agile Nudge+, a real-world platform, to ground LLM-generated interventions in peer-reviewed evidence. This work demonstrates interpretable, domain-specific retrieval pipelines for evidence synthesis and personalized healthcare.
format Preprint
id arxiv_https___arxiv_org_abs_2602_10345
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Identifying Evidence-Based Nudges in Biomedical Literature with Large Language Models
Chauhan, Jaydeep
Seidman, Mark
Parvari, Pezhman Raeisian
Zheng, Zhi
Ben-Miled, Zina
Barboi, Cristina
Gonzalez, Andrew
Boustani, Malaz
Machine Learning
We present a scalable, AI-powered system that identifies and extracts evidence-based behavioral nudges from unstructured biomedical literature. Nudges are subtle, non-coercive interventions that influence behavior without limiting choice, showing strong impact on health outcomes like medication adherence. However, identifying these interventions from PubMed's 8 million+ articles is a bottleneck. Our system uses a novel multi-stage pipeline: first, hybrid filtering (keywords, TF-IDF, cosine similarity, and a "nudge-term bonus") reduces the corpus to about 81,000 candidates. Second, we use OpenScholar (quantized LLaMA 3.1 8B) to classify papers and extract structured fields like nudge type and target behavior in a single pass, validated against a JSON schema. We evaluated four configurations on a labeled test set (N=197). The best setup (Title/Abstract/Intro) achieved a 67.0% F1 score and 72.0% recall, ideal for discovery. A high-precision variant using self-consistency (7 randomized passes) achieved 100% precision with 12% recall, demonstrating a tunable trade-off for high-trust use cases. This system is being integrated into Agile Nudge+, a real-world platform, to ground LLM-generated interventions in peer-reviewed evidence. This work demonstrates interpretable, domain-specific retrieval pipelines for evidence synthesis and personalized healthcare.
title Identifying Evidence-Based Nudges in Biomedical Literature with Large Language Models
topic Machine Learning
url https://arxiv.org/abs/2602.10345