BPQA Dataset: Evaluating How Well Language Models Leverage Blood Pressures to Answer Biomedical Questions
Fuente:
arXiv
Saved in:
| Main Authors: | Hang, Chi, Deng, Ruiqi, Jiang, Lavender Yao, Yang, Zihao, Alyakin, Anton, Alber, Daniel, Oermann, Eric Karl |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MedMobile: A mobile-sized language model with clinical capabilities
by: Vishwanath, Krithik, et al.
Published: (2024)
by: Vishwanath, Krithik, et al.
Published: (2024)
Generalist Large Language Models Outperform Clinical Tools on Medical Benchmarks
by: Vishwanath, Krithik, et al.
Published: (2025)
by: Vishwanath, Krithik, et al.
Published: (2025)
Medical large language models are easily distracted
by: Vishwanath, Krithik, et al.
Published: (2025)
by: Vishwanath, Krithik, et al.
Published: (2025)
Refining Packing and Shuffling Strategies for Enhanced Performance in Generative Language Models
by: Chen, Yanbing, et al.
Published: (2024)
by: Chen, Yanbing, et al.
Published: (2024)
Evaluating the performance and fragility of large language models on the self-assessment for neurological surgeons
by: Vishwanath, Krithik, et al.
Published: (2025)
by: Vishwanath, Krithik, et al.
Published: (2025)
Generalization in Healthcare AI: Evaluation of a Clinical Large Language Model
by: Rahman, Salman, et al.
Published: (2024)
by: Rahman, Salman, et al.
Published: (2024)
It is Too Many Options: Pitfalls of Multiple-Choice Questions in Generative AI and Medical Education
by: Singh, Shrutika, et al.
Published: (2025)
by: Singh, Shrutika, et al.
Published: (2025)
"All You Need" is Not All You Need for a Paper Title: On the Origins of a Scientific Meme
by: Alyakin, Anton
Published: (2025)
by: Alyakin, Anton
Published: (2025)
MedG-KRP: Medical Graph Knowledge Representation Probing
by: Rosenbaum, Gabriel R., et al.
Published: (2024)
by: Rosenbaum, Gabriel R., et al.
Published: (2024)
Paradox of De-identification: A Critique of HIPAA Safe Harbour in the Age of LLMs
by: Jiang, Lavender Y., et al.
Published: (2026)
by: Jiang, Lavender Y., et al.
Published: (2026)
Using Pretrained Large Language Model with Prompt Engineering to Answer Biomedical Questions
by: Zhou, Wenxin, et al.
Published: (2024)
by: Zhou, Wenxin, et al.
Published: (2024)
Generalist Foundation Models Are Not Clinical Enough for Hospital Operations
by: Jiang, Lavender Y., et al.
Published: (2025)
by: Jiang, Lavender Y., et al.
Published: (2025)
On the Relationship Between the Choice of Representation and In-Context Learning
by: Marinescu, Ioana, et al.
Published: (2025)
by: Marinescu, Ioana, et al.
Published: (2025)
How do you know that? Teaching Generative Language Models to Reference Answers to Biomedical Questions
by: Bašaragin, Bojana, et al.
Published: (2024)
by: Bašaragin, Bojana, et al.
Published: (2024)
BioPIE: A Biomedical Protocol Information Extraction Dataset for High-Reasoning-Complexity Experiment Question Answer
by: Hou, Haofei, et al.
Published: (2026)
by: Hou, Haofei, et al.
Published: (2026)
What's More Important: The Questions or the Answers?
by: Morgan, Eric Lease
Published: (1999)
by: Morgan, Eric Lease
Published: (1999)
Large Language Models Predict Functional Outcomes after Acute Ischemic Stroke
by: Kapoor, Anjali K., et al.
Published: (2026)
by: Kapoor, Anjali K., et al.
Published: (2026)
When Answers Stray from Questions: Hallucination Detection via Question-Answer Orthogonal Decomposition
by: Yao, Siyang, et al.
Published: (2026)
by: Yao, Siyang, et al.
Published: (2026)
Lebensgeschichten alter Eltern kognitiv beeinträchtigter Menschen
by: Oermann, Lisa
Published: (2023)
by: Oermann, Lisa
Published: (2023)
To Answer or to Refuse? Investigating the Effect of Refusal to Answer Privacy‐Invasive Question on Applicants' Perceived Hireability
by: Wanlu Li, et al.
Published: (2024)
by: Wanlu Li, et al.
Published: (2024)
Clinically Grounded Agent-based Report Evaluation: An Interpretable Metric for Radiology Report Generation
by: Dua, Radhika, et al.
Published: (2025)
by: Dua, Radhika, et al.
Published: (2025)
Answer, Assemble, Ace: Understanding How LMs Answer Multiple Choice Questions
by: Wiegreffe, Sarah, et al.
Published: (2024)
by: Wiegreffe, Sarah, et al.
Published: (2024)
Gateformer: Advancing Multivariate Time Series Forecasting through Temporal and Variate-Wise Attention with Gated Representations
by: Lan, Yu-Hsiang, et al.
Published: (2025)
by: Lan, Yu-Hsiang, et al.
Published: (2025)
BioACE: An Automated Framework for Biomedical Answer and Citation Evaluations
by: Gupta, Deepak, et al.
Published: (2026)
by: Gupta, Deepak, et al.
Published: (2026)
No Answer Needed: Predicting LLM Answer Accuracy from Question-Only Linear Probes
by: Cencerrado, Iván Vicente Moreno, et al.
Published: (2025)
by: Cencerrado, Iván Vicente Moreno, et al.
Published: (2025)
EduAdapt: A Question Answer Benchmark Dataset for Evaluating Grade-Level Adaptability in LLMs
by: Naeem, Numaan, et al.
Published: (2025)
by: Naeem, Numaan, et al.
Published: (2025)
A Study on Question-Answer Dataset for LLM Safety Evaluation with a Focus on Illegal Activities
by: Imamura, Kenji, et al.
Published: (2026)
by: Imamura, Kenji, et al.
Published: (2026)
Social Question and Answer Services versus Library Virtual Reference: Evaluation and Comparison from the Users' Perspective
by: Zhang, Yin, et al.
Published: (2014)
by: Zhang, Yin, et al.
Published: (2014)
Rehearsing Answers to Probable Questions with Perspective-Taking
by: Shih, Yung-Yu, et al.
Published: (2024)
by: Shih, Yung-Yu, et al.
Published: (2024)
From Answers to Questions: EQGBench for Evaluating LLMs' Educational Question Generation
by: Zhou, Chengliang, et al.
Published: (2025)
by: Zhou, Chengliang, et al.
Published: (2025)
An Empirical Study on How Video-LLMs Answer Video Questions
by: Gou, Chenhui, et al.
Published: (2025)
by: Gou, Chenhui, et al.
Published: (2025)
Question: How do Large Language Models perform on the Question Answering tasks? Answer:
by: Fischer, Kevin, et al.
Published: (2024)
by: Fischer, Kevin, et al.
Published: (2024)
Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?
by: Balepur, Nishant, et al.
Published: (2024)
by: Balepur, Nishant, et al.
Published: (2024)
MiRAGE: A Multiagent Framework for Generating Multimodal Multihop Question-Answer Dataset for RAG Evaluation
by: Sahu, Chandan Kumar, et al.
Published: (2026)
by: Sahu, Chandan Kumar, et al.
Published: (2026)
Subjective Question Generation and Answer Evaluation using NLP
by: Islam, G. M. Refatul, et al.
Published: (2025)
by: Islam, G. M. Refatul, et al.
Published: (2025)
What Questions Should Robots Be Able to Answer? A Dataset of User Questions for Explainable Robotics
by: Wachowiak, Lennart, et al.
Published: (2025)
by: Wachowiak, Lennart, et al.
Published: (2025)
A Dataset of Open-Domain Question Answering with Multiple-Span Answers
by: Luo, Zhiyi, et al.
Published: (2024)
by: Luo, Zhiyi, et al.
Published: (2024)
Questions, Answers, and Presuppositions
by: Marie Duží
Published: (2015)
by: Marie Duží
Published: (2015)
Correcting a Nonparametric Two-sample Graph Hypothesis Test for Graphs with Different Numbers of Vertices with Applications to Connectomics
by: Alyakin, Anton A., et al.
Published: (2020)
by: Alyakin, Anton A., et al.
Published: (2020)
GETALP@AutoMin 2025: Leveraging RAG to Answer Questions based on Meeting Transcripts
by: Kang, Jeongwoo, et al.
Published: (2025)
by: Kang, Jeongwoo, et al.
Published: (2025)
Similar Items
-
MedMobile: A mobile-sized language model with clinical capabilities
by: Vishwanath, Krithik, et al.
Published: (2024) -
Generalist Large Language Models Outperform Clinical Tools on Medical Benchmarks
by: Vishwanath, Krithik, et al.
Published: (2025) -
Medical large language models are easily distracted
by: Vishwanath, Krithik, et al.
Published: (2025) -
Refining Packing and Shuffling Strategies for Enhanced Performance in Generative Language Models
by: Chen, Yanbing, et al.
Published: (2024) -
Evaluating the performance and fragility of large language models on the self-assessment for neurological surgeons
by: Vishwanath, Krithik, et al.
Published: (2025)