From Conversation to Query Execution: Benchmarking User and Tool Interactions for EHR Database Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Gyubok, Chay, Woosog, Kwak, Heeyoung, Kim, Yeong Hwa, Yoo, Haanju, Jeong, Oksoon, Son, Meong Hi, Choi, Edward |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SCARE: A Benchmark for SQL Correction and Question Answerability Classification for Reliable EHR Question Answering
by: Lee, Gyubok, et al.
Published: (2025)
by: Lee, Gyubok, et al.
Published: (2025)
TrustSQL: Benchmarking Text-to-SQL Reliability with Penalty-Based Scoring
by: Lee, Gyubok, et al.
Published: (2024)
by: Lee, Gyubok, et al.
Published: (2024)
H-AdminSim: A Multi-Agent Simulator for Realistic Hospital Administrative Workflows with FHIR Integration
by: Lee, Jun-Min, et al.
Published: (2026)
by: Lee, Jun-Min, et al.
Published: (2026)
EHR-SeqSQL : A Sequential Text-to-SQL Dataset For Interactively Exploring Electronic Health Records
by: Ryu, Jaehee, et al.
Published: (2024)
by: Ryu, Jaehee, et al.
Published: (2024)
ProPerSim: Developing Proactive and Personalized AI Assistants through User-Assistant Simulation
by: Kim, Jiho, et al.
Published: (2025)
by: Kim, Jiho, et al.
Published: (2025)
Doppelganger Method: Breaking Role Consistency in LLM Agent via Prompt-based Transferable Adversarial Attack
by: Kang, Daewon, et al.
Published: (2025)
by: Kang, Daewon, et al.
Published: (2025)
DialSim: A Dialogue Simulator for Evaluating Long-Term Multi-Party Dialogue Understanding of Conversational Agents
by: Kim, Jiho, et al.
Published: (2024)
by: Kim, Jiho, et al.
Published: (2024)
FHIR-AgentBench: Benchmarking LLM Agents for Realistic Interoperable EHR Question Answering
by: Lee, Gyubok, et al.
Published: (2025)
by: Lee, Gyubok, et al.
Published: (2025)
Multi-lingual Multi-institutional Electronic Health Record based Predictive Model
by: Hur, Kyunghoon, et al.
Published: (2026)
by: Hur, Kyunghoon, et al.
Published: (2026)
Overview of the EHRSQL 2024 Shared Task on Reliable Text-to-SQL Modeling on Electronic Health Records
by: Lee, Gyubok, et al.
Published: (2024)
by: Lee, Gyubok, et al.
Published: (2024)
EHRNoteQA: An LLM Benchmark for Real-World Clinical Practice Using Discharge Summaries
by: Kweon, Sunjun, et al.
Published: (2024)
by: Kweon, Sunjun, et al.
Published: (2024)
CANVAS: A Benchmark for Vision-Language Models on Tool-Based User Interface Design
by: Jeong, Daeheon, et al.
Published: (2025)
by: Jeong, Daeheon, et al.
Published: (2025)
Comparison of the Parkinson Anxiety Scale and Other Tools to Screen Anxiety in Patients With Parkinson's Disease
by: Seong‐Hi Park
Published: (2024)
by: Seong‐Hi Park
Published: (2024)
Query Generation Pipeline with Enhanced Answerability Assessment for Financial Information Retrieval
by: Kim, Hyunkyu, et al.
Published: (2025)
by: Kim, Hyunkyu, et al.
Published: (2025)
CoX-MoE: Coalesced Expert Execution for High-Throughput MoE Inference with AMX-Enabled CPU-GPU Co-Execution
by: Son, Muyoung, et al.
Published: (2026)
by: Son, Muyoung, et al.
Published: (2026)
Synthesizing Document Database Queries using Collection Abstractions
by: Liu, Qikang, et al.
Published: (2024)
by: Liu, Qikang, et al.
Published: (2024)
Routing End User Queries to Enterprise Databases
by: Sudarshan, Saikrishna, et al.
Published: (2026)
by: Sudarshan, Saikrishna, et al.
Published: (2026)
Uniqueness and Existence of Linear Equilibrium with a Constrained Trader
by: Kwon, Heeyoung, et al.
Published: (2025)
by: Kwon, Heeyoung, et al.
Published: (2025)
Joint streaming model for backchannel prediction and automatic speech recognition
by: Yong‐Seok Choi, et al.
Published: (2024)
by: Yong‐Seok Choi, et al.
Published: (2024)
Moisture‐Tolerant, Lithiophilic Artificial Solid Electrolyte Interphase Enables Ambient‐Processable Lithium Metal Anodes
by: Yeong Hun Jeong, et al.
Published: (2025)
by: Yeong Hun Jeong, et al.
Published: (2025)
EHR Interaction Between Patients and AI: NoteAid EHR Interaction
by: Zhang, Xiaocheng, et al.
Published: (2023)
by: Zhang, Xiaocheng, et al.
Published: (2023)
DBRouting: Routing End User Queries to Databases for Answerability
by: Mandal, Priyangshu, et al.
Published: (2025)
by: Mandal, Priyangshu, et al.
Published: (2025)
OCLC's Database Conversion: A User's Perspective.
by: Wajenberg, Arnold, et al.
Published: (1981)
by: Wajenberg, Arnold, et al.
Published: (1981)
EHRSQL: A Practical Text-to-SQL Benchmark for Electronic Health Records
by: Lee, Gyubok, et al.
Published: (2023)
by: Lee, Gyubok, et al.
Published: (2023)
Using LLMs to Investigate Correlations of Conversational Follow-up Queries with User Satisfaction
by: Kim, Hyunwoo, et al.
Published: (2024)
by: Kim, Hyunwoo, et al.
Published: (2024)
VoicePrompter: Robust Zero-Shot Voice Conversion with Voice Prompt and Conditional Flow Matching
by: Choi, Ha-Yeong, et al.
Published: (2025)
by: Choi, Ha-Yeong, et al.
Published: (2025)
Oligohistidine‐Functionalized Single‐Walled Carbon Nanotube‐Guided RNA Delivery to Improve Shoot Regeneration Efficiency in Plant Calli
by: Yeong Yeop Jeong, et al.
Published: (2025)
by: Yeong Yeop Jeong, et al.
Published: (2025)
Database-Augmented Query Representation for Information Retrieval
by: Jeong, Soyeong, et al.
Published: (2024)
by: Jeong, Soyeong, et al.
Published: (2024)
Expression Patterns of Deubiquitinating Enzymes in Paclitaxel‐Treated Lung Cancer Cells
by: Hwa‐Yeong Kim, et al.
Published: (2025)
by: Hwa‐Yeong Kim, et al.
Published: (2025)
An Executable Benchmarking Suite for Tool-Using Agents
by: Zhong, Zhiqing, et al.
Published: (2026)
by: Zhong, Zhiqing, et al.
Published: (2026)
ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
by: Lu, Jiarui, et al.
Published: (2024)
by: Lu, Jiarui, et al.
Published: (2024)
QueryGym: Step-by-Step Interaction with Relational Databases
by: Ananthakrishnan, Haritha, et al.
Published: (2025)
by: Ananthakrishnan, Haritha, et al.
Published: (2025)
Querying Databases with Function Calling
by: Shorten, Connor, et al.
Published: (2025)
by: Shorten, Connor, et al.
Published: (2025)
Searching Databases without Query-Building Aids: Implications for Dyslexic Users
by: Berget, Gerd, et al.
Published: (2015)
by: Berget, Gerd, et al.
Published: (2015)
Perturb-and-Compare Approach for Detecting Out-of-Distribution Samples in Constrained Access Environments
by: Lee, Heeyoung, et al.
Published: (2024)
by: Lee, Heeyoung, et al.
Published: (2024)
AMBROSIA: A Benchmark for Parsing Ambiguous Questions into Database Queries
by: Saparina, Irina, et al.
Published: (2024)
by: Saparina, Irina, et al.
Published: (2024)
Towards Unbiased Evaluation of Detecting Unanswerable Questions in EHRSQL
by: Yang, Yongjin, et al.
Published: (2024)
by: Yang, Yongjin, et al.
Published: (2024)
Requerimientos energéticos de ovinos de pelo en las regiones tropicales de Latinoamérica. Revisión
by: Alfonso Juventino Chay Canul
Published: (2016)
by: Alfonso Juventino Chay Canul
Published: (2016)
Unlocking Health Insights with SDoH Data: A Comprehensive Open-Access Database and SDoH-EHR Linkage Tool
by: Hu, Zhenhong, et al.
Published: (2025)
by: Hu, Zhenhong, et al.
Published: (2025)
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache
by: Son, Donghyun, et al.
Published: (2025)
by: Son, Donghyun, et al.
Published: (2025)
Similar Items
-
SCARE: A Benchmark for SQL Correction and Question Answerability Classification for Reliable EHR Question Answering
by: Lee, Gyubok, et al.
Published: (2025) -
TrustSQL: Benchmarking Text-to-SQL Reliability with Penalty-Based Scoring
by: Lee, Gyubok, et al.
Published: (2024) -
H-AdminSim: A Multi-Agent Simulator for Realistic Hospital Administrative Workflows with FHIR Integration
by: Lee, Jun-Min, et al.
Published: (2026) -
EHR-SeqSQL : A Sequential Text-to-SQL Dataset For Interactively Exploring Electronic Health Records
by: Ryu, Jaehee, et al.
Published: (2024) -
ProPerSim: Developing Proactive and Personalized AI Assistants through User-Assistant Simulation
by: Kim, Jiho, et al.
Published: (2025)