Can LLM Agents Generate Real-World Evidence? Evaluating Observational Studies in Medical Databases
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Dubai, He, Yuxiang, Hu, Yan, Tian, Yu, Li, Jingsong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Foundation Models to Unlock Real-World Evidence from Nationwide Medical Claims
von: Ma, Fan, et al.
Veröffentlicht: (2026)
von: Ma, Fan, et al.
Veröffentlicht: (2026)
Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence
von: Dong, Guanting, et al.
Veröffentlicht: (2026)
von: Dong, Guanting, et al.
Veröffentlicht: (2026)
Wide-Horizon Thinking and Simulation-Based Evaluation for Real-World LLM Planning with Multifaceted Constraints
von: Yang, Dongjie, et al.
Veröffentlicht: (2025)
von: Yang, Dongjie, et al.
Veröffentlicht: (2025)
Streamlining evidence based clinical recommendations with large language models
von: Li, Dubai, et al.
Veröffentlicht: (2025)
von: Li, Dubai, et al.
Veröffentlicht: (2025)
FinMCP-Bench: Benchmarking LLM Agents for Real-World Financial Tool Use under the Model Context Protocol
von: Zhu, Jie, et al.
Veröffentlicht: (2026)
von: Zhu, Jie, et al.
Veröffentlicht: (2026)
Beyond Perfect APIs: A Comprehensive Evaluation of LLM Agents Under Real-World API Complexity
von: Kim, Doyoung, et al.
Veröffentlicht: (2026)
von: Kim, Doyoung, et al.
Veröffentlicht: (2026)
Can Large Language Models Understand Real-World Complex Instructions?
von: He, Qianyu, et al.
Veröffentlicht: (2023)
von: He, Qianyu, et al.
Veröffentlicht: (2023)
GroundAct: Can LLM Agents Ground Actions in Environmental States?
von: Wang, Zixuan, et al.
Veröffentlicht: (2025)
von: Wang, Zixuan, et al.
Veröffentlicht: (2025)
OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation
von: Hu, Mengkang, et al.
Veröffentlicht: (2025)
von: Hu, Mengkang, et al.
Veröffentlicht: (2025)
COMAP: Co-Evolving World Models and Agent Policies for LLM Agents
von: Liu, Youwei, et al.
Veröffentlicht: (2026)
von: Liu, Youwei, et al.
Veröffentlicht: (2026)
Agent2World: Learning to Generate Symbolic World Models via Adaptive Multi-Agent Feedback
von: Hu, Mengkang, et al.
Veröffentlicht: (2025)
von: Hu, Mengkang, et al.
Veröffentlicht: (2025)
Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions
von: Hu, Yuanzhe, et al.
Veröffentlicht: (2025)
von: Hu, Yuanzhe, et al.
Veröffentlicht: (2025)
LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks
von: Long, Xiang, et al.
Veröffentlicht: (2026)
von: Long, Xiang, et al.
Veröffentlicht: (2026)
MobileBench-OL: A Comprehensive Chinese Benchmark for Evaluating Mobile GUI Agents in Real-World Environment
von: Wu, Qinzhuo, et al.
Veröffentlicht: (2026)
von: Wu, Qinzhuo, et al.
Veröffentlicht: (2026)
ConsintBench: Evaluating Language Models on Real-World Consumer Intent Understanding
von: Li, Xiaozhe, et al.
Veröffentlicht: (2025)
von: Li, Xiaozhe, et al.
Veröffentlicht: (2025)
CoreEval: Automatically Building Contamination-Resilient Datasets with Real-World Knowledge toward Reliable LLM Evaluation
von: Zhao, Jingqian, et al.
Veröffentlicht: (2025)
von: Zhao, Jingqian, et al.
Veröffentlicht: (2025)
AD-Bench: A Real-World, Trajectory-Aware Advertising Analytics Benchmark for LLM Agents
von: Hu, Lingxiang, et al.
Veröffentlicht: (2026)
von: Hu, Lingxiang, et al.
Veröffentlicht: (2026)
From Physician Expertise to Clinical Agents: Preserving, Standardizing, and Scaling Physicians' Medical Expertise with Lightweight LLM
von: Luo, Chanyong, et al.
Veröffentlicht: (2026)
von: Luo, Chanyong, et al.
Veröffentlicht: (2026)
ClinicalLab: Aligning Agents for Multi-Departmental Clinical Diagnostics in the Real World
von: Yan, Weixiang, et al.
Veröffentlicht: (2024)
von: Yan, Weixiang, et al.
Veröffentlicht: (2024)
On the Role of Model Prior in Real-World Inductive Reasoning
von: Liu, Zhuo, et al.
Veröffentlicht: (2024)
von: Liu, Zhuo, et al.
Veröffentlicht: (2024)
DetectRL: Benchmarking LLM-Generated Text Detection in Real-World Scenarios
von: Wu, Junchao, et al.
Veröffentlicht: (2024)
von: Wu, Junchao, et al.
Veröffentlicht: (2024)
OpenWebVoyager: Building Multimodal Web Agents via Iterative Real-World Exploration, Feedback and Optimization
von: He, Hongliang, et al.
Veröffentlicht: (2024)
von: He, Hongliang, et al.
Veröffentlicht: (2024)
Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering Tasks with Generative Optimization
von: Chi, Yizhe, et al.
Veröffentlicht: (2026)
von: Chi, Yizhe, et al.
Veröffentlicht: (2026)
Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents
von: Deng, Shihan, et al.
Veröffentlicht: (2024)
von: Deng, Shihan, et al.
Veröffentlicht: (2024)
MIRAI: Evaluating LLM Agents for Event Forecasting
von: Ye, Chenchen, et al.
Veröffentlicht: (2024)
von: Ye, Chenchen, et al.
Veröffentlicht: (2024)
World-Model-Augmented Web Agents with Action Correction
von: Shen, Zhouzhou, et al.
Veröffentlicht: (2026)
von: Shen, Zhouzhou, et al.
Veröffentlicht: (2026)
How to Detect and Defeat Molecular Mirage: A Metric-Driven Benchmark for Hallucination in LLM-based Molecular Comprehension
von: Li, Hao, et al.
Veröffentlicht: (2025)
von: Li, Hao, et al.
Veröffentlicht: (2025)
Predicting the Big Five Personality Traits in Chinese Counselling Dialogues Using Large Language Models
von: Yan, Yang, et al.
Veröffentlicht: (2024)
von: Yan, Yang, et al.
Veröffentlicht: (2024)
From Scores to Steps: Diagnosing and Improving LLM Performance in Evidence-Based Medical Calculations
von: Wang, Benlu, et al.
Veröffentlicht: (2025)
von: Wang, Benlu, et al.
Veröffentlicht: (2025)
Can Generative Agents Predict Emotion?
von: Regan, Ciaran, et al.
Veröffentlicht: (2024)
von: Regan, Ciaran, et al.
Veröffentlicht: (2024)
DSAEval: Evaluating Data Science Agents on a Wide Range of Real-World Data Science Problems
von: Sun, Maojun, et al.
Veröffentlicht: (2026)
von: Sun, Maojun, et al.
Veröffentlicht: (2026)
Apollo: A Lightweight Multilingual Medical LLM towards Democratizing Medical AI to 6B People
von: Wang, Xidong, et al.
Veröffentlicht: (2024)
von: Wang, Xidong, et al.
Veröffentlicht: (2024)
SearchGym: Bootstrapping Real-World Search Agents via Cost-Effective and High-Fidelity Environment Simulation
von: Zhang, Xichen, et al.
Veröffentlicht: (2026)
von: Zhang, Xichen, et al.
Veröffentlicht: (2026)
Locomo-Plus: Beyond-Factual Cognitive Memory Evaluation Framework for LLM Agents
von: Li, Yifei, et al.
Veröffentlicht: (2026)
von: Li, Yifei, et al.
Veröffentlicht: (2026)
Synthetic Dialogue Dataset Generation using LLM Agents
von: Abdullin, Yelaman, et al.
Veröffentlicht: (2024)
von: Abdullin, Yelaman, et al.
Veröffentlicht: (2024)
Inclusion Arena: An Open Platform for Evaluating Large Foundation Models with Real-World Apps
von: Wang, Kangyu, et al.
Veröffentlicht: (2025)
von: Wang, Kangyu, et al.
Veröffentlicht: (2025)
RealFactBench: A Benchmark for Evaluating Large Language Models in Real-World Fact-Checking
von: Yang, Shuo, et al.
Veröffentlicht: (2025)
von: Yang, Shuo, et al.
Veröffentlicht: (2025)
LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset
von: Zheng, Lianmin, et al.
Veröffentlicht: (2023)
von: Zheng, Lianmin, et al.
Veröffentlicht: (2023)
RWESummary: A Framework and Test for Choosing Large Language Models to Summarize Real-World Evidence (RWE) Studies
von: Mukerji, Arjun, et al.
Veröffentlicht: (2025)
von: Mukerji, Arjun, et al.
Veröffentlicht: (2025)
Can LLMs Generate Reliable Test Case Generators? A Study on Competition-Level Programming Problems
von: Cao, Yuhan, et al.
Veröffentlicht: (2025)
von: Cao, Yuhan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Foundation Models to Unlock Real-World Evidence from Nationwide Medical Claims
von: Ma, Fan, et al.
Veröffentlicht: (2026) -
Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence
von: Dong, Guanting, et al.
Veröffentlicht: (2026) -
Wide-Horizon Thinking and Simulation-Based Evaluation for Real-World LLM Planning with Multifaceted Constraints
von: Yang, Dongjie, et al.
Veröffentlicht: (2025) -
Streamlining evidence based clinical recommendations with large language models
von: Li, Dubai, et al.
Veröffentlicht: (2025) -
FinMCP-Bench: Benchmarking LLM Agents for Real-World Financial Tool Use under the Model Context Protocol
von: Zhu, Jie, et al.
Veröffentlicht: (2026)