RealICU: Do LLM Agents Understand Long-Context ICU Data? A Benchmark Beyond Behavior Imitation
Fuente:
arXiv
Saved in:
| Main Authors: | Shen, Chengzhi, Shen, Weixiang, Susetzky, Tobias, Chen, Li, Jun, Liu, Yuyuan, Zhang, Xuepeng, Gong, Zhenyu, Rueckert, Daniel, Pan, Jiazhen |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Breaking the Pre-Planning Barrier: Adaptive Real-Time Coordination of Heterogeneous UAVs
by: Hu, Yuhan, et al.
Published: (2025)
by: Hu, Yuhan, et al.
Published: (2025)
Beyond Curve Fitting: Neuro-Symbolic Agents for Context-Aware Epidemic Forecasting
by: Chae, Joongwon, et al.
Published: (2025)
by: Chae, Joongwon, et al.
Published: (2025)
AgentWebBench: Benchmarking Multi-Agent Coordination in Agentic Web
by: Zhong, Shanshan, et al.
Published: (2026)
by: Zhong, Shanshan, et al.
Published: (2026)
Personality-Driven Student Agent-Based Modeling in Mathematics Education: How Well Do Student Agents Align with Human Learners?
by: Xiao, Bushi, et al.
Published: (2026)
by: Xiao, Bushi, et al.
Published: (2026)
Mean Field Correlated Imitation Learning
by: Zhao, Zhiyu, et al.
Published: (2024)
by: Zhao, Zhiyu, et al.
Published: (2024)
Q-ITAGS: Quality-Optimized Spatio-Temporal Heterogeneous Task Allocation with a Time Budget
by: Neville, Glen, et al.
Published: (2024)
by: Neville, Glen, et al.
Published: (2024)
Hierarchical Imitation Learning of Team Behavior from Heterogeneous Demonstrations
by: Seo, Sangwon, et al.
Published: (2025)
by: Seo, Sangwon, et al.
Published: (2025)
Understanding the Information Propagation Effects of Communication Topologies in LLM-based Multi-Agent Systems
by: Shen, Xu, et al.
Published: (2025)
by: Shen, Xu, et al.
Published: (2025)
Work Smarter Not Harder: Simple Imitation Learning with CS-PIBT Outperforms Large Scale Imitation Learning for MAPF
by: Veerapaneni, Rishi, et al.
Published: (2024)
by: Veerapaneni, Rishi, et al.
Published: (2024)
Context-Aware Agent-based Model for Smart Long Distance Transport System
by: Raees, Muhammad, et al.
Published: (2024)
by: Raees, Muhammad, et al.
Published: (2024)
Beyond Single-Agent Alignment: Preventing Context-Fragmented Violations in Multi-Agent Systems
by: Wu, Jie, et al.
Published: (2026)
by: Wu, Jie, et al.
Published: (2026)
YOLO-MARL: You Only LLM Once for Multi-Agent Reinforcement Learning
by: Zhuang, Yuan, et al.
Published: (2024)
by: Zhuang, Yuan, et al.
Published: (2024)
StoryBox: Collaborative Multi-Agent Simulation for Hybrid Bottom-Up Long-Form Story Generation Using Large Language Models
by: Chen, Zehao, et al.
Published: (2025)
by: Chen, Zehao, et al.
Published: (2025)
Agent-Based Modelling for Real-World Stock Markets under Behavioral Economic Principles
by: He, Tianlang, et al.
Published: (2023)
by: He, Tianlang, et al.
Published: (2023)
Understanding Persuasion in Long-Running Agents
by: Jeong, Hyejun, et al.
Published: (2026)
by: Jeong, Hyejun, et al.
Published: (2026)
PREVENT-JACK: Context Steering for Swarms of Long Heavy Articulated Vehicles
by: Baruck, Adrian, et al.
Published: (2026)
by: Baruck, Adrian, et al.
Published: (2026)
Beyond the Individual: Virtualizing Multi-Disciplinary Reasoning for Clinical Intake via Collaborative Agents
by: Chen, Huangwei, et al.
Published: (2026)
by: Chen, Huangwei, et al.
Published: (2026)
Learning to Imitate Spatial Organization in Multi-robot Systems
by: Agunloye, Ayomide O., et al.
Published: (2024)
by: Agunloye, Ayomide O., et al.
Published: (2024)
Beyond Input Guardrails: Reconstructing Cross-Agent Semantic Flows for Execution-Aware Attack Detection
by: Wei, Yangyang, et al.
Published: (2026)
by: Wei, Yangyang, et al.
Published: (2026)
Complex Instruction Following with Diverse Style Policies in Football Games
by: Sun, Chenglu, et al.
Published: (2025)
by: Sun, Chenglu, et al.
Published: (2025)
FairStream: Fair Multimedia Streaming Benchmark for Reinforcement Learning Agents
by: Weil, Jannis, et al.
Published: (2024)
by: Weil, Jannis, et al.
Published: (2024)
How Far Have Medical Vision-Language Models Come? A Comprehensive Benchmarking Study
by: Liu, Che, et al.
Published: (2025)
by: Liu, Che, et al.
Published: (2025)
Cooperative bots exhibit nuanced effects on cooperation across strategic frameworks
by: Si, Zehua, et al.
Published: (2024)
by: Si, Zehua, et al.
Published: (2024)
Food4All: A Multi-Agent Framework for Real-time Free Food Discovery with Integrated Nutritional Metadata
by: Yuan, Zhengqing, et al.
Published: (2025)
by: Yuan, Zhengqing, et al.
Published: (2025)
Beyond Black-Box Benchmarking: Observability, Analytics, and Optimization of Agentic Systems
by: Moshkovich, Dany, et al.
Published: (2025)
by: Moshkovich, Dany, et al.
Published: (2025)
G-UBS: Towards Robust Understanding of Implicit Feedback via Group-Aware User Behavior Simulation
by: Chen, Boyu, et al.
Published: (2025)
by: Chen, Boyu, et al.
Published: (2025)
CataractSurg-80K: Knowledge-Driven Benchmarking for Structured Reasoning in Ophthalmic Surgery Planning
by: Meng, Yang, et al.
Published: (2025)
by: Meng, Yang, et al.
Published: (2025)
LongCLI-Bench: A Preliminary Benchmark and Study for Long-horizon Agentic Programming in Command-Line Interfaces
by: Feng, Yukang, et al.
Published: (2026)
by: Feng, Yukang, et al.
Published: (2026)
Encirclement Guaranteed Cooperative Pursuit with Robust Model Predictive Control
by: Wang, Chen, et al.
Published: (2021)
by: Wang, Chen, et al.
Published: (2021)
AMUSE: Audio-Visual Benchmark and Alignment Framework for Agentic Multi-Speaker Understanding
by: Chowdhury, Sanjoy, et al.
Published: (2025)
by: Chowdhury, Sanjoy, et al.
Published: (2025)
Bearing-based Simultaneous Localization and Affine Formation Tracking for Fixed-wing Unmanned Aerial Vehicles
by: Huiming, Li, et al.
Published: (2023)
by: Huiming, Li, et al.
Published: (2023)
The Subtle Art of Defection: Understanding Uncooperative Behaviors in LLM based Multi-Agent Systems
by: Kulshreshtha, Devang, et al.
Published: (2025)
by: Kulshreshtha, Devang, et al.
Published: (2025)
Benchmarking LLM-based agents for single-cell omics analysis
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
Breaking the Communication-Accuracy Trade-off: A Sparsified Information Diffusion Framework for Multi-Agent Collaborative Perception
by: Zha, Jirong, et al.
Published: (2026)
by: Zha, Jirong, et al.
Published: (2026)
Trace2Skill: Verifier-Guided Skill Evolution for Long-Context EDA Agents
by: Du, Zijian, et al.
Published: (2026)
by: Du, Zijian, et al.
Published: (2026)
Seeing the Whole Elephant: A Benchmark for Failure Attribution in LLM-based Multi-Agent Systems
by: Chen, Mengzhuo, et al.
Published: (2026)
by: Chen, Mengzhuo, et al.
Published: (2026)
Slipstream: Trajectory-Grounded Compaction Validation for Long-Horizon Agents
by: Chen, Zhuofu, et al.
Published: (2026)
by: Chen, Zhuofu, et al.
Published: (2026)
JaxRobotarium: Training and Deploying Multi-Robot Policies in 10 Minutes
by: Jain, Shalin Anand, et al.
Published: (2025)
by: Jain, Shalin Anand, et al.
Published: (2025)
Beyond the All-in-One Agent: Benchmarking Role-Specialized Multi-Agent Collaboration in Enterprise Workflows
by: Yu, Tao, et al.
Published: (2026)
by: Yu, Tao, et al.
Published: (2026)
DSBC : Data Science task Benchmarking with Context engineering
by: Kadiyala, Ram Mohan Rao, et al.
Published: (2025)
by: Kadiyala, Ram Mohan Rao, et al.
Published: (2025)
Similar Items
-
Breaking the Pre-Planning Barrier: Adaptive Real-Time Coordination of Heterogeneous UAVs
by: Hu, Yuhan, et al.
Published: (2025) -
Beyond Curve Fitting: Neuro-Symbolic Agents for Context-Aware Epidemic Forecasting
by: Chae, Joongwon, et al.
Published: (2025) -
AgentWebBench: Benchmarking Multi-Agent Coordination in Agentic Web
by: Zhong, Shanshan, et al.
Published: (2026) -
Personality-Driven Student Agent-Based Modeling in Mathematics Education: How Well Do Student Agents Align with Human Learners?
by: Xiao, Bushi, et al.
Published: (2026) -
Mean Field Correlated Imitation Learning
by: Zhao, Zhiyu, et al.
Published: (2024)