AURA: A Multi-Modal Medical Agent for Understanding, Reasoning & Annotation
Fuente:
arXiv
Saved in:
| Main Authors: | Fathi, Nima, Kumar, Amar, Arbel, Tal |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scene-Aware Vectorized Memory Multi-Agent Framework with Cross-Modal Differentiated Quantization VLMs for Visually Impaired Assistance
by: Wang, Xiangxiang, et al.
Published: (2025)
by: Wang, Xiangxiang, et al.
Published: (2025)
LongVideoAgent: Multi-Agent Reasoning with Long Videos
by: Liu, Runtao, et al.
Published: (2025)
by: Liu, Runtao, et al.
Published: (2025)
MAG-3D: Multi-Agent Grounded Reasoning for 3D Understanding
by: Zheng, Henry, et al.
Published: (2026)
by: Zheng, Henry, et al.
Published: (2026)
Multi-Agent Dynamic Relational Reasoning for Social Robot Navigation
by: Li, Jiachen, et al.
Published: (2024)
by: Li, Jiachen, et al.
Published: (2024)
MedRoute: RL-Based Dynamic Specialist Routing in Multi-Agent Medical Diagnosis
by: Vayani, Ashmal, et al.
Published: (2026)
by: Vayani, Ashmal, et al.
Published: (2026)
Chain of Questions: Guiding Multimodal Curiosity in Language Models
by: Iji, Nima, et al.
Published: (2025)
by: Iji, Nima, et al.
Published: (2025)
MedChat: A Multi-Agent Framework for Multimodal Diagnosis with Large Language Models
by: Liu, Philip R., et al.
Published: (2025)
by: Liu, Philip R., et al.
Published: (2025)
TranSPORTmer: A Holistic Approach to Trajectory Understanding in Multi-Agent Sports
by: Capellera, Guillem, et al.
Published: (2024)
by: Capellera, Guillem, et al.
Published: (2024)
Hollywood Town: Long-Video Generation via Cross-Modal Multi-Agent Orchestration
by: Wei, Zheng, et al.
Published: (2025)
by: Wei, Zheng, et al.
Published: (2025)
Experience-Driven Multi-Agent Systems Are Training-free Context-aware Earth Observers
by: Dai, Pengyu, et al.
Published: (2026)
by: Dai, Pengyu, et al.
Published: (2026)
A Multi-Agent Perception-Action Alliance for Efficient Long Video Reasoning
by: Xu, Yichang, et al.
Published: (2026)
by: Xu, Yichang, et al.
Published: (2026)
RADAR: A Risk-Aware Dynamic Multi-Agent Framework for LLM Safety Evaluation via Role-Specialized Collaboration
by: Chen, Xiuyuan, et al.
Published: (2025)
by: Chen, Xiuyuan, et al.
Published: (2025)
FEAT: A Multi-Agent Forensic AI System with Domain-Adapted Large Language Model for Automated Cause-of-Death Analysis
by: Shen, Chen, et al.
Published: (2025)
by: Shen, Chen, et al.
Published: (2025)
OctoTools: An Agentic Framework with Extensible Tools for Complex Reasoning
by: Lu, Pan, et al.
Published: (2025)
by: Lu, Pan, et al.
Published: (2025)
AgentCVR: Active Multi-Agent Cross-Video Reasoning via Script-Simulated Reinforcement Learning
by: Qiu, Yilun, et al.
Published: (2026)
by: Qiu, Yilun, et al.
Published: (2026)
Multi-Agent VQA: Exploring Multi-Agent Foundation Models in Zero-Shot Visual Question Answering
by: Jiang, Bowen, et al.
Published: (2024)
by: Jiang, Bowen, et al.
Published: (2024)
CMP: Cooperative Motion Prediction with Multi-Agent Communication
by: Wang, Zehao, et al.
Published: (2024)
by: Wang, Zehao, et al.
Published: (2024)
MATRIX: Multi-Agent Trajectory Generation with Diverse Contexts
by: Xu, Zhuo, et al.
Published: (2024)
by: Xu, Zhuo, et al.
Published: (2024)
STROOBnet Optimization via GPU-Accelerated Proximal Recurrence Strategies
by: Holmberg, Ted Edward, et al.
Published: (2024)
by: Holmberg, Ted Edward, et al.
Published: (2024)
PixelFlowCast: Latent-Free Precipitation Nowcasting via Pixel Mean Flows
by: Zhu, Yufeng, et al.
Published: (2026)
by: Zhu, Yufeng, et al.
Published: (2026)
On the Road to Clarity: Exploring Explainable AI for World Models in a Driver Assistance System
by: Roshdi, Mohamed, et al.
Published: (2024)
by: Roshdi, Mohamed, et al.
Published: (2024)
DiffCP: Ultra-Low Bit Collaborative Perception via Diffusion Model
by: Mao, Ruiqing, et al.
Published: (2024)
by: Mao, Ruiqing, et al.
Published: (2024)
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering
by: Kugo, Noriyuki, et al.
Published: (2025)
by: Kugo, Noriyuki, et al.
Published: (2025)
Spatio-Temporal Graph Dual-Attention Network for Multi-Agent Prediction and Tracking
by: Li, Jiachen, et al.
Published: (2021)
by: Li, Jiachen, et al.
Published: (2021)
PathFinder: A Multi-Modal Multi-Agent System for Medical Diagnostic Decision-Making Applied to Histopathology
by: Ghezloo, Fatemeh, et al.
Published: (2025)
by: Ghezloo, Fatemeh, et al.
Published: (2025)
VideoChat-M1: Collaborative Policy Planning for Video Understanding via Multi-Agent Reinforcement Learning
by: Chen, Boyu, et al.
Published: (2025)
by: Chen, Boyu, et al.
Published: (2025)
DeCoDEx: Confounder Detector Guidance for Improved Diffusion-based Counterfactual Explanations
by: Fathi, Nima, et al.
Published: (2024)
by: Fathi, Nima, et al.
Published: (2024)
ToolTok: Tool Tokenization for Efficient and Generalizable GUI Agents
by: Wang, Xiaoce, et al.
Published: (2026)
by: Wang, Xiaoce, et al.
Published: (2026)
CommCP: Efficient Multi-Agent Coordination via LLM-Based Communication with Conformal Prediction
by: Zhang, Xiaopan, et al.
Published: (2026)
by: Zhang, Xiaopan, et al.
Published: (2026)
EmboTeam: Grounding LLM Reasoning into Reactive Behavior Trees via PDDL for Embodied Multi-Robot Collaboration
by: Zeng, Haishan, et al.
Published: (2026)
by: Zeng, Haishan, et al.
Published: (2026)
ProtoMedAgent: Multimodal Clinical Interpretability via Privacy-Aware Agentic Workflows
by: Pellicer, Alvaro Lopez, et al.
Published: (2026)
by: Pellicer, Alvaro Lopez, et al.
Published: (2026)
Many Heads Are Better Than One: Improved Scientific Idea Generation by A LLM-Based Multi-Agent System
by: Su, Haoyang, et al.
Published: (2024)
by: Su, Haoyang, et al.
Published: (2024)
LaMMA-P: Generalizable Multi-Agent Long-Horizon Task Allocation and Planning with LM-Driven PDDL Planner
by: Zhang, Xiaopan, et al.
Published: (2024)
by: Zhang, Xiaopan, et al.
Published: (2024)
FetalAgents: A Multi-Agent System for Fetal Ultrasound Image and Video Analysis
by: Hu, Xiaotian, et al.
Published: (2026)
by: Hu, Xiaotian, et al.
Published: (2026)
Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast
by: Gu, Xiangming, et al.
Published: (2024)
by: Gu, Xiangming, et al.
Published: (2024)
LVRPO: Language-Visual Alignment with GRPO for Multimodal Understanding and Generation
by: Mo, Shentong, et al.
Published: (2026)
by: Mo, Shentong, et al.
Published: (2026)
FedAgentBench: Towards Automating Real-world Federated Medical Image Analysis with Server-Client LLM Agents
by: Saha, Pramit, et al.
Published: (2025)
by: Saha, Pramit, et al.
Published: (2025)
Agent Planning with World Knowledge Model
by: Qiao, Shuofei, et al.
Published: (2024)
by: Qiao, Shuofei, et al.
Published: (2024)
Towards Reliable Fetal Ultrasound Interpretation with Multi-Agent Collaboration
by: Hu, Xiaotian, et al.
Published: (2026)
by: Hu, Xiaotian, et al.
Published: (2026)
From Perception to Action: Spatial AI Agents and World Models
by: Felicia, Gloria, et al.
Published: (2026)
by: Felicia, Gloria, et al.
Published: (2026)
Similar Items
-
Scene-Aware Vectorized Memory Multi-Agent Framework with Cross-Modal Differentiated Quantization VLMs for Visually Impaired Assistance
by: Wang, Xiangxiang, et al.
Published: (2025) -
LongVideoAgent: Multi-Agent Reasoning with Long Videos
by: Liu, Runtao, et al.
Published: (2025) -
MAG-3D: Multi-Agent Grounded Reasoning for 3D Understanding
by: Zheng, Henry, et al.
Published: (2026) -
Multi-Agent Dynamic Relational Reasoning for Social Robot Navigation
by: Li, Jiachen, et al.
Published: (2024) -
MedRoute: RL-Based Dynamic Specialist Routing in Multi-Agent Medical Diagnosis
by: Vayani, Ashmal, et al.
Published: (2026)