Magma: A Foundation Model for Multimodal AI Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Jianwei, Tan, Reuben, Wu, Qianhui, Zheng, Ruijie, Peng, Baolin, Liang, Yongyuan, Gu, Yu, Cai, Mu, Ye, Seonghyeon, Jang, Joel, Deng, Yuquan, Liden, Lars, Gao, Jianfeng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Latent Action Pretraining from Videos
by: Ye, Seonghyeon, et al.
Published: (2024)
by: Ye, Seonghyeon, et al.
Published: (2024)
GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents
by: Wu, Qianhui, et al.
Published: (2025)
by: Wu, Qianhui, et al.
Published: (2025)
AsgardBench -- Evaluating Visually Grounded Interactive Planning Under Minimal Feedback
by: Tupini, Andrea, et al.
Published: (2026)
by: Tupini, Andrea, et al.
Published: (2026)
LLM-based Multimodal Feedback Produces Equivalent Learning and Better Student Perceptions than Educator Feedback
by: Zhao, Chloe Qianhui, et al.
Published: (2026)
by: Zhao, Chloe Qianhui, et al.
Published: (2026)
From First Draft to Final Insight: A Multi-Agent Approach for Feedback Generation
by: Cao, Jie, et al.
Published: (2025)
by: Cao, Jie, et al.
Published: (2025)
Access Denied: Meaningful Data Access for Quantitative Algorithm Audits
by: Zaccour, Juliette, et al.
Published: (2025)
by: Zaccour, Juliette, et al.
Published: (2025)
Hear You in Silence: Designing for Active Listening in Human Interaction with Conversational Agents Using Context-Aware Pacing
by: Jiang, Zhihan, et al.
Published: (2026)
by: Jiang, Zhihan, et al.
Published: (2026)
Trustworthy and Practical AI for Healthcare: A Guided Deferral System with Large Language Models
by: Strong, Joshua, et al.
Published: (2024)
by: Strong, Joshua, et al.
Published: (2024)
Toward Automated Qualitative Analysis: Leveraging Large Language Models for Tutoring Dialogue Evaluation
by: Gu, Megan, et al.
Published: (2025)
by: Gu, Megan, et al.
Published: (2025)
How Users Understand Robot Foundation Model Performance through Task Success Rates and Beyond
by: Sheidlower, Isaac, et al.
Published: (2026)
by: Sheidlower, Isaac, et al.
Published: (2026)
Task-Aware Delegation Cues for LLM Agents
by: Gu, Xingrui
Published: (2026)
by: Gu, Xingrui
Published: (2026)
Gensors: Authoring Personalized Visual Sensors with Multimodal Foundation Models and Reasoning
by: Liu, Michael Xieyang, et al.
Published: (2025)
by: Liu, Michael Xieyang, et al.
Published: (2025)
Beyond the AI Tutor: Social Learning with LLM Agents
by: Kumar, Harsh, et al.
Published: (2026)
by: Kumar, Harsh, et al.
Published: (2026)
Exploring the Feasibility of Multimodal Chatbot AI as Copilot in Pathology Diagnostics: Generalist Model's Pitfall
by: Liu, Mianxin, et al.
Published: (2024)
by: Liu, Mianxin, et al.
Published: (2024)
"Diversity is Having the Diversity": Unpacking and Designing for Diversity in Applicant Selection
by: Natarajan, Neil, et al.
Published: (2024)
by: Natarajan, Neil, et al.
Published: (2024)
Control-Theoretic Analysis of Shared Control Systems
by: Aronson, Reuben M., et al.
Published: (2024)
by: Aronson, Reuben M., et al.
Published: (2024)
SlideItRight: Using AI to Find Relevant Slides and Provide Feedback for Open-Ended Questions
by: Zhao, Chloe Qianhui, et al.
Published: (2025)
by: Zhao, Chloe Qianhui, et al.
Published: (2025)
Can Virtual Agents Care? Designing an Empathetic and Personalized LLM-Driven Conversational Agent
by: Toan, Truong Le Minh, et al.
Published: (2026)
by: Toan, Truong Le Minh, et al.
Published: (2026)
Respectful Things: Adding Social Intelligence to 'Smart' Devices
by: Van Kleek, Max, et al.
Published: (2026)
by: Van Kleek, Max, et al.
Published: (2026)
Operationalizing Perceptions of Agent Gender: Foundations and Guidelines
by: Seaborn, Katie, et al.
Published: (2026)
by: Seaborn, Katie, et al.
Published: (2026)
Explainable Iterative Data Visualisation Refinement via an LLM Agent
by: Susam, Burak, et al.
Published: (2026)
by: Susam, Burak, et al.
Published: (2026)
The Visualization JUDGE : Can Multimodal Foundation Models Guide Visualization Design Through Visual Perception?
by: Berger, Matthew, et al.
Published: (2024)
by: Berger, Matthew, et al.
Published: (2024)
Not Even Nice Work If You Can Get It; A Longitudinal Study of Uber's Algorithmic Pay and Pricing
by: Binns, Reuben, et al.
Published: (2025)
by: Binns, Reuben, et al.
Published: (2025)
Agent AI: Surveying the Horizons of Multimodal Interaction
by: Durante, Zane, et al.
Published: (2024)
by: Durante, Zane, et al.
Published: (2024)
PhysicsSolutionAgent: Towards Multimodal Explanations for Numerical Physics Problem Solving
by: Thole, Aditya, et al.
Published: (2026)
by: Thole, Aditya, et al.
Published: (2026)
Matryoshka Multimodal Models
by: Cai, Mu, et al.
Published: (2024)
by: Cai, Mu, et al.
Published: (2024)
See What I Mean? Expressiveness and Clarity in Robot Display Design
by: Ebisu, Matthew, et al.
Published: (2025)
by: Ebisu, Matthew, et al.
Published: (2025)
ARCADE: An Augmented Reality Display Environment for Multimodal Interaction with Conversational Agents
by: Schindler, Carolin, et al.
Published: (2024)
by: Schindler, Carolin, et al.
Published: (2024)
Exploring General-Purpose Autonomous Multimodal Agents for Pathology Report Generation
by: Aubreville, Marc, et al.
Published: (2025)
by: Aubreville, Marc, et al.
Published: (2025)
Unit-Based Agent for Semi-Cascaded Full-Duplex Dialogue Systems
by: Yu, Haoyuan, et al.
Published: (2026)
by: Yu, Haoyuan, et al.
Published: (2026)
LegalWebAgent: Empowering Access to Justice via LLM-Based Web Agents
by: Tan, Jinzhe, et al.
Published: (2025)
by: Tan, Jinzhe, et al.
Published: (2025)
Customized FinGPT Search Agents Using Foundation Models
by: Tian, Felix, et al.
Published: (2024)
by: Tian, Felix, et al.
Published: (2024)
EyeAgent: An Agentic AI System for Multimodal Clinical Decision Support in Ophthalmology
by: Shi, Danli, et al.
Published: (2025)
by: Shi, Danli, et al.
Published: (2025)
Design and Evaluation of a Culturally Adapted Multimodal Virtual Agent for PTSD Screening
by: Ozel, Cengiz, et al.
Published: (2026)
by: Ozel, Cengiz, et al.
Published: (2026)
The Benefits of Prosociality towards AI Agents: Examining the Effects of Helping AI Agents on Human Well-Being
by: Zhu, Zicheng, et al.
Published: (2025)
by: Zhu, Zicheng, et al.
Published: (2025)
Bridging Culture and Finance: A Multimodal Analysis of Memecoins in the Web3 Ecosystem
by: Long, Hou-Wan, et al.
Published: (2024)
by: Long, Hou-Wan, et al.
Published: (2024)
A-MEM: Agentic Memory for LLM Agents
by: Xu, Wujiang, et al.
Published: (2025)
by: Xu, Wujiang, et al.
Published: (2025)
SIAgent: Spatial Interaction Agent via LLM-powered Eye-Hand Motion Intent Understanding in VR
by: Wang, Zhimin, et al.
Published: (2026)
by: Wang, Zhimin, et al.
Published: (2026)
Assessing the Impact and Underlying Pathways of Sequenced AI feedback on Student Learning
by: Cao, Jie, et al.
Published: (2026)
by: Cao, Jie, et al.
Published: (2026)
GUI Agents with Foundation Models: A Comprehensive Survey
by: Wang, Shuai, et al.
Published: (2024)
by: Wang, Shuai, et al.
Published: (2024)
Similar Items
-
Latent Action Pretraining from Videos
by: Ye, Seonghyeon, et al.
Published: (2024) -
GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents
by: Wu, Qianhui, et al.
Published: (2025) -
AsgardBench -- Evaluating Visually Grounded Interactive Planning Under Minimal Feedback
by: Tupini, Andrea, et al.
Published: (2026) -
LLM-based Multimodal Feedback Produces Equivalent Learning and Better Student Perceptions than Educator Feedback
by: Zhao, Chloe Qianhui, et al.
Published: (2026) -
From First Draft to Final Insight: A Multi-Agent Approach for Feedback Generation
by: Cao, Jie, et al.
Published: (2025)