An Interactive Agent Foundation Model
Fuente:
arXiv
Saved in:
| Main Authors: | Durante, Zane, Sarkar, Bidipta, Gong, Ran, Taori, Rohan, Noda, Yusuke, Tang, Paul, Adeli, Ehsan, Lakshmikanth, Shrinidhi Kowshika, Schulman, Kevin, Milstein, Arnold, Terzopoulos, Demetri, Famoti, Ade, Kuno, Noboru, Llorens, Ashley, Vo, Hoi, Ikeuchi, Katsu, Fei-Fei, Li, Gao, Jianfeng, Wake, Naoki, Huang, Qiuyuan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Position Paper: Agent AI Towards a Holistic Intelligence
by: Huang, Qiuyuan, et al.
Published: (2024)
by: Huang, Qiuyuan, et al.
Published: (2024)
Agent AI: Surveying the Horizons of Multimodal Interaction
by: Durante, Zane, et al.
Published: (2024)
by: Durante, Zane, et al.
Published: (2024)
Towards Fine-Grained Video Question Answering
by: Dai, Wei, et al.
Published: (2025)
by: Dai, Wei, et al.
Published: (2025)
The Language of Motion: Unifying Verbal and Non-verbal Language of 3D Human Motion
by: Chen, Changan, et al.
Published: (2024)
by: Chen, Changan, et al.
Published: (2024)
Prompting Medical Large Vision-Language Models to Diagnose Pathologies by Visual Question Answering
by: Guo, Danfeng, et al.
Published: (2024)
by: Guo, Danfeng, et al.
Published: (2024)
ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body
by: Zhang, Juze, et al.
Published: (2025)
by: Zhang, Juze, et al.
Published: (2025)
Few-Shot Classification of Interactive Activities of Daily Living (InteractADL)
by: Durante, Zane, et al.
Published: (2024)
by: Durante, Zane, et al.
Published: (2024)
VideoWeave: A Data-Centric Approach for Efficient Video Understanding
by: Durante, Zane, et al.
Published: (2026)
by: Durante, Zane, et al.
Published: (2026)
Inverse Attention Agents for Multi-Agent Systems
by: Long, Qian, et al.
Published: (2024)
by: Long, Qian, et al.
Published: (2024)
Learning Neural Force Manifolds for Sim2Real Robotic Symmetrical Paper Folding
by: Choi, Andrew, et al.
Published: (2023)
by: Choi, Andrew, et al.
Published: (2023)
A Human-centric Framework for Debating the Ethics of AI Consciousness Under Uncertainty
by: Ziheng, Zhou, et al.
Published: (2025)
by: Ziheng, Zhou, et al.
Published: (2025)
Trajectories of second language student classroom engagement: Profiles and correlates
by: Hoi Vo
Published: (2026)
by: Hoi Vo
Published: (2026)
TeamCraft: A Benchmark for Multi-Modal Multi-Agent Systems in Minecraft
by: Long, Qian, et al.
Published: (2024)
by: Long, Qian, et al.
Published: (2024)
mBEST: Realtime Deformable Linear Object Detection Through Minimal Bending Energy Skeleton Pixel Traversals
by: Choi, Andrew, et al.
Published: (2023)
by: Choi, Andrew, et al.
Published: (2023)
T*: Re-thinking Temporal Search for Long-Form Video Understanding
by: Ye, Jinhui, et al.
Published: (2025)
by: Ye, Jinhui, et al.
Published: (2025)
Agreeing to Interact in Human-Robot Interaction using Large Language Models and Vision Language Models
by: Sasabuchi, Kazuhiro, et al.
Published: (2025)
by: Sasabuchi, Kazuhiro, et al.
Published: (2025)
RL-Driven Data Generation for Robust Vision-Based Dexterous Grasping
by: Kanehira, Atsushi, et al.
Published: (2025)
by: Kanehira, Atsushi, et al.
Published: (2025)
Plan-and-Act using Large Language Models for Interactive Agreement
by: Sasabuchi, Kazuhiro, et al.
Published: (2025)
by: Sasabuchi, Kazuhiro, et al.
Published: (2025)
GPT-4V(ision) for Robotics: Multimodal Task Planning from Human Demonstration
by: Wake, Naoki, et al.
Published: (2023)
by: Wake, Naoki, et al.
Published: (2023)
Open-Vocabulary Action Localization with Iterative Visual Prompting
by: Wake, Naoki, et al.
Published: (2024)
by: Wake, Naoki, et al.
Published: (2024)
VLM-driven Behavior Tree for Context-aware Task Planning
by: Wake, Naoki, et al.
Published: (2025)
by: Wake, Naoki, et al.
Published: (2025)
A Taxonomy of Self-Handover
by: Wake, Naoki, et al.
Published: (2025)
by: Wake, Naoki, et al.
Published: (2025)
IK Seed Generator for Dual-Arm Human-like Physicality Robot with Mobile Base
by: Takamatsu, Jun, et al.
Published: (2025)
by: Takamatsu, Jun, et al.
Published: (2025)
Unstructured Moving Least Squares Material Point Methods: A Stable Kernel Approach With Continuous Gradient Reconstruction on General Unstructured Tessellations
by: Cao, Yadi, et al.
Published: (2023)
by: Cao, Yadi, et al.
Published: (2023)
Cross-Slice Attention and Evidential Critical Loss for Uncertainty-Aware Prostate Cancer Detection
by: Hung, Alex Ling Yu, et al.
Published: (2024)
by: Hung, Alex Ling Yu, et al.
Published: (2024)
On Approximating the Weighted Region Problem in Square Tessellations
by: Kakimura, Naonori, et al.
Published: (2024)
by: Kakimura, Naonori, et al.
Published: (2024)
Pneumatic bladder links with wide range of motion joints for articulated inflatable robots
by: Uchiyama, Katsu, et al.
Published: (2025)
by: Uchiyama, Katsu, et al.
Published: (2025)
Review of Deep Learning Applications to Structural Proteomics Enabled by Cryogenic Electron Microscopy and Tomography
by: Zhou, Brady K., et al.
Published: (2025)
by: Zhou, Brady K., et al.
Published: (2025)
AnimaMimic: Imitating 3D Animation from Video Priors
by: Xie, Tianyi, et al.
Published: (2025)
by: Xie, Tianyi, et al.
Published: (2025)
CFSum: A Transformer-Based Multi-Modal Video Summarization Framework With Coarse-Fine Fusion
by: Guo, Yaowei, et al.
Published: (2025)
by: Guo, Yaowei, et al.
Published: (2025)
Designing Library of Skill-Agents for Hardware-Level Reusability
by: Takamatsu, Jun, et al.
Published: (2024)
by: Takamatsu, Jun, et al.
Published: (2024)
On non-existence of bifurcations in one-dimensional Bratu equation
by: Pandurangi, Shrinidhi S., et al.
Published: (2025)
by: Pandurangi, Shrinidhi S., et al.
Published: (2025)
A back-linked Fabry-Perot interferometer for space-borne gravitational wave observations
by: Izumi, Kiwamu, et al.
Published: (2020)
by: Izumi, Kiwamu, et al.
Published: (2020)
OccFusion: Rendering Occluded Humans with Generative Diffusion Priors
by: Sun, Adam, et al.
Published: (2024)
by: Sun, Adam, et al.
Published: (2024)
Latest Object Memory Management for Temporally Consistent Video Instance Segmentation
by: Lee, Seunghun, et al.
Published: (2025)
by: Lee, Seunghun, et al.
Published: (2025)
The Essential Role of Causality in Foundation World Models for Embodied AI
by: Gupta, Tarun, et al.
Published: (2024)
by: Gupta, Tarun, et al.
Published: (2024)
Global risk management: The role of collective cognition in response to COVID‐19. By LouiseComfort, Mary LeeRhodes, New York and London: Routledge. 2022. pp. 291. ISBN: 978‐1‐032‐18182‐0
by: Paul Schulman
Published: (2024)
by: Paul Schulman
Published: (2024)
A Look Inside Those Shiny Covers: Mass-Market Children's Books.
by: Schulman, Janet
Published: (1982)
by: Schulman, Janet
Published: (1982)
A suggested reorganization of the Florida marine fisheries laws
by: Schulman, M.
Published: (1953)
by: Schulman, M.
Published: (1953)
Simple Role Assignment is Extraordinarily Effective for Safety Alignment
by: Ziheng, Zhou, et al.
Published: (2026)
by: Ziheng, Zhou, et al.
Published: (2026)
Similar Items
-
Position Paper: Agent AI Towards a Holistic Intelligence
by: Huang, Qiuyuan, et al.
Published: (2024) -
Agent AI: Surveying the Horizons of Multimodal Interaction
by: Durante, Zane, et al.
Published: (2024) -
Towards Fine-Grained Video Question Answering
by: Dai, Wei, et al.
Published: (2025) -
The Language of Motion: Unifying Verbal and Non-verbal Language of 3D Human Motion
by: Chen, Changan, et al.
Published: (2024) -
Prompting Medical Large Vision-Language Models to Diagnose Pathologies by Visual Question Answering
by: Guo, Danfeng, et al.
Published: (2024)