Agent AI: Surveying the Horizons of Multimodal Interaction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Durante, Zane, Huang, Qiuyuan, Wake, Naoki, Gong, Ran, Park, Jae Sung, Sarkar, Bidipta, Taori, Rohan, Noda, Yusuke, Terzopoulos, Demetri, Choi, Yejin, Ikeuchi, Katsushi, Vo, Hoi, Fei-Fei, Li, Gao, Jianfeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Position Paper: Agent AI Towards a Holistic Intelligence
von: Huang, Qiuyuan, et al.
Veröffentlicht: (2024)
von: Huang, Qiuyuan, et al.
Veröffentlicht: (2024)
An Interactive Agent Foundation Model
von: Durante, Zane, et al.
Veröffentlicht: (2024)
von: Durante, Zane, et al.
Veröffentlicht: (2024)
Agreeing to Interact in Human-Robot Interaction using Large Language Models and Vision Language Models
von: Sasabuchi, Kazuhiro, et al.
Veröffentlicht: (2025)
von: Sasabuchi, Kazuhiro, et al.
Veröffentlicht: (2025)
RL-Driven Data Generation for Robust Vision-Based Dexterous Grasping
von: Kanehira, Atsushi, et al.
Veröffentlicht: (2025)
von: Kanehira, Atsushi, et al.
Veröffentlicht: (2025)
Plan-and-Act using Large Language Models for Interactive Agreement
von: Sasabuchi, Kazuhiro, et al.
Veröffentlicht: (2025)
von: Sasabuchi, Kazuhiro, et al.
Veröffentlicht: (2025)
GPT-4V(ision) for Robotics: Multimodal Task Planning from Human Demonstration
von: Wake, Naoki, et al.
Veröffentlicht: (2023)
von: Wake, Naoki, et al.
Veröffentlicht: (2023)
Open-Vocabulary Action Localization with Iterative Visual Prompting
von: Wake, Naoki, et al.
Veröffentlicht: (2024)
von: Wake, Naoki, et al.
Veröffentlicht: (2024)
VLM-driven Behavior Tree for Context-aware Task Planning
von: Wake, Naoki, et al.
Veröffentlicht: (2025)
von: Wake, Naoki, et al.
Veröffentlicht: (2025)
A Taxonomy of Self-Handover
von: Wake, Naoki, et al.
Veröffentlicht: (2025)
von: Wake, Naoki, et al.
Veröffentlicht: (2025)
IK Seed Generator for Dual-Arm Human-like Physicality Robot with Mobile Base
von: Takamatsu, Jun, et al.
Veröffentlicht: (2025)
von: Takamatsu, Jun, et al.
Veröffentlicht: (2025)
Prompting Medical Large Vision-Language Models to Diagnose Pathologies by Visual Question Answering
von: Guo, Danfeng, et al.
Veröffentlicht: (2024)
von: Guo, Danfeng, et al.
Veröffentlicht: (2024)
Designing Library of Skill-Agents for Hardware-Level Reusability
von: Takamatsu, Jun, et al.
Veröffentlicht: (2024)
von: Takamatsu, Jun, et al.
Veröffentlicht: (2024)
Modality-Driven Design for Multi-Step Dexterous Manipulation: Insights from Neuroscience
von: Wake, Naoki, et al.
Veröffentlicht: (2024)
von: Wake, Naoki, et al.
Veröffentlicht: (2024)
APriCoT: Action Primitives based on Contact-state Transition for In-Hand Tool Manipulation
von: Saito, Daichi, et al.
Veröffentlicht: (2024)
von: Saito, Daichi, et al.
Veröffentlicht: (2024)
TeamCraft: A Benchmark for Multi-Modal Multi-Agent Systems in Minecraft
von: Long, Qian, et al.
Veröffentlicht: (2024)
von: Long, Qian, et al.
Veröffentlicht: (2024)
Manipulación mecánica de partes aleatoriamente orientadas
von: Horn, Berthold K. y Ikeuchi, Katsushi
Veröffentlicht: (1984)
von: Horn, Berthold K. y Ikeuchi, Katsushi
Veröffentlicht: (1984)
VideoWeave: A Data-Centric Approach for Efficient Video Understanding
von: Durante, Zane, et al.
Veröffentlicht: (2026)
von: Durante, Zane, et al.
Veröffentlicht: (2026)
Inverse Attention Agents for Multi-Agent Systems
von: Long, Qian, et al.
Veröffentlicht: (2024)
von: Long, Qian, et al.
Veröffentlicht: (2024)
Cross-Slice Attention and Evidential Critical Loss for Uncertainty-Aware Prostate Cancer Detection
von: Hung, Alex Ling Yu, et al.
Veröffentlicht: (2024)
von: Hung, Alex Ling Yu, et al.
Veröffentlicht: (2024)
Learning Neural Force Manifolds for Sim2Real Robotic Symmetrical Paper Folding
von: Choi, Andrew, et al.
Veröffentlicht: (2023)
von: Choi, Andrew, et al.
Veröffentlicht: (2023)
A Human-centric Framework for Debating the Ethics of AI Consciousness Under Uncertainty
von: Ziheng, Zhou, et al.
Veröffentlicht: (2025)
von: Ziheng, Zhou, et al.
Veröffentlicht: (2025)
Trajectories of second language student classroom engagement: Profiles and correlates
von: Hoi Vo
Veröffentlicht: (2026)
von: Hoi Vo
Veröffentlicht: (2026)
mBEST: Realtime Deformable Linear Object Detection Through Minimal Bending Energy Skeleton Pixel Traversals
von: Choi, Andrew, et al.
Veröffentlicht: (2023)
von: Choi, Andrew, et al.
Veröffentlicht: (2023)
Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning
von: Sarkar, Bidipta, et al.
Veröffentlicht: (2025)
von: Sarkar, Bidipta, et al.
Veröffentlicht: (2025)
Unstructured Moving Least Squares Material Point Methods: A Stable Kernel Approach With Continuous Gradient Reconstruction on General Unstructured Tessellations
von: Cao, Yadi, et al.
Veröffentlicht: (2023)
von: Cao, Yadi, et al.
Veröffentlicht: (2023)
Physically Grounded Vision-Language Models for Robotic Manipulation
von: Gao, Jensen, et al.
Veröffentlicht: (2023)
von: Gao, Jensen, et al.
Veröffentlicht: (2023)
Towards Fine-Grained Video Question Answering
von: Dai, Wei, et al.
Veröffentlicht: (2025)
von: Dai, Wei, et al.
Veröffentlicht: (2025)
Few-Shot Classification of Interactive Activities of Daily Living (InteractADL)
von: Durante, Zane, et al.
Veröffentlicht: (2024)
von: Durante, Zane, et al.
Veröffentlicht: (2024)
Review of Deep Learning Applications to Structural Proteomics Enabled by Cryogenic Electron Microscopy and Tomography
von: Zhou, Brady K., et al.
Veröffentlicht: (2025)
von: Zhou, Brady K., et al.
Veröffentlicht: (2025)
AnimaMimic: Imitating 3D Animation from Video Priors
von: Xie, Tianyi, et al.
Veröffentlicht: (2025)
von: Xie, Tianyi, et al.
Veröffentlicht: (2025)
CFSum: A Transformer-Based Multi-Modal Video Summarization Framework With Coarse-Fine Fusion
von: Guo, Yaowei, et al.
Veröffentlicht: (2025)
von: Guo, Yaowei, et al.
Veröffentlicht: (2025)
Sub-Horizon Amplification of Curvature Perturbations: A New Route to Primordial Black Holes and Gravitational Waves
von: Nandi, Debottam, et al.
Veröffentlicht: (2025)
von: Nandi, Debottam, et al.
Veröffentlicht: (2025)
Simple Role Assignment is Extraordinarily Effective for Safety Alignment
von: Ziheng, Zhou, et al.
Veröffentlicht: (2026)
von: Ziheng, Zhou, et al.
Veröffentlicht: (2026)
MPM Lite: Linear Kernels and Integration without Particles
von: Feng, Xiang, et al.
Veröffentlicht: (2026)
von: Feng, Xiang, et al.
Veröffentlicht: (2026)
Fractional heat content asymptotics for Carnot groups
von: Sarkar, Rohan
Veröffentlicht: (2026)
von: Sarkar, Rohan
Veröffentlicht: (2026)
Small time asymptotics of spectral heat content of isotropic processes
von: Sarkar, Rohan
Veröffentlicht: (2025)
von: Sarkar, Rohan
Veröffentlicht: (2025)
Spectral theory of non-local Ornstein-Uhlenbeck operators
von: Sarkar, Rohan
Veröffentlicht: (2025)
von: Sarkar, Rohan
Veröffentlicht: (2025)
HourVideo: 1-Hour Video-Language Understanding
von: Chandrasegaran, Keshigeyan, et al.
Veröffentlicht: (2024)
von: Chandrasegaran, Keshigeyan, et al.
Veröffentlicht: (2024)
Wonderland: Navigating 3D Scenes from a Single Image
von: Liang, Hanwen, et al.
Veröffentlicht: (2024)
von: Liang, Hanwen, et al.
Veröffentlicht: (2024)
Task Value, Teacher Enthusiasm, and Student Engagement in Online Second Language Learning: A Latent Moderated Model
von: Hoi Vo, et al.
Veröffentlicht: (2026)
von: Hoi Vo, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Position Paper: Agent AI Towards a Holistic Intelligence
von: Huang, Qiuyuan, et al.
Veröffentlicht: (2024) -
An Interactive Agent Foundation Model
von: Durante, Zane, et al.
Veröffentlicht: (2024) -
Agreeing to Interact in Human-Robot Interaction using Large Language Models and Vision Language Models
von: Sasabuchi, Kazuhiro, et al.
Veröffentlicht: (2025) -
RL-Driven Data Generation for Robust Vision-Based Dexterous Grasping
von: Kanehira, Atsushi, et al.
Veröffentlicht: (2025) -
Plan-and-Act using Large Language Models for Interactive Agreement
von: Sasabuchi, Kazuhiro, et al.
Veröffentlicht: (2025)