Embodied Web Agents: Bridging Physical-Digital Realms for Integrated Agent Intelligence
Fuente:
arXiv
Saved in:
| Main Authors: | Hong, Yining, Sun, Rui, Li, Bingxuan, Yao, Xingcheng, Wu, Maxine, Chien, Alexander, Yin, Da, Wu, Ying Nian, Wang, Zhecan James, Chang, Kai-Wei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Detecting Multimodal Situations with Insufficient Context and Abstaining from Baseless Predictions
by: Liu, Junzhang, et al.
Published: (2024)
by: Liu, Junzhang, et al.
Published: (2024)
Memory-Centric Embodied Question Answering
by: Zhai, Mingliang, et al.
Published: (2025)
by: Zhai, Mingliang, et al.
Published: (2025)
CoNav: Collaborative Cross-Modal Reasoning for Embodied Navigation
by: Hao, Haihong, et al.
Published: (2025)
by: Hao, Haihong, et al.
Published: (2025)
SentiAvatar: Towards Expressive and Interactive Digital Humans
by: Jin, Chuhao, et al.
Published: (2026)
by: Jin, Chuhao, et al.
Published: (2026)
Integrating Multi-Modal Sensors: A Review of Fusion Techniques for Intelligent Vehicles
by: Wei, Chuheng, et al.
Published: (2025)
by: Wei, Chuheng, et al.
Published: (2025)
Unify-Agent: A Unified Multimodal Agent for World-Grounded Image Synthesis
by: Chen, Shuang, et al.
Published: (2026)
by: Chen, Shuang, et al.
Published: (2026)
Livia: An Emotion-Aware AR Companion Powered by Modular AI Agents and Progressive Memory Compression
by: Xi, Rui, et al.
Published: (2025)
by: Xi, Rui, et al.
Published: (2025)
The Interspeech 2026 Audio Reasoning Challenge: Evaluating Reasoning Process Quality for Audio Reasoning Models and Agents
by: Ma, Ziyang, et al.
Published: (2026)
by: Ma, Ziyang, et al.
Published: (2026)
TeleAntiFraud-28k: An Audio-Text Slow-Thinking Dataset for Telecom Fraud Detection
by: Ma, Zhiming, et al.
Published: (2025)
by: Ma, Zhiming, et al.
Published: (2025)
Simulacra Naturae: Generative Ecosystem driven by Agent-Based Simulations and Brain Organoid Collective Intelligence
by: Manoudaki, Nefeli, et al.
Published: (2025)
by: Manoudaki, Nefeli, et al.
Published: (2025)
Spoken Humanoid Embodied Conversational Agents in Mobile Serious Games: A Usability Assessment
by: Korre, Danai, et al.
Published: (2023)
by: Korre, Danai, et al.
Published: (2023)
SSTFormer: Bridging Spiking Neural Network and Memory Support Transformer for Frame-Event based Recognition
by: Wang, Xiao, et al.
Published: (2023)
by: Wang, Xiao, et al.
Published: (2023)
Modular Conversational Agents for Surveys and Interviews
by: Yu, Jiangbo, et al.
Published: (2024)
by: Yu, Jiangbo, et al.
Published: (2024)
Contrastive Visual Data Augmentation
by: Zhou, Yu, et al.
Published: (2025)
by: Zhou, Yu, et al.
Published: (2025)
HaloQuest: A Visual Hallucination Dataset for Advancing Multimodal Reasoning
by: Wang, Zhecan, et al.
Published: (2024)
by: Wang, Zhecan, et al.
Published: (2024)
Agency Among Agents: Designing with Hypertextual Friction in the Algorithmic Web
by: Liu, Sophia, et al.
Published: (2025)
by: Liu, Sophia, et al.
Published: (2025)
EditEmoTalk: Controllable Speech-Driven 3D Facial Animation with Continuous Expression Editing
by: Jiang, Diqiong, et al.
Published: (2026)
by: Jiang, Diqiong, et al.
Published: (2026)
EVA: An Embodied World Model for Future Video Anticipation
by: Chi, Xiaowei, et al.
Published: (2024)
by: Chi, Xiaowei, et al.
Published: (2024)
EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models
by: Du, Mengfei, et al.
Published: (2024)
by: Du, Mengfei, et al.
Published: (2024)
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces
by: E, Shaojun, et al.
Published: (2025)
by: E, Shaojun, et al.
Published: (2025)
WoW: Towards a World omniscient World model Through Embodied Interaction
by: Chi, Xiaowei, et al.
Published: (2025)
by: Chi, Xiaowei, et al.
Published: (2025)
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
by: Fang, Xinyu, et al.
Published: (2024)
by: Fang, Xinyu, et al.
Published: (2024)
Mixture-of-Prompt-Experts for Multi-modal Semantic Understanding
by: Wu, Zichen, et al.
Published: (2024)
by: Wu, Zichen, et al.
Published: (2024)
RealBench: A Chinese Multi-image Understanding Benchmark Close to Real-world Scenarios
by: Zhao, Fei, et al.
Published: (2025)
by: Zhao, Fei, et al.
Published: (2025)
Dynamic Multimodal Expression Generation for LLM-Driven Pedagogical Agents: From User Experience Perspective
by: Wan, Ninghao, et al.
Published: (2026)
by: Wan, Ninghao, et al.
Published: (2026)
KG-RAG: Enhancing GUI Agent Decision-Making via Knowledge Graph-Driven Retrieval-Augmented Generation
by: Guan, Ziyi, et al.
Published: (2025)
by: Guan, Ziyi, et al.
Published: (2025)
ResearchPulse: Building Method-Experiment Chains through Multi-Document Scientific Inference
by: Chen, Qi, et al.
Published: (2025)
by: Chen, Qi, et al.
Published: (2025)
Bridging the Gap: Sketch-Aware Interpolation Network for High-Quality Animation Sketch Inbetweening
by: Shen, Jiaming, et al.
Published: (2023)
by: Shen, Jiaming, et al.
Published: (2023)
Using Technology in Digital Humanities for Learning and Knowledge Dissemination
by: Rodrigues, Armanda, et al.
Published: (2025)
by: Rodrigues, Armanda, et al.
Published: (2025)
Multimodal Digital Sensing of Early-Life Laying Hens: A Pilot Study Integrating Thermal, Acoustic, Optical-Flow and Environmental Data
by: Dhaliwal, Yashan, et al.
Published: (2026)
by: Dhaliwal, Yashan, et al.
Published: (2026)
AesopAgent: Agent-driven Evolutionary System on Story-to-Video Production
by: Wang, Jiuniu, et al.
Published: (2024)
by: Wang, Jiuniu, et al.
Published: (2024)
SD-DiT: Unleashing the Power of Self-supervised Discrimination in Diffusion Transformer
by: Zhu, Rui, et al.
Published: (2024)
by: Zhu, Rui, et al.
Published: (2024)
Immersive Fantasy Based on Digital Nostalgia: Environmental Narratives for the Korean Millennials and Gen Z
by: Doh, Yerin, et al.
Published: (2025)
by: Doh, Yerin, et al.
Published: (2025)
RiverEcho: Real-Time Interactive Digital System for Ancient Yellow River Culture
by: Wang, Haofeng, et al.
Published: (2025)
by: Wang, Haofeng, et al.
Published: (2025)
AgentDisCo: Towards Disentanglement and Collaboration in Open-ended Deep Research Agents
by: Jin, Jiarui, et al.
Published: (2026)
by: Jin, Jiarui, et al.
Published: (2026)
VIoTGPT: Learning to Schedule Vision Tools in LLMs towards Intelligent Video Internet of Things
by: Zhong, Yaoyao, et al.
Published: (2023)
by: Zhong, Yaoyao, et al.
Published: (2023)
Speech as a Multimodal Digital Phenotype for Multi-Task LLM-based Mental Health Prediction
by: Ali, Mai, et al.
Published: (2025)
by: Ali, Mai, et al.
Published: (2025)
Anim-Director: A Large Multimodal Model Powered Agent for Controllable Animation Video Generation
by: Li, Yunxin, et al.
Published: (2024)
by: Li, Yunxin, et al.
Published: (2024)
Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents
by: Zhang, Xueqiao, et al.
Published: (2025)
by: Zhang, Xueqiao, et al.
Published: (2025)
Kubrick: Multimodal Agent Collaborations for Synthetic Video Generation
by: He, Liu, et al.
Published: (2024)
by: He, Liu, et al.
Published: (2024)
Similar Items
-
Detecting Multimodal Situations with Insufficient Context and Abstaining from Baseless Predictions
by: Liu, Junzhang, et al.
Published: (2024) -
Memory-Centric Embodied Question Answering
by: Zhai, Mingliang, et al.
Published: (2025) -
CoNav: Collaborative Cross-Modal Reasoning for Embodied Navigation
by: Hao, Haihong, et al.
Published: (2025) -
SentiAvatar: Towards Expressive and Interactive Digital Humans
by: Jin, Chuhao, et al.
Published: (2026) -
Integrating Multi-Modal Sensors: A Review of Fusion Techniques for Intelligent Vehicles
by: Wei, Chuheng, et al.
Published: (2025)