Visually-grounded Humanoid Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Ye, Hang, Ma, Xiaoxuan, Lu, Fan, Wu, Wayne, Lin, Kwan-Yee, Wang, Yizhou |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Let Humanoids Hike! Integrative Skill Development on Complex Trails
by: Lin, Kwan-Yee, et al.
Published: (2025)
by: Lin, Kwan-Yee, et al.
Published: (2025)
Humanoid-VLA: Towards Universal Humanoid Control with Visual Integration
by: Ding, Pengxiang, et al.
Published: (2025)
by: Ding, Pengxiang, et al.
Published: (2025)
Real-time Holistic Robot Pose Estimation with Unknown States
by: Ban, Shikun, et al.
Published: (2024)
by: Ban, Shikun, et al.
Published: (2024)
SUGAR: A Scalable Human-Video-Driven Generalizable Humanoid Loco-Manipulation Learning Framework
by: Wu, Tianshu, et al.
Published: (2026)
by: Wu, Tianshu, et al.
Published: (2026)
Emergent Active Perception and Dexterity of Simulated Humanoids from Visual Reinforcement Learning
by: Luo, Zhengyi, et al.
Published: (2025)
by: Luo, Zhengyi, et al.
Published: (2025)
Visual Imitation Enables Contextual Humanoid Control
by: Allshire, Arthur, et al.
Published: (2025)
by: Allshire, Arthur, et al.
Published: (2025)
InsMapper: Exploring Inner-instance Information for Vectorized HD Mapping
by: Xu, Zhenhua, et al.
Published: (2023)
by: Xu, Zhenhua, et al.
Published: (2023)
UniAct: Unified Motion Generation and Action Streaming for Humanoid Robots
by: Jiang, Nan, et al.
Published: (2025)
by: Jiang, Nan, et al.
Published: (2025)
SynAgent: Generalizable Cooperative Humanoid Manipulation via Solo-to-Cooperative Agent Synergy
by: Yao, Wei, et al.
Published: (2026)
by: Yao, Wei, et al.
Published: (2026)
High-Precision Transformer-Based Visual Servoing for Humanoid Robots in Aligning Tiny Objects
by: Xue, Jialong, et al.
Published: (2025)
by: Xue, Jialong, et al.
Published: (2025)
SMPLOlympics: Sports Environments for Physically Simulated Humanoids
by: Luo, Zhengyi, et al.
Published: (2024)
by: Luo, Zhengyi, et al.
Published: (2024)
Learning Humanoid End-Effector Control for Open-Vocabulary Visual Loco-Manipulation
by: Dong, Runpei, et al.
Published: (2026)
by: Dong, Runpei, et al.
Published: (2026)
VisualMimic: Visual Humanoid Loco-Manipulation via Motion Tracking and Generation
by: Yin, Shaofeng, et al.
Published: (2025)
by: Yin, Shaofeng, et al.
Published: (2025)
TrajBooster: Boosting Humanoid Whole-Body Manipulation via Trajectory-Centric Learning
by: Liu, Jiacheng, et al.
Published: (2025)
by: Liu, Jiacheng, et al.
Published: (2025)
EgoActor: Grounding Task Planning into Spatial-aware Egocentric Actions for Humanoid Robots via Visual-Language Models
by: Bai, Yu, et al.
Published: (2026)
by: Bai, Yu, et al.
Published: (2026)
From Seeing to Experiencing: Scaling Navigation Foundation Models with Reinforcement Learning
by: He, Honglin, et al.
Published: (2025)
by: He, Honglin, et al.
Published: (2025)
Hierarchical World Models as Visual Whole-Body Humanoid Controllers
by: Hansen, Nicklas, et al.
Published: (2024)
by: Hansen, Nicklas, et al.
Published: (2024)
RichControl: Structure- and Appearance-Rich Training-Free Spatial Control for Text-to-Image Generation
by: Pang, Lexi, et al.
Published: (2025)
by: Pang, Lexi, et al.
Published: (2025)
Opening the Sim-to-Real Door for Humanoid Pixel-to-Action Policy Transfer
by: Xue, Haoru, et al.
Published: (2025)
by: Xue, Haoru, et al.
Published: (2025)
VISC: mmWave Radar Scene Flow Estimation using Pervasive Visual-Inertial Supervision
by: Liu, Kezhong, et al.
Published: (2025)
by: Liu, Kezhong, et al.
Published: (2025)
Attentive Feature Aggregation or: How Policies Learn to Stop Worrying about Robustness and Attend to Task-Relevant Visual Cues
by: Tsagkas, Nikolaos, et al.
Published: (2025)
by: Tsagkas, Nikolaos, et al.
Published: (2025)
Iterative Closed-Loop Motion Synthesis for Scaling the Capabilities of Humanoid Control
by: Xu, Weisheng, et al.
Published: (2026)
by: Xu, Weisheng, et al.
Published: (2026)
Robust 3D Object Detection from LiDAR-Radar Point Clouds via Cross-Modal Feature Augmentation
by: Deng, Jianning, et al.
Published: (2023)
by: Deng, Jianning, et al.
Published: (2023)
Before the Body Moves: Learning Anticipatory Joint Intent for Language-Conditioned Humanoid Control
by: Jia, Haozhe, et al.
Published: (2026)
by: Jia, Haozhe, et al.
Published: (2026)
VR-Robo: A Real-to-Sim-to-Real Framework for Visual Robot Navigation and Locomotion
by: Zhu, Shaoting, et al.
Published: (2025)
by: Zhu, Shaoting, et al.
Published: (2025)
PDF-HR: Pose Distance Fields for Humanoid Robots
by: Gu, Yi, et al.
Published: (2026)
by: Gu, Yi, et al.
Published: (2026)
From Motion to Behavior: Hierarchical Modeling of Humanoid Generative Behavior Control
by: Zhang, Jusheng, et al.
Published: (2025)
by: Zhang, Jusheng, et al.
Published: (2025)
Learning Sidewalk Autopilot from Multi-Scale Imitation with Corrective Behavior Expansion
by: He, Honglin, et al.
Published: (2026)
by: He, Honglin, et al.
Published: (2026)
VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
by: Xu, Runsen, et al.
Published: (2024)
by: Xu, Runsen, et al.
Published: (2024)
Humanoid Occupancy: Enabling A Generalized Multimodal Occupancy Perception System on Humanoid Robots
by: Cui, Wei, et al.
Published: (2025)
by: Cui, Wei, et al.
Published: (2025)
RoboMirror: Understand Before You Imitate for Video to Humanoid Locomotion
by: Li, Zhe, et al.
Published: (2025)
by: Li, Zhe, et al.
Published: (2025)
Click to Grasp: Zero-Shot Precise Manipulation via Visual Diffusion Descriptors
by: Tsagkas, Nikolaos, et al.
Published: (2024)
by: Tsagkas, Nikolaos, et al.
Published: (2024)
RoboView-Bias: Benchmarking Visual Bias in Embodied Agents for Robotic Manipulation
by: Liu, Enguang, et al.
Published: (2025)
by: Liu, Enguang, et al.
Published: (2025)
From Language to Locomotion: Retargeting-free Humanoid Control via Motion Latent Guidance
by: Li, Zhe, et al.
Published: (2025)
by: Li, Zhe, et al.
Published: (2025)
ULTRA: Unified Multimodal Control for Autonomous Humanoid Whole-Body Loco-Manipulation
by: He, Xialin, et al.
Published: (2026)
by: He, Xialin, et al.
Published: (2026)
Image Generation as a Visual Planner for Robotic Manipulation
by: Pang, Ye
Published: (2025)
by: Pang, Ye
Published: (2025)
MapGPT: Map-Guided Prompting with Adaptive Path Planning for Vision-and-Language Navigation
by: Chen, Jiaqi, et al.
Published: (2024)
by: Chen, Jiaqi, et al.
Published: (2024)
DriveGPT4: Interpretable End-to-end Autonomous Driving via Large Language Model
by: Xu, Zhenhua, et al.
Published: (2023)
by: Xu, Zhenhua, et al.
Published: (2023)
Mimicking-Bench: A Benchmark for Generalizable Humanoid-Scene Interaction Learning via Human Mimicking
by: Liu, Yun, et al.
Published: (2024)
by: Liu, Yun, et al.
Published: (2024)
WeatherPrompt: Multi-modality Representation Learning for All-Weather Drone Visual Geo-Localization
by: Wen, Jiahao, et al.
Published: (2025)
by: Wen, Jiahao, et al.
Published: (2025)
Similar Items
-
Let Humanoids Hike! Integrative Skill Development on Complex Trails
by: Lin, Kwan-Yee, et al.
Published: (2025) -
Humanoid-VLA: Towards Universal Humanoid Control with Visual Integration
by: Ding, Pengxiang, et al.
Published: (2025) -
Real-time Holistic Robot Pose Estimation with Unknown States
by: Ban, Shikun, et al.
Published: (2024) -
SUGAR: A Scalable Human-Video-Driven Generalizable Humanoid Loco-Manipulation Learning Framework
by: Wu, Tianshu, et al.
Published: (2026) -
Emergent Active Perception and Dexterity of Simulated Humanoids from Visual Reinforcement Learning
by: Luo, Zhengyi, et al.
Published: (2025)