From Instruction to Event: Sound-Triggered Mobile Manipulation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ju, Hao, Huang, Shaofei, Li, Hongyu, Ding, Zihan, Liu, Si, Wang, Meng, Zheng, Zhedong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Video2BEV: Transforming Drone Videos to BEVs for Video-based Geo-localization
von: Ju, Hao, et al.
Veröffentlicht: (2024)
von: Ju, Hao, et al.
Veröffentlicht: (2024)
Echo Planning for Autonomous Driving: From Current Observations to Future Trajectories and Back
von: Sun, Jintao, et al.
Veröffentlicht: (2025)
von: Sun, Jintao, et al.
Veröffentlicht: (2025)
Learning Actionable Manipulation Recovery via Counterfactual Failure Synthesis
von: Li, Dayou, et al.
Veröffentlicht: (2026)
von: Li, Dayou, et al.
Veröffentlicht: (2026)
WeatherPrompt: Multi-modality Representation Learning for All-Weather Drone Visual Geo-Localization
von: Wen, Jiahao, et al.
Veröffentlicht: (2025)
von: Wen, Jiahao, et al.
Veröffentlicht: (2025)
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
von: Yang, Shuai, et al.
Veröffentlicht: (2025)
von: Yang, Shuai, et al.
Veröffentlicht: (2025)
Object-Centric Instruction Augmentation for Robotic Manipulation
von: Wen, Junjie, et al.
Veröffentlicht: (2024)
von: Wen, Junjie, et al.
Veröffentlicht: (2024)
Event Camera Meets Mobile Embodied Perception: Abstraction, Algorithm, Acceleration, Application
von: Wang, Haoyang, et al.
Veröffentlicht: (2025)
von: Wang, Haoyang, et al.
Veröffentlicht: (2025)
Scalable Trajectory Generation for Whole-Body Mobile Manipulation
von: Niu, Yida, et al.
Veröffentlicht: (2026)
von: Niu, Yida, et al.
Veröffentlicht: (2026)
BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation
von: Wang, Hongyu, et al.
Veröffentlicht: (2025)
von: Wang, Hongyu, et al.
Veröffentlicht: (2025)
eKalibr: Dynamic Intrinsic Calibration for Event Cameras From First Principles of Events
von: Chen, Shuolong, et al.
Veröffentlicht: (2025)
von: Chen, Shuolong, et al.
Veröffentlicht: (2025)
Towards Open-World Mobile Manipulation in Homes: Lessons from the Neurips 2023 HomeRobot Open Vocabulary Mobile Manipulation Challenge
von: Yenamandra, Sriram, et al.
Veröffentlicht: (2024)
von: Yenamandra, Sriram, et al.
Veröffentlicht: (2024)
Robotic Programmer: Video Instructed Policy Code Generation for Robotic Manipulation
von: Xie, Senwei, et al.
Veröffentlicht: (2025)
von: Xie, Senwei, et al.
Veröffentlicht: (2025)
Neural Assembler: Learning to Generate Fine-Grained Robotic Assembly Instructions from Multi-View Images
von: Yan, Hongyu, et al.
Veröffentlicht: (2024)
von: Yan, Hongyu, et al.
Veröffentlicht: (2024)
BFA: Best-Feature-Aware Fusion for Multi-View Fine-grained Manipulation
von: Lan, Zihan, et al.
Veröffentlicht: (2025)
von: Lan, Zihan, et al.
Veröffentlicht: (2025)
MobileVLA-R1: Reinforcing Vision-Language-Action for Mobile Robots
von: Huang, Ting, et al.
Veröffentlicht: (2025)
von: Huang, Ting, et al.
Veröffentlicht: (2025)
ODTFormer: Efficient Obstacle Detection and Tracking with Stereo Cameras Based on Transformer
von: Ding, Tianye, et al.
Veröffentlicht: (2024)
von: Ding, Tianye, et al.
Veröffentlicht: (2024)
Pretrain-then-Adapt: Uncertainty-Aware Test-Time Adaptation for Text-based Person Search
von: Zhang, Jiahao, et al.
Veröffentlicht: (2026)
von: Zhang, Jiahao, et al.
Veröffentlicht: (2026)
RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation
von: Han, Songhao, et al.
Veröffentlicht: (2025)
von: Han, Songhao, et al.
Veröffentlicht: (2025)
Towards the Vision-Sound-Language-Action Paradigm: The HEAR Framework for Sound-Centric Manipulation
von: Nie, Chang, et al.
Veröffentlicht: (2026)
von: Nie, Chang, et al.
Veröffentlicht: (2026)
MoMa-Kitchen: A 100K+ Benchmark for Affordance-Grounded Last-Mile Navigation in Mobile Manipulation
von: Zhang, Pingrui, et al.
Veröffentlicht: (2025)
von: Zhang, Pingrui, et al.
Veröffentlicht: (2025)
ODYSSEY: Open-World Quadrupeds Exploration and Manipulation for Long-Horizon Tasks
von: Wang, Kaijun, et al.
Veröffentlicht: (2025)
von: Wang, Kaijun, et al.
Veröffentlicht: (2025)
Neuro-Symbolic Manipulation Understanding with Enriched Semantic Event Chains
von: Ziaeetabar, Fatemeh
Veröffentlicht: (2026)
von: Ziaeetabar, Fatemeh
Veröffentlicht: (2026)
ESCAPE: Episodic Spatial Memory and Adaptive Execution Policy for Long-Horizon Mobile Manipulation
von: Qian, Jingjing, et al.
Veröffentlicht: (2026)
von: Qian, Jingjing, et al.
Veröffentlicht: (2026)
GAMMA: Generalizable Articulation Modeling and Manipulation for Articulated Objects
von: Yu, Qiaojun, et al.
Veröffentlicht: (2023)
von: Yu, Qiaojun, et al.
Veröffentlicht: (2023)
MoManipVLA: Transferring Vision-language-action Models for General Mobile Manipulation
von: Wu, Zhenyu, et al.
Veröffentlicht: (2025)
von: Wu, Zhenyu, et al.
Veröffentlicht: (2025)
EV-MGDispNet: Motion-Guided Event-Based Stereo Disparity Estimation Network with Left-Right Consistency
von: Jiang, Junjie, et al.
Veröffentlicht: (2024)
von: Jiang, Junjie, et al.
Veröffentlicht: (2024)
M4Diffuser: Multi-View Diffusion Policy with Manipulability-Aware Control for Robust Mobile Manipulation
von: Dong, Ju, et al.
Veröffentlicht: (2025)
von: Dong, Ju, et al.
Veröffentlicht: (2025)
Demystifying Action Space Design for Robotic Manipulation Policies
von: Feng, Yuchun, et al.
Veröffentlicht: (2026)
von: Feng, Yuchun, et al.
Veröffentlicht: (2026)
RoboGround: Robotic Manipulation with Grounded Vision-Language Priors
von: Huang, Haifeng, et al.
Veröffentlicht: (2025)
von: Huang, Haifeng, et al.
Veröffentlicht: (2025)
EgoSpot:Egocentric Multimodal Control for Hands-Free Mobile Manipulation
von: Zhang, Ganlin, et al.
Veröffentlicht: (2023)
von: Zhang, Ganlin, et al.
Veröffentlicht: (2023)
SpatialActor: Exploring Disentangled Spatial Representations for Robust Robotic Manipulation
von: Shi, Hao, et al.
Veröffentlicht: (2025)
von: Shi, Hao, et al.
Veröffentlicht: (2025)
ManiSoft: Towards Vision-Language Manipulation for Soft Continuum Robotics
von: Wei, Ziyu, et al.
Veröffentlicht: (2026)
von: Wei, Ziyu, et al.
Veröffentlicht: (2026)
TLA: Tactile-Language-Action Model for Contact-Rich Manipulation
von: Hao, Peng, et al.
Veröffentlicht: (2025)
von: Hao, Peng, et al.
Veröffentlicht: (2025)
SEBVS: Synthetic Event-based Visual Servoing for Robot Navigation and Manipulation
von: Vinod, Krishna, et al.
Veröffentlicht: (2025)
von: Vinod, Krishna, et al.
Veröffentlicht: (2025)
Event-Based Visual Odometry on Non-Holonomic Ground Vehicles
von: Xu, Wanting, et al.
Veröffentlicht: (2024)
von: Xu, Wanting, et al.
Veröffentlicht: (2024)
ForceVLA: Enhancing VLA Models with a Force-aware MoE for Contact-rich Manipulation
von: Yu, Jiawen, et al.
Veröffentlicht: (2025)
von: Yu, Jiawen, et al.
Veröffentlicht: (2025)
ManipTrans: Efficient Dexterous Bimanual Manipulation Transfer via Residual Learning
von: Li, Kailin, et al.
Veröffentlicht: (2025)
von: Li, Kailin, et al.
Veröffentlicht: (2025)
AnomalyLMM: Bridging Generative Knowledge and Discriminative Retrieval for Text-Based Person Anomaly Search
von: Ju, Hao, et al.
Veröffentlicht: (2025)
von: Ju, Hao, et al.
Veröffentlicht: (2025)
Learning Manipulation by Predicting Interaction
von: Zeng, Jia, et al.
Veröffentlicht: (2024)
von: Zeng, Jia, et al.
Veröffentlicht: (2024)
SEM: Enhancing Spatial Understanding for Robust Robot Manipulation
von: Lin, Xuewu, et al.
Veröffentlicht: (2025)
von: Lin, Xuewu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Video2BEV: Transforming Drone Videos to BEVs for Video-based Geo-localization
von: Ju, Hao, et al.
Veröffentlicht: (2024) -
Echo Planning for Autonomous Driving: From Current Observations to Future Trajectories and Back
von: Sun, Jintao, et al.
Veröffentlicht: (2025) -
Learning Actionable Manipulation Recovery via Counterfactual Failure Synthesis
von: Li, Dayou, et al.
Veröffentlicht: (2026) -
WeatherPrompt: Multi-modality Representation Learning for All-Weather Drone Visual Geo-Localization
von: Wen, Jiahao, et al.
Veröffentlicht: (2025) -
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
von: Yang, Shuai, et al.
Veröffentlicht: (2025)