OmniGuide: Universal Guidance Fields for Enhancing Generalist Robot Policies
Fuente:
arXiv
Saved in:
| Main Authors: | Song, Yunzhou, Le, Long, Park, Yong-Hyun, Wang, Jie, Shi, Junyao, Liu, Lingjie, Gu, Jiatao, Eaton, Eric, Jayaraman, Dinesh, Daniilidis, Kostas |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Maestro: Orchestrating Robotics Modules with Vision-Language Models for Zero-Shot Generalist Robots
by: Shi, Junyao, et al.
Published: (2025)
by: Shi, Junyao, et al.
Published: (2025)
HDGS: Textured 2D Gaussian Splatting for Enhanced Scene Rendering
by: Song, Yunzhou, et al.
Published: (2024)
by: Song, Yunzhou, et al.
Published: (2024)
Track Everything Everywhere Fast and Robustly
by: Song, Yunzhou, et al.
Published: (2024)
by: Song, Yunzhou, et al.
Published: (2024)
Zero-1-to-G: Taming Pretrained 2D Diffusion Model for Direct 3D Generation
by: Meng, Xuyi, et al.
Published: (2025)
by: Meng, Xuyi, et al.
Published: (2025)
Pixie: Fast and Generalizable Supervised Learning of 3D Physics from Pixels
by: Le, Long, et al.
Published: (2025)
by: Le, Long, et al.
Published: (2025)
TRAM: Global Trajectory and Motion of 3D Humans from in-the-wild Videos
by: Wang, Yufu, et al.
Published: (2024)
by: Wang, Yufu, et al.
Published: (2024)
Composing Pre-Trained Object-Centric Representations for Robotics From "What" and "Where" Foundation Models
by: Shi, Junyao, et al.
Published: (2024)
by: Shi, Junyao, et al.
Published: (2024)
RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies
by: Atreya, Pranav, et al.
Published: (2025)
by: Atreya, Pranav, et al.
Published: (2025)
StereoDiff: Stereo-Diffusion Synergy for Video Depth Estimation
by: Li, Haodong, et al.
Published: (2025)
by: Li, Haodong, et al.
Published: (2025)
DIMO: Diverse 3D Motion Generation for Arbitrary Objects
by: Mou, Linzhan, et al.
Published: (2025)
by: Mou, Linzhan, et al.
Published: (2025)
ActiveGrasp: Information-Guided Active Grasping with Calibrated Energy-based Model
by: Lei, Boshu, et al.
Published: (2025)
by: Lei, Boshu, et al.
Published: (2025)
Fast Feature Field ($\text{F}^3$): A Predictive Representation of Events
by: Das, Richeek, et al.
Published: (2025)
by: Das, Richeek, et al.
Published: (2025)
FisherRF: Active View Selection and Uncertainty Quantification for Radiance Fields using Fisher Information
by: Jiang, Wen, et al.
Published: (2023)
by: Jiang, Wen, et al.
Published: (2023)
REGENT: A Retrieval-Augmented Generalist Agent That Can Act In-Context in New Environments
by: Sridhar, Kaustubh, et al.
Published: (2024)
by: Sridhar, Kaustubh, et al.
Published: (2024)
PhysHMR: Learning Humanoid Control Policies from Vision for Physically Plausible Human Motion Reconstruction
by: Feng, Qiao, et al.
Published: (2025)
by: Feng, Qiao, et al.
Published: (2025)
GECO: Generative Image-to-3D within a SECOnd
by: Wang, Chen, et al.
Published: (2024)
by: Wang, Chen, et al.
Published: (2024)
Zero-shot Reconstruction of In-Scene Object Manipulation from Video
by: Lin, Dixuan, et al.
Published: (2025)
by: Lin, Dixuan, et al.
Published: (2025)
Multimodal LLM Guided Exploration and Active Mapping using Fisher Information
by: Jiang, Wen, et al.
Published: (2024)
by: Jiang, Wen, et al.
Published: (2024)
OmniPlanner: Universal Exploration and Inspection Path Planning across Robot Morphologies
by: Zacharia, Angelos, et al.
Published: (2026)
by: Zacharia, Angelos, et al.
Published: (2026)
VLMgineer: Vision Language Models as Robotic Toolsmiths
by: Gao, George Jiayuan, et al.
Published: (2025)
by: Gao, George Jiayuan, et al.
Published: (2025)
Un-EVIMO: Unsupervised Event-Based Independent Motion Segmentation
by: Wang, Ziyun, et al.
Published: (2023)
by: Wang, Ziyun, et al.
Published: (2023)
Next Best View Selections for Semantic and Dynamic 3D Gaussian Splatting
by: Li, Yiqian, et al.
Published: (2025)
by: Li, Yiqian, et al.
Published: (2025)
BiEquiFormer: Bi-Equivariant Representations for Global Point Cloud Registration
by: Pertigkiozoglou, Stefanos, et al.
Published: (2024)
by: Pertigkiozoglou, Stefanos, et al.
Published: (2024)
DynMF: Neural Motion Factorization for Real-time Dynamic View Synthesis with 3D Gaussian Splatting
by: Kratimenos, Agelos, et al.
Published: (2023)
by: Kratimenos, Agelos, et al.
Published: (2023)
ZeroMimic: Distilling Robotic Manipulation Skills from Web Videos
by: Shi, Junyao, et al.
Published: (2025)
by: Shi, Junyao, et al.
Published: (2025)
Grokking of Diffusion Models: Case Study on Modular Addition
by: Kim, Joon Hyeok, et al.
Published: (2026)
by: Kim, Joon Hyeok, et al.
Published: (2026)
Recasting Generic Pretrained Vision Transformers As Object-Centric Scene Encoders For Manipulation Policies
by: Qian, Jianing, et al.
Published: (2024)
by: Qian, Jianing, et al.
Published: (2024)
Distributed Continual Learning
by: Le, Long, et al.
Published: (2024)
by: Le, Long, et al.
Published: (2024)
Continuous-Time Human Motion Field from Events
by: Wang, Ziyun, et al.
Published: (2024)
by: Wang, Ziyun, et al.
Published: (2024)
Uncertainty-Aware Deployment of Pre-trained Language-Conditioned Imitation Learning Policies
by: Wu, Bo, et al.
Published: (2024)
by: Wu, Bo, et al.
Published: (2024)
TLControl: Trajectory and Language Control for Human Motion Synthesis
by: Wan, Weilin, et al.
Published: (2023)
by: Wan, Weilin, et al.
Published: (2023)
Octo: An Open-Source Generalist Robot Policy
by: Octo Model Team, et al.
Published: (2024)
by: Octo Model Team, et al.
Published: (2024)
Turning Video Models into Generalist Robot Policies
by: Li, Sizhe Lester, et al.
Published: (2026)
by: Li, Sizhe Lester, et al.
Published: (2026)
ReFineVLA: Multimodal Reasoning-Aware Generalist Robotic Policies via Teacher-Guided Fine-Tuning
by: Van Vo, Tuan, et al.
Published: (2026)
by: Van Vo, Tuan, et al.
Published: (2026)
OmniVTON++: Training-Free Universal Virtual Try-On with Principal Pose Guidance
by: Yang, Zhaotong, et al.
Published: (2026)
by: Yang, Zhaotong, et al.
Published: (2026)
Articulate-Anything: Automatic Modeling of Articulated Objects via a Vision-Language Foundation Model
by: Le, Long, et al.
Published: (2024)
by: Le, Long, et al.
Published: (2024)
Omni-Reward: Towards Generalist Omni-Modal Reward Modeling with Free-Form Preferences
by: Jin, Zhuoran, et al.
Published: (2025)
by: Jin, Zhuoran, et al.
Published: (2025)
Symmetries-enhanced Multi-Agent Reinforcement Learning
by: Bousias, Nikolaos, et al.
Published: (2025)
by: Bousias, Nikolaos, et al.
Published: (2025)
Recurrent Equivariant Constraint Modulation: Learning Per-Layer Symmetry Relaxation from Data
by: Pertigkiozoglou, Stefanos, et al.
Published: (2026)
by: Pertigkiozoglou, Stefanos, et al.
Published: (2026)
Match-Any-Events: Zero-Shot Motion-Robust Feature Matching Across Wide Baselines for Event Cameras
by: Zhang, Ruijun, et al.
Published: (2026)
by: Zhang, Ruijun, et al.
Published: (2026)
Similar Items
-
Maestro: Orchestrating Robotics Modules with Vision-Language Models for Zero-Shot Generalist Robots
by: Shi, Junyao, et al.
Published: (2025) -
HDGS: Textured 2D Gaussian Splatting for Enhanced Scene Rendering
by: Song, Yunzhou, et al.
Published: (2024) -
Track Everything Everywhere Fast and Robustly
by: Song, Yunzhou, et al.
Published: (2024) -
Zero-1-to-G: Taming Pretrained 2D Diffusion Model for Direct 3D Generation
by: Meng, Xuyi, et al.
Published: (2025) -
Pixie: Fast and Generalizable Supervised Learning of 3D Physics from Pixels
by: Le, Long, et al.
Published: (2025)