Instruction-Guided Visual Masking
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zheng, Jinliang, Li, Jianxiong, Cheng, Sijie, Zheng, Yinan, Li, Jiaming, Liu, Jihao, Liu, Yu, Liu, Jingjing, Zhan, Xianyuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DecisionNCE: Embodied Multimodal Representations via Implicit Preference Learning
von: Li, Jianxiong, et al.
Veröffentlicht: (2024)
von: Li, Jianxiong, et al.
Veröffentlicht: (2024)
Universal Actions for Enhanced Embodied Foundation Models
von: Zheng, Jinliang, et al.
Veröffentlicht: (2025)
von: Zheng, Jinliang, et al.
Veröffentlicht: (2025)
Efficient Robotic Policy Learning via Latent Space Backward Planning
von: Liu, Dongxiu, et al.
Veröffentlicht: (2025)
von: Liu, Dongxiu, et al.
Veröffentlicht: (2025)
X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model
von: Zheng, Jinliang, et al.
Veröffentlicht: (2025)
von: Zheng, Jinliang, et al.
Veröffentlicht: (2025)
Safe Offline Reinforcement Learning with Feasibility-Guided Diffusion Model
von: Zheng, Yinan, et al.
Veröffentlicht: (2024)
von: Zheng, Yinan, et al.
Veröffentlicht: (2024)
MM-Instruct: Generated Visual Instructions for Large Multimodal Model Alignment
von: Liu, Jihao, et al.
Veröffentlicht: (2024)
von: Liu, Jihao, et al.
Veröffentlicht: (2024)
Diffusion-Based Planning for Autonomous Driving with Flexible Guidance
von: Zheng, Yinan, et al.
Veröffentlicht: (2025)
von: Zheng, Yinan, et al.
Veröffentlicht: (2025)
Flow Matching-Based Autonomous Driving Planning with Advanced Interactive Behavior Modeling
von: Tan, Tianyi, et al.
Veröffentlicht: (2025)
von: Tan, Tianyi, et al.
Veröffentlicht: (2025)
Demystifying Action Space Design for Robotic Manipulation Policies
von: Feng, Yuchun, et al.
Veröffentlicht: (2026)
von: Feng, Yuchun, et al.
Veröffentlicht: (2026)
Robo-MUTUAL: Robotic Multimodal Task Specification via Unimodal Learning
von: Li, Jianxiong, et al.
Veröffentlicht: (2024)
von: Li, Jianxiong, et al.
Veröffentlicht: (2024)
Skill Expansion and Composition in Parameter Space
von: Liu, Tenglong, et al.
Veröffentlicht: (2025)
von: Liu, Tenglong, et al.
Veröffentlicht: (2025)
PhysiAgent: An Embodied Agent Framework in Physical World
von: Wang, Zhihao, et al.
Veröffentlicht: (2025)
von: Wang, Zhihao, et al.
Veröffentlicht: (2025)
Robotic Visual Instruction
von: Li, Yanbang, et al.
Veröffentlicht: (2025)
von: Li, Yanbang, et al.
Veröffentlicht: (2025)
Vega: Learning to Drive with Natural Language Instructions
von: Zuo, Sicheng, et al.
Veröffentlicht: (2026)
von: Zuo, Sicheng, et al.
Veröffentlicht: (2026)
Towards Robust Zero-Shot Reinforcement Learning
von: Zheng, Kexin, et al.
Veröffentlicht: (2025)
von: Zheng, Kexin, et al.
Veröffentlicht: (2025)
IKEA Manuals at Work: 4D Grounding of Assembly Instructions on Internet Videos
von: Liu, Yunong, et al.
Veröffentlicht: (2024)
von: Liu, Yunong, et al.
Veröffentlicht: (2024)
FloNa: Floor Plan Guided Embodied Visual Navigation
von: Li, Jiaxin, et al.
Veröffentlicht: (2024)
von: Li, Jiaxin, et al.
Veröffentlicht: (2024)
Good Token Hunting: A Hitchhiker's Guide to Token Selection for Visual Geometry Transformers
von: Zheng, Shuhong, et al.
Veröffentlicht: (2026)
von: Zheng, Shuhong, et al.
Veröffentlicht: (2026)
ComposableNav: Instruction-Following Navigation in Dynamic Environments via Composable Diffusion
von: Hu, Zichao, et al.
Veröffentlicht: (2025)
von: Hu, Zichao, et al.
Veröffentlicht: (2025)
Faster or Stronger: Towards Flexible Visual Place Recognition via Weighted Aggregation and Token Pruning
von: Zeng, Zichao, et al.
Veröffentlicht: (2026)
von: Zeng, Zichao, et al.
Veröffentlicht: (2026)
SyncDiff: Synchronized Motion Diffusion for Multi-Body Human-Object Interaction Synthesis
von: He, Wenkun, et al.
Veröffentlicht: (2024)
von: He, Wenkun, et al.
Veröffentlicht: (2024)
PhysInOne: Visual Physics Learning and Reasoning in One Suite
von: Zhou, Siyuan, et al.
Veröffentlicht: (2026)
von: Zhou, Siyuan, et al.
Veröffentlicht: (2026)
StarVLA-$α$: Reducing Complexity in Vision-Language-Action Systems
von: Ye, Jinhui, et al.
Veröffentlicht: (2026)
von: Ye, Jinhui, et al.
Veröffentlicht: (2026)
Closed Loop Dynamic Driving Data Mixture for Real-Synthetic Co-Training
von: Ruan, Hongzhi, et al.
Veröffentlicht: (2026)
von: Ruan, Hongzhi, et al.
Veröffentlicht: (2026)
Visual Sync: Multi-Camera Synchronization via Cross-View Object Motion
von: Liu, Shaowei, et al.
Veröffentlicht: (2025)
von: Liu, Shaowei, et al.
Veröffentlicht: (2025)
Tidiness Score-Guided Monte Carlo Tree Search for Visual Tabletop Rearrangement
von: Kee, Hogun, et al.
Veröffentlicht: (2025)
von: Kee, Hogun, et al.
Veröffentlicht: (2025)
ReconDreamer: Crafting World Models for Driving Scene Reconstruction via Online Restoration
von: Ni, Chaojun, et al.
Veröffentlicht: (2024)
von: Ni, Chaojun, et al.
Veröffentlicht: (2024)
CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models
von: Zhao, Qingqing, et al.
Veröffentlicht: (2025)
von: Zhao, Qingqing, et al.
Veröffentlicht: (2025)
Adversarial Exploitation of Data Diversity Improves Visual Localization
von: Li, Sihang, et al.
Veröffentlicht: (2024)
von: Li, Sihang, et al.
Veröffentlicht: (2024)
SIMPL: A Simple and Efficient Multi-agent Motion Prediction Baseline for Autonomous Driving
von: Zhang, Lu, et al.
Veröffentlicht: (2024)
von: Zhang, Lu, et al.
Veröffentlicht: (2024)
Attention-Guided Integration of CLIP and SAM for Precise Object Masking in Robotic Manipulation
von: Muttaqien, Muhammad A., et al.
Veröffentlicht: (2025)
von: Muttaqien, Muhammad A., et al.
Veröffentlicht: (2025)
LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding
von: Wang, Shihao, et al.
Veröffentlicht: (2026)
von: Wang, Shihao, et al.
Veröffentlicht: (2026)
Self-Correcting VLA: Online Action Refinement via Sparse World Imagination
von: Liu, Chenyv, et al.
Veröffentlicht: (2026)
von: Liu, Chenyv, et al.
Veröffentlicht: (2026)
When Should We Prefer State-to-Visual DAgger Over Visual Reinforcement Learning?
von: Mu, Tongzhou, et al.
Veröffentlicht: (2024)
von: Mu, Tongzhou, et al.
Veröffentlicht: (2024)
EgoTraj-Bench: Towards Robust Trajectory Prediction Under Ego-view Noisy Observations
von: Liu, Jiayi, et al.
Veröffentlicht: (2025)
von: Liu, Jiayi, et al.
Veröffentlicht: (2025)
Gradient-Guided Parameter Mask for Multi-Scenario Image Restoration Under Adverse Weather
von: Guo, Jilong, et al.
Veröffentlicht: (2024)
von: Guo, Jilong, et al.
Veröffentlicht: (2024)
From Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation Priors
von: Zhang, Zhengshen, et al.
Veröffentlicht: (2025)
von: Zhang, Zhengshen, et al.
Veröffentlicht: (2025)
RoboFactory: Exploring Embodied Agent Collaboration with Compositional Constraints
von: Qin, Yiran, et al.
Veröffentlicht: (2025)
von: Qin, Yiran, et al.
Veröffentlicht: (2025)
SlotLifter: Slot-guided Feature Lifting for Learning Object-centric Radiance Fields
von: Liu, Yu, et al.
Veröffentlicht: (2024)
von: Liu, Yu, et al.
Veröffentlicht: (2024)
DexTrack: Towards Generalizable Neural Tracking Control for Dexterous Manipulation from Human References
von: Liu, Xueyi, et al.
Veröffentlicht: (2025)
von: Liu, Xueyi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
DecisionNCE: Embodied Multimodal Representations via Implicit Preference Learning
von: Li, Jianxiong, et al.
Veröffentlicht: (2024) -
Universal Actions for Enhanced Embodied Foundation Models
von: Zheng, Jinliang, et al.
Veröffentlicht: (2025) -
Efficient Robotic Policy Learning via Latent Space Backward Planning
von: Liu, Dongxiu, et al.
Veröffentlicht: (2025) -
X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model
von: Zheng, Jinliang, et al.
Veröffentlicht: (2025) -
Safe Offline Reinforcement Learning with Feasibility-Guided Diffusion Model
von: Zheng, Yinan, et al.
Veröffentlicht: (2024)