Control-oriented Clustering of Visual Latent Representation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Qi, Han, Yin, Haocheng, Yang, Heng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Inference-Time Enhancement of Generative Robot Policies via Predictive World Modeling
von: Qi, Han, et al.
Veröffentlicht: (2025)
von: Qi, Han, et al.
Veröffentlicht: (2025)
Language-Enhanced Latent Representations for Out-of-Distribution Detection in Autonomous Driving
von: Mao, Zhenjiang, et al.
Veröffentlicht: (2024)
von: Mao, Zhenjiang, et al.
Veröffentlicht: (2024)
Latent Object Characteristics Recognition with Visual to Haptic-Audio Cross-modal Transfer Learning
von: Saito, Namiko, et al.
Veröffentlicht: (2024)
von: Saito, Namiko, et al.
Veröffentlicht: (2024)
Visual Whole-Body Control for Legged Loco-Manipulation
von: Liu, Minghuan, et al.
Veröffentlicht: (2024)
von: Liu, Minghuan, et al.
Veröffentlicht: (2024)
Grounding Bodily Awareness in Visual Representations for Efficient Policy Learning
von: Wang, Junlin, et al.
Veröffentlicht: (2025)
von: Wang, Junlin, et al.
Veröffentlicht: (2025)
V-MORALS: Visual Morse Graph-Aided Estimation of Regions of Attraction in a Learned Latent Space
von: Aladin, Faiz, et al.
Veröffentlicht: (2026)
von: Aladin, Faiz, et al.
Veröffentlicht: (2026)
VisualMimic: Visual Humanoid Loco-Manipulation via Motion Tracking and Generation
von: Yin, Shaofeng, et al.
Veröffentlicht: (2025)
von: Yin, Shaofeng, et al.
Veröffentlicht: (2025)
Hierarchical World Models as Visual Whole-Body Humanoid Controllers
von: Hansen, Nicklas, et al.
Veröffentlicht: (2024)
von: Hansen, Nicklas, et al.
Veröffentlicht: (2024)
ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning
von: Huang, Chi-Pin, et al.
Veröffentlicht: (2025)
von: Huang, Chi-Pin, et al.
Veröffentlicht: (2025)
Motus: A Unified Latent Action World Model
von: Bi, Hongzhe, et al.
Veröffentlicht: (2025)
von: Bi, Hongzhe, et al.
Veröffentlicht: (2025)
Twisting Lids Off with Two Hands
von: Lin, Toru, et al.
Veröffentlicht: (2024)
von: Lin, Toru, et al.
Veröffentlicht: (2024)
CoIRL-AD: Collaborative-Competitive Imitation-Reinforcement Learning in Latent World Models for Autonomous Driving
von: Zheng, Xiaoji, et al.
Veröffentlicht: (2025)
von: Zheng, Xiaoji, et al.
Veröffentlicht: (2025)
Visual Representation Learning with Stochastic Frame Prediction
von: Jang, Huiwon, et al.
Veröffentlicht: (2024)
von: Jang, Huiwon, et al.
Veröffentlicht: (2024)
Dense-depth map guided deep Lidar-Visual Odometry with Sparse Point Clouds and Images
von: Huang, JunYing, et al.
Veröffentlicht: (2025)
von: Huang, JunYing, et al.
Veröffentlicht: (2025)
Vision-Based Runtime Monitoring under Varying Specifications using Semantic Latent Representations
von: Hoxha, Bardh, et al.
Veröffentlicht: (2026)
von: Hoxha, Bardh, et al.
Veröffentlicht: (2026)
Learning Visual Feature-Based World Models via Residual Latent Action
von: Zhang, Xinyu, et al.
Veröffentlicht: (2026)
von: Zhang, Xinyu, et al.
Veröffentlicht: (2026)
Sensor-Invariant Tactile Representation
von: Gupta, Harsh, et al.
Veröffentlicht: (2025)
von: Gupta, Harsh, et al.
Veröffentlicht: (2025)
ACE: A Cross-Platform Visual-Exoskeletons System for Low-Cost Dexterous Teleoperation
von: Yang, Shiqi, et al.
Veröffentlicht: (2024)
von: Yang, Shiqi, et al.
Veröffentlicht: (2024)
Reconstruction or Semantics? What Makes a Latent Space Useful for Robotic World Models
von: Nilaksh, et al.
Veröffentlicht: (2026)
von: Nilaksh, et al.
Veröffentlicht: (2026)
Latent Action Pretraining from Videos
von: Ye, Seonghyeon, et al.
Veröffentlicht: (2024)
von: Ye, Seonghyeon, et al.
Veröffentlicht: (2024)
COT-FM: Cluster-wise Optimal Transport Flow Matching
von: Chiang, Chiensheng, et al.
Veröffentlicht: (2026)
von: Chiang, Chiensheng, et al.
Veröffentlicht: (2026)
Being-H0.7: A Latent World-Action Model from Egocentric Videos
von: Luo, Hao, et al.
Veröffentlicht: (2026)
von: Luo, Hao, et al.
Veröffentlicht: (2026)
LoRA3D: Low-Rank Self-Calibration of 3D Geometric Foundation Models
von: Lu, Ziqi, et al.
Veröffentlicht: (2024)
von: Lu, Ziqi, et al.
Veröffentlicht: (2024)
Feudal Networks for Visual Navigation
von: Johnson, Faith, et al.
Veröffentlicht: (2024)
von: Johnson, Faith, et al.
Veröffentlicht: (2024)
Minimalist Visual Inertial Odometry
von: Pasti, Francesco, et al.
Veröffentlicht: (2026)
von: Pasti, Francesco, et al.
Veröffentlicht: (2026)
Robot Synesthesia: In-Hand Manipulation with Visuotactile Sensing
von: Yuan, Ying, et al.
Veröffentlicht: (2023)
von: Yuan, Ying, et al.
Veröffentlicht: (2023)
Latent Representations for Visual Proprioception in Inexpensive Robots
von: Sheikholeslami, Sahara, et al.
Veröffentlicht: (2025)
von: Sheikholeslami, Sahara, et al.
Veröffentlicht: (2025)
Safety Certification in the Latent space using Control Barrier Functions and World Models
von: Anand, Mehul, et al.
Veröffentlicht: (2025)
von: Anand, Mehul, et al.
Veröffentlicht: (2025)
Bi-Manual Joint Camera Calibration and Scene Representation
von: Tang, Haozhan, et al.
Veröffentlicht: (2025)
von: Tang, Haozhan, et al.
Veröffentlicht: (2025)
Quantization-Aware Imitation-Learning for Resource-Efficient Robotic Control
von: Park, Seongmin, et al.
Veröffentlicht: (2024)
von: Park, Seongmin, et al.
Veröffentlicht: (2024)
4D Contrastive Superflows are Dense 3D Representation Learners
von: Xu, Xiang, et al.
Veröffentlicht: (2024)
von: Xu, Xiang, et al.
Veröffentlicht: (2024)
Transferable Tactile Transformers for Representation Learning Across Diverse Sensors and Tasks
von: Zhao, Jialiang, et al.
Veröffentlicht: (2024)
von: Zhao, Jialiang, et al.
Veröffentlicht: (2024)
The Temporal Trap: Entanglement in Pre-Trained Visual Representations for Visuomotor Policy Learning
von: Tsagkas, Nikolaos, et al.
Veröffentlicht: (2025)
von: Tsagkas, Nikolaos, et al.
Veröffentlicht: (2025)
Time-Archival Camera Virtualization for Sports and Visual Performances
von: Zhang, Yunxiao, et al.
Veröffentlicht: (2026)
von: Zhang, Yunxiao, et al.
Veröffentlicht: (2026)
OSN: Infinite Representations of Dynamic 3D Scenes from Monocular Videos
von: Song, Ziyang, et al.
Veröffentlicht: (2024)
von: Song, Ziyang, et al.
Veröffentlicht: (2024)
Scaling Proprioceptive-Visual Learning with Heterogeneous Pre-trained Transformers
von: Wang, Lirui, et al.
Veröffentlicht: (2024)
von: Wang, Lirui, et al.
Veröffentlicht: (2024)
A Recipe for Unbounded Data Augmentation in Visual Reinforcement Learning
von: Almuzairee, Abdulaziz, et al.
Veröffentlicht: (2024)
von: Almuzairee, Abdulaziz, et al.
Veröffentlicht: (2024)
Point Cloud Models Improve Visual Robustness in Robotic Learners
von: Peri, Skand, et al.
Veröffentlicht: (2024)
von: Peri, Skand, et al.
Veröffentlicht: (2024)
Squint: Fast Visual Reinforcement Learning for Sim-to-Real Robotics
von: Almuzairee, Abdulaziz, et al.
Veröffentlicht: (2026)
von: Almuzairee, Abdulaziz, et al.
Veröffentlicht: (2026)
Merging and Disentangling Views in Visual Reinforcement Learning for Robotic Manipulation
von: Almuzairee, Abdulaziz, et al.
Veröffentlicht: (2025)
von: Almuzairee, Abdulaziz, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Inference-Time Enhancement of Generative Robot Policies via Predictive World Modeling
von: Qi, Han, et al.
Veröffentlicht: (2025) -
Language-Enhanced Latent Representations for Out-of-Distribution Detection in Autonomous Driving
von: Mao, Zhenjiang, et al.
Veröffentlicht: (2024) -
Latent Object Characteristics Recognition with Visual to Haptic-Audio Cross-modal Transfer Learning
von: Saito, Namiko, et al.
Veröffentlicht: (2024) -
Visual Whole-Body Control for Legged Loco-Manipulation
von: Liu, Minghuan, et al.
Veröffentlicht: (2024) -
Grounding Bodily Awareness in Visual Representations for Efficient Policy Learning
von: Wang, Junlin, et al.
Veröffentlicht: (2025)