SCAR: Self-Supervised Continuous Action Representation Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Hongjia, Feng, Fan, Fu, Minghao, Wang, Xinyue, Lu, Haofei, Huang, Biwei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VANP: Learning Where to See for Navigation with Self-Supervised Vision-Action Pre-Training
by: Nazeri, Mohammad, et al.
Published: (2024)
by: Nazeri, Mohammad, et al.
Published: (2024)
SCAR: Satellite Imagery-Based Calibration for Aerial Recordings
by: Hölzemann, Henry, et al.
Published: (2026)
by: Hölzemann, Henry, et al.
Published: (2026)
CARE: Multi-Task Pretraining for Latent Continuous Action Representation in Robot Control
by: Shi, Jiaqi, et al.
Published: (2026)
by: Shi, Jiaqi, et al.
Published: (2026)
DST-Calib: A Dual-Path, Self-Supervised, Target-Free LiDAR-Camera Extrinsic Calibration Network
by: Huang, Zhiwei, et al.
Published: (2026)
by: Huang, Zhiwei, et al.
Published: (2026)
Ag2Manip: Learning Novel Manipulation Skills with Agent-Agnostic Visual and Action Representations
by: Li, Puhao, et al.
Published: (2024)
by: Li, Puhao, et al.
Published: (2024)
Self-Supervised Representation Learning with Joint Embedding Predictive Architecture for Automotive LiDAR Object Detection
by: Zhu, Haoran, et al.
Published: (2025)
by: Zhu, Haoran, et al.
Published: (2025)
Action Draft and Verify: A Self-Verifying Framework for Vision-Language-Action Model
by: Zhao, Chen, et al.
Published: (2026)
by: Zhao, Chen, et al.
Published: (2026)
OCRA: Object-Centric Learning with 3D and Tactile Priors for Human-to-Robot Action Transfer
by: Wang, Kuanning, et al.
Published: (2026)
by: Wang, Kuanning, et al.
Published: (2026)
Self Supervised Deep Learning for Robot Grasping
by: Saqib, Danyal, et al.
Published: (2024)
by: Saqib, Danyal, et al.
Published: (2024)
DynVLA: Learning World Dynamics for Action Reasoning in Autonomous Driving
by: Shang, Shuyao, et al.
Published: (2026)
by: Shang, Shuyao, et al.
Published: (2026)
QueryOcc: Query-based Self-Supervision for 3D Semantic Occupancy
by: Lilja, Adam, et al.
Published: (2025)
by: Lilja, Adam, et al.
Published: (2025)
GesVLA: Gesture-Aware Vision-Language-Action Model Embedded Representations
by: Guo, Wenxuan, et al.
Published: (2026)
by: Guo, Wenxuan, et al.
Published: (2026)
SM$^3$: Self-Supervised Multi-task Modeling with Multi-view 2D Images for Articulated Objects
by: Wang, Haowen, et al.
Published: (2024)
by: Wang, Haowen, et al.
Published: (2024)
CricaVPR: Cross-image Correlation-aware Representation Learning for Visual Place Recognition
by: Lu, Feng, et al.
Published: (2024)
by: Lu, Feng, et al.
Published: (2024)
Object-Centric Action-Enhanced Representations for Robot Visuo-Motor Policy Learning
by: Giannakakis, Nikos, et al.
Published: (2025)
by: Giannakakis, Nikos, et al.
Published: (2025)
Kalib: Easy Hand-Eye Calibration with Reference Point Tracking
by: Tang, Tutian, et al.
Published: (2024)
by: Tang, Tutian, et al.
Published: (2024)
Variational Dynamic for Self-Supervised Exploration in Deep Reinforcement Learning
by: Bai, Chenjia, et al.
Published: (2020)
by: Bai, Chenjia, et al.
Published: (2020)
INoD: Injected Noise Discriminator for Self-Supervised Representation Learning in Agricultural Fields
by: Hindel, Julia, et al.
Published: (2023)
by: Hindel, Julia, et al.
Published: (2023)
Bootstrap Dynamic-Aware 3D Visual Representation for Scalable Robot Learning
by: Liang, Qiwei, et al.
Published: (2025)
by: Liang, Qiwei, et al.
Published: (2025)
From Pixels to Tokens: A Systematic Study of Latent Action Supervision for Vision-Language-Action Models
by: Lin, Yihan, et al.
Published: (2026)
by: Lin, Yihan, et al.
Published: (2026)
Angle Robustness Unmanned Aerial Vehicle Navigation in GNSS-Denied Scenarios
by: Wang, Yuxin, et al.
Published: (2024)
by: Wang, Yuxin, et al.
Published: (2024)
CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos
by: Zhang, Chubin, et al.
Published: (2026)
by: Zhang, Chubin, et al.
Published: (2026)
LaMP: Learning Vision-Language-Action Policies with 3D Scene Flow as Latent Motion Prior
by: Wang, Xinkai, et al.
Published: (2026)
by: Wang, Xinkai, et al.
Published: (2026)
DINOv3-Diffusion Policy: Self-Supervised Large Visual Model for Visuomotor Diffusion Policy Learning
by: Egbe, ThankGod, et al.
Published: (2025)
by: Egbe, ThankGod, et al.
Published: (2025)
MotionHint: Self-Supervised Monocular Visual Odometry with Motion Constraints
by: Wang, Cong, et al.
Published: (2021)
by: Wang, Cong, et al.
Published: (2021)
LARY: A Latent Action Representation Yielding Benchmark for Generalizable Vision-to-Action Alignment
by: Nie, Dujun, et al.
Published: (2026)
by: Nie, Dujun, et al.
Published: (2026)
Investigating Location-Regularised Self-Supervised Feature Learning for Seafloor Visual Imagery
by: Liang, Cailei, et al.
Published: (2025)
by: Liang, Cailei, et al.
Published: (2025)
Continuous Vision-Language-Action Co-Learning with Semantic-Physical Alignment for Behavioral Cloning
by: Qi, Xiuxiu, et al.
Published: (2025)
by: Qi, Xiuxiu, et al.
Published: (2025)
GASP: Unifying Geometric and Semantic Self-Supervised Pre-training for Autonomous Driving
by: Ljungbergh, William, et al.
Published: (2025)
by: Ljungbergh, William, et al.
Published: (2025)
OTTER: A Vision-Language-Action Model with Text-Aware Visual Feature Extraction
by: Huang, Huang, et al.
Published: (2025)
by: Huang, Huang, et al.
Published: (2025)
ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models
by: Ye, Wencheng, et al.
Published: (2025)
by: Ye, Wencheng, et al.
Published: (2025)
Lookahead Exploration with Neural Radiance Representation for Continuous Vision-Language Navigation
by: Wang, Zihan, et al.
Published: (2024)
by: Wang, Zihan, et al.
Published: (2024)
On-Device Self-Supervised Learning of Low-Latency Monocular Depth from Only Events
by: Hagenaars, Jesse, et al.
Published: (2024)
by: Hagenaars, Jesse, et al.
Published: (2024)
MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
by: Shi, Hao, et al.
Published: (2025)
by: Shi, Hao, et al.
Published: (2025)
M2CURL: Sample-Efficient Multimodal Reinforcement Learning via Self-Supervised Representation Learning for Robotic Manipulation
by: Lygerakis, Fotios, et al.
Published: (2024)
by: Lygerakis, Fotios, et al.
Published: (2024)
VAT: Vision Action Transformer by Unlocking Full Representation of ViT
by: Li, Wenhao, et al.
Published: (2025)
by: Li, Wenhao, et al.
Published: (2025)
Demystifying Action Space Design for Robotic Manipulation Policies
by: Feng, Yuchun, et al.
Published: (2026)
by: Feng, Yuchun, et al.
Published: (2026)
SpatialActor: Exploring Disentangled Spatial Representations for Robust Robotic Manipulation
by: Shi, Hao, et al.
Published: (2025)
by: Shi, Hao, et al.
Published: (2025)
NeSLAM: Neural Implicit Mapping and Self-Supervised Feature Tracking With Depth Completion and Denoising
by: Deng, Tianchen, et al.
Published: (2024)
by: Deng, Tianchen, et al.
Published: (2024)
Self-Improving Vision-Language-Action Models with Data Generation via Residual RL
by: Xiao, Wenli, et al.
Published: (2025)
by: Xiao, Wenli, et al.
Published: (2025)
Similar Items
-
VANP: Learning Where to See for Navigation with Self-Supervised Vision-Action Pre-Training
by: Nazeri, Mohammad, et al.
Published: (2024) -
SCAR: Satellite Imagery-Based Calibration for Aerial Recordings
by: Hölzemann, Henry, et al.
Published: (2026) -
CARE: Multi-Task Pretraining for Latent Continuous Action Representation in Robot Control
by: Shi, Jiaqi, et al.
Published: (2026) -
DST-Calib: A Dual-Path, Self-Supervised, Target-Free LiDAR-Camera Extrinsic Calibration Network
by: Huang, Zhiwei, et al.
Published: (2026) -
Ag2Manip: Learning Novel Manipulation Skills with Agent-Agnostic Visual and Action Representations
by: Li, Puhao, et al.
Published: (2024)