Mash, Spread, Slice! Learning to Manipulate Object States via Visual Spatial Progress
Fuente:
arXiv
Saved in:
| Main Authors: | Mandikal, Priyanka, Hu, Jiaheng, Dass, Shivin, Majumder, Sagnik, Martín-Martín, Roberto, Grauman, Kristen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SPOC: Spatially-Progressing Object State Change Segmentation in Video
by: Mandikal, Priyanka, et al.
Published: (2025)
by: Mandikal, Priyanka, et al.
Published: (2025)
ActiveRIR: Active Audio-Visual Exploration for Acoustic Environment Modeling
by: Somayazulu, Arjun, et al.
Published: (2024)
by: Somayazulu, Arjun, et al.
Published: (2024)
Learning to Look: Seeking Information for Decision Making via Policy Factorization
by: Dass, Shivin, et al.
Published: (2024)
by: Dass, Shivin, et al.
Published: (2024)
Learning Spatial Features from Audio-Visual Correspondence in Egocentric Videos
by: Majumder, Sagnik, et al.
Published: (2023)
by: Majumder, Sagnik, et al.
Published: (2023)
Audio-Visual Camera Pose Estimation with Passive Scene Sounds and In-the-Wild Video
by: Adebi, Daniel, et al.
Published: (2025)
by: Adebi, Daniel, et al.
Published: (2025)
Model-Based Runtime Monitoring with Interactive Imitation Learning
by: Liu, Huihan, et al.
Published: (2023)
by: Liu, Huihan, et al.
Published: (2023)
COLLAGE: Adaptive Fusion-based Retrieval for Augmented Policy Learning
by: Kumar, Sateesh, et al.
Published: (2025)
by: Kumar, Sateesh, et al.
Published: (2025)
TeleMoMa: A Modular and Versatile Teleoperation System for Mobile Manipulation
by: Dass, Shivin, et al.
Published: (2024)
by: Dass, Shivin, et al.
Published: (2024)
ScrewMimic: Bimanual Imitation from Human Videos with Screw Space Projection
by: Bahety, Arpit, et al.
Published: (2024)
by: Bahety, Arpit, et al.
Published: (2024)
DataMIL: Selecting Data for Robot Imitation Learning with Datamodels
by: Dass, Shivin, et al.
Published: (2025)
by: Dass, Shivin, et al.
Published: (2025)
MistExit: Learning to Exit for Early Mistake Detection in Procedural Videos
by: Majumder, Sagnik, et al.
Published: (2026)
by: Majumder, Sagnik, et al.
Published: (2026)
Switch-a-View: View Selection Learned from Unlabeled In-the-wild Videos
by: Majumder, Sagnik, et al.
Published: (2024)
by: Majumder, Sagnik, et al.
Published: (2024)
Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric Perspective
by: Chung, Nhat, et al.
Published: (2025)
by: Chung, Nhat, et al.
Published: (2025)
Materialistic RIR: Material Conditioned Realistic RIR Generation
by: Saad, Mahnoor Fatima, et al.
Published: (2026)
by: Saad, Mahnoor Fatima, et al.
Published: (2026)
Learning Object State Changes in Videos: An Open-World Perspective
by: Xue, Zihui, et al.
Published: (2023)
by: Xue, Zihui, et al.
Published: (2023)
Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos
by: Majumder, Sagnik, et al.
Published: (2024)
by: Majumder, Sagnik, et al.
Published: (2024)
EgoExo-WM: Unlocking Exo Video for Ego World Models
by: Tran, Danny, et al.
Published: (2026)
by: Tran, Danny, et al.
Published: (2026)
RoboSSM: Scalable In-context Imitation Learning via State-Space Models
by: Yoo, Youngju, et al.
Published: (2025)
by: Yoo, Youngju, et al.
Published: (2025)
FLaRe: Achieving Masterful and Adaptive Robot Policies with Large-Scale Reinforcement Learning Fine-Tuning
by: Hu, Jiaheng, et al.
Published: (2024)
by: Hu, Jiaheng, et al.
Published: (2024)
Diagnose, Correct, and Learn from Manipulation Failures via Visual Symbols
by: Zeng, Xianchao, et al.
Published: (2025)
by: Zeng, Xianchao, et al.
Published: (2025)
6D Object Pose Tracking in Internet Videos for Robotic Manipulation
by: Ponimatkin, Georgy, et al.
Published: (2025)
by: Ponimatkin, Georgy, et al.
Published: (2025)
Disentangled Unsupervised Skill Discovery for Efficient Hierarchical Reinforcement Learning
by: Hu, Jiaheng, et al.
Published: (2024)
by: Hu, Jiaheng, et al.
Published: (2024)
HieraMamba: Video Temporal Grounding via Hierarchical Anchor-Mamba Pooling
by: An, Joungbin, et al.
Published: (2025)
by: An, Joungbin, et al.
Published: (2025)
Rapid Adaptation of Particle Dynamics for Generalized Deformable Object Mobile Manipulation
by: Wu, Bohan, et al.
Published: (2026)
by: Wu, Bohan, et al.
Published: (2026)
SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and Object Manipulation
by: Qi, Zekun, et al.
Published: (2025)
by: Qi, Zekun, et al.
Published: (2025)
FUNCanon: Learning Pose-Aware Action Primitives via Functional Object Canonicalization for Generalizable Robotic Manipulation
by: Xu, Hongli, et al.
Published: (2025)
by: Xu, Hongli, et al.
Published: (2025)
MG-Nav: Dual-Scale Visual Navigation via Sparse Spatial Memory
by: Wang, Bo, et al.
Published: (2025)
by: Wang, Bo, et al.
Published: (2025)
SLAC: Simulation-Pretrained Latent Action Space for Whole-Body Real-World RL
by: Hu, Jiaheng, et al.
Published: (2025)
by: Hu, Jiaheng, et al.
Published: (2025)
Simultaneous Tactile-Visual Perception for Learning Multimodal Robot Manipulation
by: Li, Yuyang, et al.
Published: (2025)
by: Li, Yuyang, et al.
Published: (2025)
DoughNet: A Visual Predictive Model for Topological Manipulation of Deformable Objects
by: Bauer, Dominik, et al.
Published: (2024)
by: Bauer, Dominik, et al.
Published: (2024)
Spatially Visual Perception for End-to-End Robotic Learning
by: Davies, Travis, et al.
Published: (2024)
by: Davies, Travis, et al.
Published: (2024)
Residual-NeRF: Learning Residual NeRFs for Transparent Object Manipulation
by: Duisterhof, Bardienus P., et al.
Published: (2024)
by: Duisterhof, Bardienus P., et al.
Published: (2024)
VTAO-BiManip: Masked Visual-Tactile-Action Pre-training with Object Understanding for Bimanual Dexterous Manipulation
by: Sun, Zhengnan, et al.
Published: (2025)
by: Sun, Zhengnan, et al.
Published: (2025)
SpatialActor: Exploring Disentangled Spatial Representations for Robust Robotic Manipulation
by: Shi, Hao, et al.
Published: (2025)
by: Shi, Hao, et al.
Published: (2025)
Learning Skill-Attributes for Transferable Assessment in Video
by: Ashutosh, Kumar, et al.
Published: (2025)
by: Ashutosh, Kumar, et al.
Published: (2025)
Collision-Aware Object-Goal Visual Navigation via Two-Stage Deep Reinforcement Learning
by: Wang, Hongwu, et al.
Published: (2025)
by: Wang, Hongwu, et al.
Published: (2025)
An Integrated Approach to Robotic Object Grasping and Manipulation
by: Ahmed, Owais, et al.
Published: (2024)
by: Ahmed, Owais, et al.
Published: (2024)
Object-Centric Instruction Augmentation for Robotic Manipulation
by: Wen, Junjie, et al.
Published: (2024)
by: Wen, Junjie, et al.
Published: (2024)
Progress-Aware Video Frame Captioning
by: Xue, Zihui, et al.
Published: (2024)
by: Xue, Zihui, et al.
Published: (2024)
Two by Two: Learning Multi-Task Pairwise Objects Assembly for Generalizable Robot Manipulation
by: Qi, Yu, et al.
Published: (2025)
by: Qi, Yu, et al.
Published: (2025)
Similar Items
-
SPOC: Spatially-Progressing Object State Change Segmentation in Video
by: Mandikal, Priyanka, et al.
Published: (2025) -
ActiveRIR: Active Audio-Visual Exploration for Acoustic Environment Modeling
by: Somayazulu, Arjun, et al.
Published: (2024) -
Learning to Look: Seeking Information for Decision Making via Policy Factorization
by: Dass, Shivin, et al.
Published: (2024) -
Learning Spatial Features from Audio-Visual Correspondence in Egocentric Videos
by: Majumder, Sagnik, et al.
Published: (2023) -
Audio-Visual Camera Pose Estimation with Passive Scene Sounds and In-the-Wild Video
by: Adebi, Daniel, et al.
Published: (2025)