MVSA-Net: Multi-View State-Action Recognition for Robust and Deployable Trajectory Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Asali, Ehsan, Doshi, Prashant, Sun, Jin |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Visual IRL for Human-Like Robotic Manipulation
by: Asali, Ehsan, et al.
Published: (2024)
by: Asali, Ehsan, et al.
Published: (2024)
Grounding Foundational Vision Models with 3D Human Poses for Robust Action Recognition
by: Babey, Nicholas, et al.
Published: (2025)
by: Babey, Nicholas, et al.
Published: (2025)
CLAMP: Contrastive Learning for 3D Multi-View Action-Conditioned Robotic Manipulation Pretraining
by: Liu, I-Chun Arthur, et al.
Published: (2026)
by: Liu, I-Chun Arthur, et al.
Published: (2026)
Lightweight Multimodal Artificial Intelligence Framework for Maritime Multi-Scene Recognition
by: Xi, Xinyu, et al.
Published: (2025)
by: Xi, Xinyu, et al.
Published: (2025)
PEAfowl: Perception-Enhanced Multi-View Vision-Language-Action for Bimanual Manipulation
by: Fan, Qingyu, et al.
Published: (2026)
by: Fan, Qingyu, et al.
Published: (2026)
SkillMimicGen: Automated Demonstration Generation for Efficient Skill Learning and Deployment
by: Garrett, Caelan, et al.
Published: (2024)
by: Garrett, Caelan, et al.
Published: (2024)
FALCON: Future-Aware Learning with Contextual Object-Centric Pretraining for UAV Action Recognition
by: Xian, Ruiqi, et al.
Published: (2024)
by: Xian, Ruiqi, et al.
Published: (2024)
SkelVIT: Consensus of Vision Transformers for a Lightweight Skeleton-Based Action Recognition System
by: Karadag, Ozge Oztimur
Published: (2023)
by: Karadag, Ozge Oztimur
Published: (2023)
MATRIX: Multi-Agent Trajectory Generation with Diverse Contexts
by: Xu, Zhuo, et al.
Published: (2024)
by: Xu, Zhuo, et al.
Published: (2024)
On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations
by: Guo, Jianing, et al.
Published: (2025)
by: Guo, Jianing, et al.
Published: (2025)
SplaTraj: Camera Trajectory Generation with Semantic Gaussian Splatting
by: Liu, Xinyi, et al.
Published: (2024)
by: Liu, Xinyi, et al.
Published: (2024)
Interactive Spatiotemporal Token Attention Network for Skeleton-based General Interactive Action Recognition
by: Wen, Yuhang, et al.
Published: (2023)
by: Wen, Yuhang, et al.
Published: (2023)
MS-Net: A Multi-Path Sparse Model for Motion Prediction in Multi-Scenes
by: Tang, Xiaqiang, et al.
Published: (2024)
by: Tang, Xiaqiang, et al.
Published: (2024)
Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
by: Kim, Moo Jin, et al.
Published: (2025)
by: Kim, Moo Jin, et al.
Published: (2025)
Visual Sync: Multi-Camera Synchronization via Cross-View Object Motion
by: Liu, Shaowei, et al.
Published: (2025)
by: Liu, Shaowei, et al.
Published: (2025)
VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting
by: Lin, Juyi, et al.
Published: (2025)
by: Lin, Juyi, et al.
Published: (2025)
Generative Camera Dolly: Extreme Monocular Dynamic Novel View Synthesis
by: Van Hoorick, Basile, et al.
Published: (2024)
by: Van Hoorick, Basile, et al.
Published: (2024)
Enhancing Generalization in Vision-Language-Action Models by Preserving Pretrained Representations
by: Grover, Shresth, et al.
Published: (2025)
by: Grover, Shresth, et al.
Published: (2025)
ROSA: Harnessing Robot States for Vision-Language and Action Alignment
by: Wen, Yuqing, et al.
Published: (2025)
by: Wen, Yuqing, et al.
Published: (2025)
Recognizing Actions from Robotic View for Natural Human-Robot Interaction
by: Wang, Ziyi, et al.
Published: (2025)
by: Wang, Ziyi, et al.
Published: (2025)
MARs: Multi-view Attention Regularizations for Patch-based Feature Recognition of Space Terrain
by: Chase Jr, Timothy, et al.
Published: (2024)
by: Chase Jr, Timothy, et al.
Published: (2024)
View-Invariant Policy Learning via Zero-Shot Novel View Synthesis
by: Tian, Stephen, et al.
Published: (2024)
by: Tian, Stephen, et al.
Published: (2024)
M4Diffuser: Multi-View Diffusion Policy with Manipulability-Aware Control for Robust Mobile Manipulation
by: Dong, Ju, et al.
Published: (2025)
by: Dong, Ju, et al.
Published: (2025)
RoboCurate: Harnessing Diversity with Action-Verified Neural Trajectory for Robot Learning
by: Kim, Seungku, et al.
Published: (2026)
by: Kim, Seungku, et al.
Published: (2026)
UAV-VLA: Vision-Language-Action System for Large Scale Aerial Mission Generation
by: Sautenkov, Oleg, et al.
Published: (2025)
by: Sautenkov, Oleg, et al.
Published: (2025)
A Spatio-temporal Graph Network Allowing Incomplete Trajectory Input for Pedestrian Trajectory Prediction
by: Long, Juncen, et al.
Published: (2025)
by: Long, Juncen, et al.
Published: (2025)
Lessons from Deploying CropFollow++: Under-Canopy Agricultural Navigation with Keypoints
by: Sivakumar, Arun N., et al.
Published: (2024)
by: Sivakumar, Arun N., et al.
Published: (2024)
Work Zones challenge VLM Trajectory Planning: Toward Mitigation and Robust Autonomous Driving
by: Liao, Yifan, et al.
Published: (2025)
by: Liao, Yifan, et al.
Published: (2025)
Predicate Hierarchies Improve Few-Shot State Classification
by: Jin, Emily, et al.
Published: (2025)
by: Jin, Emily, et al.
Published: (2025)
Conditional Unscented Autoencoders for Trajectory Prediction
by: Janjoš, Faris, et al.
Published: (2023)
by: Janjoš, Faris, et al.
Published: (2023)
Enhancing 3D Point Cloud Classification with ModelNet-R and Point-SkipNet
by: Saeid, Mohammad, et al.
Published: (2025)
by: Saeid, Mohammad, et al.
Published: (2025)
From Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation Priors
by: Zhang, Zhengshen, et al.
Published: (2025)
by: Zhang, Zhengshen, et al.
Published: (2025)
Hypergraph-based Multi-View Action Recognition using Event Cameras
by: Gao, Yue, et al.
Published: (2024)
by: Gao, Yue, et al.
Published: (2024)
GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving
by: Russell, Lloyd, et al.
Published: (2025)
by: Russell, Lloyd, et al.
Published: (2025)
Extrapolated Urban View Synthesis Benchmark
by: Han, Xiangyu, et al.
Published: (2024)
by: Han, Xiangyu, et al.
Published: (2024)
Distilling Knowledge for Short-to-Long Term Trajectory Prediction
by: Das, Sourav, et al.
Published: (2023)
by: Das, Sourav, et al.
Published: (2023)
RoboVIP: Multi-View Video Generation with Visual Identity Prompting Augments Robot Manipulation
by: Wang, Boyang, et al.
Published: (2026)
by: Wang, Boyang, et al.
Published: (2026)
CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models
by: Zhao, Qingqing, et al.
Published: (2025)
by: Zhao, Qingqing, et al.
Published: (2025)
Learning to Visually Connect Actions and their Effects
by: Parmar, Paritosh, et al.
Published: (2024)
by: Parmar, Paritosh, et al.
Published: (2024)
Generative Image as Action Models
by: Shridhar, Mohit, et al.
Published: (2024)
by: Shridhar, Mohit, et al.
Published: (2024)
Similar Items
-
Visual IRL for Human-Like Robotic Manipulation
by: Asali, Ehsan, et al.
Published: (2024) -
Grounding Foundational Vision Models with 3D Human Poses for Robust Action Recognition
by: Babey, Nicholas, et al.
Published: (2025) -
CLAMP: Contrastive Learning for 3D Multi-View Action-Conditioned Robotic Manipulation Pretraining
by: Liu, I-Chun Arthur, et al.
Published: (2026) -
Lightweight Multimodal Artificial Intelligence Framework for Maritime Multi-Scene Recognition
by: Xi, Xinyu, et al.
Published: (2025) -
PEAfowl: Perception-Enhanced Multi-View Vision-Language-Action for Bimanual Manipulation
by: Fan, Qingyu, et al.
Published: (2026)