Physical Autoregressive Model for Robotic Manipulation without Action Pretraining
Fuente:
arXiv
Saved in:
| Main Authors: | Song, Zijian, Qin, Sihan, Chen, Tianshui, Lin, Liang, Wang, Guangrun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SIRI-Bench: Challenging VLMs' Spatial Intelligence through Complex Reasoning Tasks
by: Song, Zijian, et al.
Published: (2025)
by: Song, Zijian, et al.
Published: (2025)
Learning Spatial-Temporal Coherent Correlations for Speech-Preserving Facial Expression Manipulation
by: Chen, Tianshui, et al.
Published: (2026)
by: Chen, Tianshui, et al.
Published: (2026)
Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation
by: Song, Zijian, et al.
Published: (2026)
by: Song, Zijian, et al.
Published: (2026)
Chain-of-Action: Trajectory Autoregressive Modeling for Robotic Manipulation
by: Zhang, Wenbo, et al.
Published: (2025)
by: Zhang, Wenbo, et al.
Published: (2025)
One Model for All: Unified Try-On and Try-Off in Any Pose via LLM-Inspired Bidirectional Tweedie Diffusion
by: Liu, Jinxi, et al.
Published: (2025)
by: Liu, Jinxi, et al.
Published: (2025)
In-Situ Tweedie Discrete Diffusion Models
by: Li, Xiao, et al.
Published: (2025)
by: Li, Xiao, et al.
Published: (2025)
Human-Centric Open-Future Task Discovery: Formulation, Benchmark, and Scalable Tree-Based Search
by: Song, Zijian, et al.
Published: (2025)
by: Song, Zijian, et al.
Published: (2025)
ActionSink: Toward Precise Robot Manipulation with Dynamic Integration of Action Flow
by: Guo, Shanshan, et al.
Published: (2025)
by: Guo, Shanshan, et al.
Published: (2025)
WildVidFit: Video Virtual Try-On in the Wild via Image-Based Controlled Diffusion Models
by: He, Zijian, et al.
Published: (2024)
by: He, Zijian, et al.
Published: (2024)
VTON 360: High-Fidelity Virtual Try-On from Any Viewing Direction
by: He, Zijian, et al.
Published: (2025)
by: He, Zijian, et al.
Published: (2025)
GS: Generative Segmentation via Label Diffusion
by: Chen, Yuhao, et al.
Published: (2025)
by: Chen, Yuhao, et al.
Published: (2025)
UML-CoT: Structured Reasoning and Planning with Unified Modeling Language for Robotic Room Cleaning
by: Chen, Hongyu, et al.
Published: (2025)
by: Chen, Hongyu, et al.
Published: (2025)
ACD: Direct Conditional Control for Video Diffusion Models via Attention Supervision
by: Li, Weiqi, et al.
Published: (2025)
by: Li, Weiqi, et al.
Published: (2025)
HumanGenesis: Agent-Based Geometric and Generative Modeling for Synthetic Human Dynamics
by: Li, Weiqi, et al.
Published: (2025)
by: Li, Weiqi, et al.
Published: (2025)
Neural Scene Designer: Self-Styled Semantic Image Manipulation
by: Lin, Jianman, et al.
Published: (2025)
by: Lin, Jianman, et al.
Published: (2025)
AnimateZoo: Zero-shot Video Generation of Cross-Species Animation via Subject Alignment
by: Xu, Yuanfeng, et al.
Published: (2024)
by: Xu, Yuanfeng, et al.
Published: (2024)
Contrastive Decoupled Representation Learning and Regularization for Speech-Preserving Facial Expression Manipulation
by: Chen, Tianshui, et al.
Published: (2025)
by: Chen, Tianshui, et al.
Published: (2025)
Personalized Cross-Modal Emotional Correlation Learning for Speech-Preserving Facial Expression Manipulation
by: Chen, Tianshui, et al.
Published: (2026)
by: Chen, Tianshui, et al.
Published: (2026)
Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
by: Li, Qixiu, et al.
Published: (2025)
by: Li, Qixiu, et al.
Published: (2025)
Autoregressive Pretraining with Mamba in Vision
by: Ren, Sucheng, et al.
Published: (2024)
by: Ren, Sucheng, et al.
Published: (2024)
TAG: Target-Agnostic Guidance for Stable Object-Centric Inference in Vision-Language-Action Models
by: Zhou, Jiaying, et al.
Published: (2026)
by: Zhou, Jiaying, et al.
Published: (2026)
BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation
by: Wang, Hongyu, et al.
Published: (2025)
by: Wang, Hongyu, et al.
Published: (2025)
VideoMAP: Toward Scalable Mamba-based Video Autoregressive Pretraining
by: Liu, Yunze, et al.
Published: (2025)
by: Liu, Yunze, et al.
Published: (2025)
Geometry aware 3D generation from in-the-wild images in ImageNet
by: Shen, Qijia, et al.
Published: (2024)
by: Shen, Qijia, et al.
Published: (2024)
DFM-VLA: Iterative Action Refinement for Robot Manipulation via Discrete Flow Matching
by: Chen, Jiayi, et al.
Published: (2026)
by: Chen, Jiayi, et al.
Published: (2026)
Exploring Talking Head Models With Adjacent Frame Prior for Speech-Preserving Facial Expression Manipulation
by: Lu, Zhenxuan, et al.
Published: (2026)
by: Lu, Zhenxuan, et al.
Published: (2026)
Dynamic Correlation Learning and Regularization for Multi-Label Confidence Calibration
by: Chen, Tianshui, et al.
Published: (2024)
by: Chen, Tianshui, et al.
Published: (2024)
Gaze-Regularized Vision-Language-Action Models for Robotic Manipulation
by: Pani, Anupam, et al.
Published: (2026)
by: Pani, Anupam, et al.
Published: (2026)
Language-free Compositional Action Generation via Decoupling Refinement
by: Liu, Xiao, et al.
Published: (2023)
by: Liu, Xiao, et al.
Published: (2023)
Making Large Language Models Better Planners with Reasoning-Decision Alignment
by: Huang, Zhijian, et al.
Published: (2024)
by: Huang, Zhijian, et al.
Published: (2024)
Pretraining Frame Preservation for Lightweight Autoregressive Video History Embedding
by: Zhang, Lvmin, et al.
Published: (2025)
by: Zhang, Lvmin, et al.
Published: (2025)
RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation
by: Liu, Jiaming, et al.
Published: (2024)
by: Liu, Jiaming, et al.
Published: (2024)
SQLNet: Scale-Modulated Query and Localization Network for Few-Shot Class-Agnostic Counting
by: Wu, Hefeng, et al.
Published: (2023)
by: Wu, Hefeng, et al.
Published: (2023)
Observe Then Act: Asynchronous Active Vision-Action Model for Robotic Manipulation
by: Wang, Guokang, et al.
Published: (2024)
by: Wang, Guokang, et al.
Published: (2024)
Lift3D Foundation Policy: Lifting 2D Large-Scale Pretrained Models for Robust 3D Robotic Manipulation
by: Jia, Yueru, et al.
Published: (2024)
by: Jia, Yueru, et al.
Published: (2024)
MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
by: Shi, Hao, et al.
Published: (2025)
by: Shi, Hao, et al.
Published: (2025)
ARVideo: Autoregressive Pretraining for Self-Supervised Video Representation Learning
by: Ren, Sucheng, et al.
Published: (2024)
by: Ren, Sucheng, et al.
Published: (2024)
Demystifying Action Space Design for Robotic Manipulation Policies
by: Feng, Yuchun, et al.
Published: (2026)
by: Feng, Yuchun, et al.
Published: (2026)
FlexVAR: Flexible Visual Autoregressive Modeling without Residual Prediction
by: Jiao, Siyu, et al.
Published: (2025)
by: Jiao, Siyu, et al.
Published: (2025)
Exploiting Temporal Audio-Visual Correlation Embedding for Audio-Driven One-Shot Talking Head Animation
by: Xu, Zhihua, et al.
Published: (2025)
by: Xu, Zhihua, et al.
Published: (2025)
Similar Items
-
SIRI-Bench: Challenging VLMs' Spatial Intelligence through Complex Reasoning Tasks
by: Song, Zijian, et al.
Published: (2025) -
Learning Spatial-Temporal Coherent Correlations for Speech-Preserving Facial Expression Manipulation
by: Chen, Tianshui, et al.
Published: (2026) -
Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation
by: Song, Zijian, et al.
Published: (2026) -
Chain-of-Action: Trajectory Autoregressive Modeling for Robotic Manipulation
by: Zhang, Wenbo, et al.
Published: (2025) -
One Model for All: Unified Try-On and Try-Off in Any Pose via LLM-Inspired Bidirectional Tweedie Diffusion
by: Liu, Jinxi, et al.
Published: (2025)