CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Qixiu, Liang, Yaobo, Wang, Zeyu, Luo, Lin, Chen, Xi, Liao, Mozheng, Wei, Fangyun, Deng, Yu, Xu, Sicheng, Zhang, Yizhong, Wang, Xiaofan, Liu, Bei, Fu, Jianlong, Bao, Jianmin, Chen, Dong, Shi, Yuanchun, Yang, Jiaolong, Guo, Baining |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
by: Li, Qixiu, et al.
Published: (2025)
by: Li, Qixiu, et al.
Published: (2025)
VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
by: Shen, Yichao, et al.
Published: (2025)
by: Shen, Yichao, et al.
Published: (2025)
VASA-3D: Lifelike Audio-Driven Gaussian Head Avatars from a Single Image
by: Xu, Sicheng, et al.
Published: (2025)
by: Xu, Sicheng, et al.
Published: (2025)
Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs
by: Xu, Sicheng, et al.
Published: (2026)
by: Xu, Sicheng, et al.
Published: (2026)
MobileManiBench: Simplifying Model Verification for Mobile Manipulation
by: Wang, Wenbo, et al.
Published: (2026)
by: Wang, Wenbo, et al.
Published: (2026)
TaskGround: Structured Executable Task Inference for Full-Scene Household Reasoning
by: Feng, ZhiYuan, et al.
Published: (2026)
by: Feng, ZhiYuan, et al.
Published: (2026)
VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time
by: Xu, Sicheng, et al.
Published: (2024)
by: Xu, Sicheng, et al.
Published: (2024)
Gaussian Variation Field Diffusion for High-fidelity Video-to-4D Synthesis
by: Zhang, Bowen, et al.
Published: (2025)
by: Zhang, Bowen, et al.
Published: (2025)
CogDPM: Diffusion Probabilistic Models via Cognitive Predictive Coding
by: Chen, Kaiyuan, et al.
Published: (2024)
by: Chen, Kaiyuan, et al.
Published: (2024)
LoLA: Long Horizon Latent Action Learning for General Robot Manipulation
by: Wang, Xiaofan, et al.
Published: (2025)
by: Wang, Xiaofan, et al.
Published: (2025)
Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes
by: Feng, Zhiyuan, et al.
Published: (2025)
by: Feng, Zhiyuan, et al.
Published: (2025)
PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction
by: Chen, Shizhe, et al.
Published: (2026)
by: Chen, Shizhe, et al.
Published: (2026)
InterACT: Inter-dependency Aware Action Chunking with Hierarchical Attention Transformers for Bimanual Manipulation
by: Lee, Andrew, et al.
Published: (2024)
by: Lee, Andrew, et al.
Published: (2024)
UniGraspTransformer: Simplified Policy Distillation for Scalable Dexterous Robotic Grasping
by: Wang, Wenbo, et al.
Published: (2024)
by: Wang, Wenbo, et al.
Published: (2024)
CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification
by: Li, Wei, et al.
Published: (2025)
by: Li, Wei, et al.
Published: (2025)
ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning
by: Yang, Yandan, et al.
Published: (2026)
by: Yang, Yandan, et al.
Published: (2026)
A Simple Baseline for Spoken Language to Sign Language Translation with 3D Avatars
by: Zuo, Ronglai, et al.
Published: (2024)
by: Zuo, Ronglai, et al.
Published: (2024)
Fast Autoregressive Models for Continuous Latent Generation
by: Hang, Tiankai, et al.
Published: (2025)
by: Hang, Tiankai, et al.
Published: (2025)
Language-Guided Face Animation by Recurrent StyleGAN-based Generator
by: Hang, Tiankai, et al.
Published: (2022)
by: Hang, Tiankai, et al.
Published: (2022)
Beyond Dense Futures: World Models as Structured Planners for Robotic Manipulation
by: Jin, Minghao, et al.
Published: (2026)
by: Jin, Minghao, et al.
Published: (2026)
HiSpatial: Taming Hierarchical 3D Spatial Understanding in Vision-Language Models
by: Liang, Huizhi, et al.
Published: (2026)
by: Liang, Huizhi, et al.
Published: (2026)
Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks
by: Zhang, Wenqi, et al.
Published: (2025)
by: Zhang, Wenqi, et al.
Published: (2025)
Diffusion Models without Classifier-free Guidance
by: Tang, Zhicong, et al.
Published: (2025)
by: Tang, Zhicong, et al.
Published: (2025)
LatBot: Distilling Universal Latent Actions for Vision-Language-Action Models
by: Li, Zuolei, et al.
Published: (2025)
by: Li, Zuolei, et al.
Published: (2025)
Physical Autoregressive Model for Robotic Manipulation without Action Pretraining
by: Song, Zijian, et al.
Published: (2025)
by: Song, Zijian, et al.
Published: (2025)
EgoBrain: Synergizing Minds and Eyes For Human Action Understanding
by: Lin, Nie, et al.
Published: (2025)
by: Lin, Nie, et al.
Published: (2025)
Aggregated Contextual Transformations for High-Resolution Image Inpainting
by: Zeng, Yanhong, et al.
Published: (2021)
by: Zeng, Yanhong, et al.
Published: (2021)
BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation
by: Wang, Hongyu, et al.
Published: (2025)
by: Wang, Hongyu, et al.
Published: (2025)
WHEN TO ACT, WHEN TO WAIT: Modeling the Intent-Action Alignment Problem in Dialogue
by: Qian, Yaoyao, et al.
Published: (2025)
by: Qian, Yaoyao, et al.
Published: (2025)
Dynamic Execution Commitment of Vision-Language-Action Models
by: Chen, Feng, et al.
Published: (2026)
by: Chen, Feng, et al.
Published: (2026)
GaussianCube: A Structured and Explicit Radiance Representation for 3D Generative Modeling
by: Zhang, Bowen, et al.
Published: (2024)
by: Zhang, Bowen, et al.
Published: (2024)
OS-ATLAS: A Foundation Action Model for Generalist GUI Agents
by: Wu, Zhiyong, et al.
Published: (2024)
by: Wu, Zhiyong, et al.
Published: (2024)
Low-Resolution Action Recognition for Tiny Actions Challenge
by: Chen, Boyu, et al.
Published: (2022)
by: Chen, Boyu, et al.
Published: (2022)
CogLM: Tracking Cognitive Development of Large Language Models
by: Wang, Xinglin, et al.
Published: (2024)
by: Wang, Xinglin, et al.
Published: (2024)
MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
by: Shi, Hao, et al.
Published: (2025)
by: Shi, Hao, et al.
Published: (2025)
CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning
by: Qi, Ji, et al.
Published: (2024)
by: Qi, Ji, et al.
Published: (2024)
Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models
by: Chen, Dong, et al.
Published: (2026)
by: Chen, Dong, et al.
Published: (2026)
LOME: Learning Human-Object Manipulation with Action-Conditioned Egocentric World Model
by: Gao, Quankai, et al.
Published: (2026)
by: Gao, Quankai, et al.
Published: (2026)
A Corrected Open Boundary Framework for Lattice Boltzmann Immiscible Pseudopotential Models
by: Chen, Yizhong, et al.
Published: (2025)
by: Chen, Yizhong, et al.
Published: (2025)
ExACT: Language-guided Conceptual Reasoning and Uncertainty Estimation for Event-based Action Recognition and More
by: Zhou, Jiazhou, et al.
Published: (2024)
by: Zhou, Jiazhou, et al.
Published: (2024)
Similar Items
-
Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
by: Li, Qixiu, et al.
Published: (2025) -
VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
by: Shen, Yichao, et al.
Published: (2025) -
VASA-3D: Lifelike Audio-Driven Gaussian Head Avatars from a Single Image
by: Xu, Sicheng, et al.
Published: (2025) -
Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs
by: Xu, Sicheng, et al.
Published: (2026) -
MobileManiBench: Simplifying Model Verification for Mobile Manipulation
by: Wang, Wenbo, et al.
Published: (2026)