Pixel Motion Diffusion is What We Need for Robot Control
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nguyen, E-Ro, Zhang, Yichi, Ranasinghe, Kanchana, Li, Xiang, Ryoo, Michael S. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Pixel Motion as Universal Representation for Robot Control
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2025)
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2025)
Robotic VLA Benefits from Joint Learning with Motion Image Diffusion
von: Fang, Yu, et al.
Veröffentlicht: (2025)
von: Fang, Yu, et al.
Veröffentlicht: (2025)
Understanding Long Videos with Multimodal Language Models
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2024)
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2024)
CoPT: Unsupervised Domain Adaptive Segmentation using Domain-Agnostic Text Embeddings
von: Mata, Cristina, et al.
Veröffentlicht: (2025)
von: Mata, Cristina, et al.
Veröffentlicht: (2025)
Language Repository for Long Video Understanding
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024)
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024)
Crossway Diffusion: Improving Diffusion-based Visuomotor Policy via Self-supervised Learning
von: Li, Xiang, et al.
Veröffentlicht: (2023)
von: Li, Xiang, et al.
Veröffentlicht: (2023)
LLaRA: Supercharging Robot Learning Data for Vision-Language Policy
von: Li, Xiang, et al.
Veröffentlicht: (2024)
von: Li, Xiang, et al.
Veröffentlicht: (2024)
Too Many Frames, Not All Useful: Efficient Strategies for Long-Form Video QA
von: Park, Jongwoo, et al.
Veröffentlicht: (2024)
von: Park, Jongwoo, et al.
Veröffentlicht: (2024)
Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2024)
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2024)
LatentCRF: Continuous CRF for Efficient Latent Diffusion
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2024)
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2024)
Future Optical Flow Prediction Improves Robot Control & Video Generation
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2026)
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2026)
Exploring Conditions for Diffusion models in Robotic Control
von: Shin, Heeseong, et al.
Veröffentlicht: (2025)
von: Shin, Heeseong, et al.
Veröffentlicht: (2025)
MambaGlue: Fast and Robust Local Feature Matching With Mamba
von: Ryoo, Kihwan, et al.
Veröffentlicht: (2025)
von: Ryoo, Kihwan, et al.
Veröffentlicht: (2025)
Goal2Pixel: Grounding Goals to Pixels for Vision-Language Navigation
von: Bao, Muyi, et al.
Veröffentlicht: (2026)
von: Bao, Muyi, et al.
Veröffentlicht: (2026)
DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control
von: Wen, Junjie, et al.
Veröffentlicht: (2025)
von: Wen, Junjie, et al.
Veröffentlicht: (2025)
Pixel-wise Smoothing for Certified Robustness against Camera Motion Perturbations
von: Hu, Hanjiang, et al.
Veröffentlicht: (2023)
von: Hu, Hanjiang, et al.
Veröffentlicht: (2023)
UniAct: Unified Motion Generation and Action Streaming for Humanoid Robots
von: Jiang, Nan, et al.
Veröffentlicht: (2025)
von: Jiang, Nan, et al.
Veröffentlicht: (2025)
From Motion to Behavior: Hierarchical Modeling of Humanoid Generative Behavior Control
von: Zhang, Jusheng, et al.
Veröffentlicht: (2025)
von: Zhang, Jusheng, et al.
Veröffentlicht: (2025)
DiffPixelFormer: Differential Pixel-Aware Transformer for RGB-D Indoor Scene Segmentation
von: Gong, Yan, et al.
Veröffentlicht: (2025)
von: Gong, Yan, et al.
Veröffentlicht: (2025)
What Matters in Building Vision-Language-Action Models for Generalist Robots
von: Li, Xinghang, et al.
Veröffentlicht: (2024)
von: Li, Xinghang, et al.
Veröffentlicht: (2024)
Radar and Event Camera Fusion for Agile Robot Ego-Motion Estimation
von: Lyu, Yang, et al.
Veröffentlicht: (2025)
von: Lyu, Yang, et al.
Veröffentlicht: (2025)
Spatial RoboGrasp: Generalized Robotic Grasping Control Policy
von: Huang, Yiqi, et al.
Veröffentlicht: (2025)
von: Huang, Yiqi, et al.
Veröffentlicht: (2025)
Pixel-level Scene Understanding in One Token: Visual States Need What-is-Where Composition
von: Lee, Seokmin, et al.
Veröffentlicht: (2026)
von: Lee, Seokmin, et al.
Veröffentlicht: (2026)
ARC-Calib: Autonomous Markerless Camera-to-Robot Calibration via Exploratory Robot Motions
von: Chanrungmaneekul, Podshara, et al.
Veröffentlicht: (2025)
von: Chanrungmaneekul, Podshara, et al.
Veröffentlicht: (2025)
PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model
von: Liang, Wenqi, et al.
Veröffentlicht: (2025)
von: Liang, Wenqi, et al.
Veröffentlicht: (2025)
RoboGrasp: A Universal Grasping Policy for Robust Robotic Control
von: Huang, Yiqi, et al.
Veröffentlicht: (2025)
von: Huang, Yiqi, et al.
Veröffentlicht: (2025)
Causal World Modeling for Robot Control
von: Li, Lin, et al.
Veröffentlicht: (2026)
von: Li, Lin, et al.
Veröffentlicht: (2026)
Iterative Closed-Loop Motion Synthesis for Scaling the Capabilities of Humanoid Control
von: Xu, Weisheng, et al.
Veröffentlicht: (2026)
von: Xu, Weisheng, et al.
Veröffentlicht: (2026)
MotionBits: Video Segmentation through Motion-Level Analysis of Rigid Bodies
von: Qian, Howard H., et al.
Veröffentlicht: (2026)
von: Qian, Howard H., et al.
Veröffentlicht: (2026)
Large Video Planner Enables Generalizable Robot Control
von: Chen, Boyuan, et al.
Veröffentlicht: (2025)
von: Chen, Boyuan, et al.
Veröffentlicht: (2025)
StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State Representation
von: Liu, Mingyu, et al.
Veröffentlicht: (2025)
von: Liu, Mingyu, et al.
Veröffentlicht: (2025)
Robotic-CLIP: Fine-tuning CLIP on Action Data for Robotic Applications
von: Nguyen, Nghia, et al.
Veröffentlicht: (2024)
von: Nguyen, Nghia, et al.
Veröffentlicht: (2024)
Kimodo: Scaling Controllable Human Motion Generation
von: Rempe, Davis, et al.
Veröffentlicht: (2026)
von: Rempe, Davis, et al.
Veröffentlicht: (2026)
Need for Speed: Zero-Shot Depth Completion with Single-Step Diffusion
von: Gregorek, Jakub, et al.
Veröffentlicht: (2026)
von: Gregorek, Jakub, et al.
Veröffentlicht: (2026)
What Is The Best 3D Scene Representation for Robotics? From Geometric to Foundation Models
von: Deng, Tianchen, et al.
Veröffentlicht: (2025)
von: Deng, Tianchen, et al.
Veröffentlicht: (2025)
FlowDreamer: A RGB-D World Model with Flow-based Motion Representations for Robot Manipulation
von: Guo, Jun, et al.
Veröffentlicht: (2025)
von: Guo, Jun, et al.
Veröffentlicht: (2025)
From Language to Locomotion: Retargeting-free Humanoid Control via Motion Latent Guidance
von: Li, Zhe, et al.
Veröffentlicht: (2025)
von: Li, Zhe, et al.
Veröffentlicht: (2025)
KAN We Flow? Advancing Robotic Manipulation with 3D Flow Matching via KAN & RWKV
von: Chen, Zhihao, et al.
Veröffentlicht: (2026)
von: Chen, Zhihao, et al.
Veröffentlicht: (2026)
AR-VRM: Imitating Human Motions for Visual Robot Manipulation with Analogical Reasoning
von: Yang, Dejie, et al.
Veröffentlicht: (2025)
von: Yang, Dejie, et al.
Veröffentlicht: (2025)
Phrase-Instance Alignment for Generalized Referring Segmentation
von: Nguyen, E-Ro, et al.
Veröffentlicht: (2024)
von: Nguyen, E-Ro, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Pixel Motion as Universal Representation for Robot Control
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2025) -
Robotic VLA Benefits from Joint Learning with Motion Image Diffusion
von: Fang, Yu, et al.
Veröffentlicht: (2025) -
Understanding Long Videos with Multimodal Language Models
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2024) -
CoPT: Unsupervised Domain Adaptive Segmentation using Domain-Agnostic Text Embeddings
von: Mata, Cristina, et al.
Veröffentlicht: (2025) -
Language Repository for Long Video Understanding
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024)