SFHand: Learning Embodied Manipulation by Streaming Egocentric 3D Hand Forecasting
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Ruicong, Huang, Yifei, Ouyang, Liangyang, Kang, Caixin, Sato, Yoichi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SocialDirector: Training-Free Social Interaction Control for Multi-Person Video Generation
by: Ouyang, Liangyang, et al.
Published: (2026)
by: Ouyang, Liangyang, et al.
Published: (2026)
Can MLLMs Read the Room? A Multimodal Benchmark for Assessing Deception in Multi-Party Social Interactions
by: Kang, Caixin, et al.
Published: (2025)
by: Kang, Caixin, et al.
Published: (2025)
Egocentric Action-aware Inertial Localization in Point Clouds with Vision-Language Guidance
by: Zhang, Mingfang, et al.
Published: (2025)
by: Zhang, Mingfang, et al.
Published: (2025)
ActionVOS: Actions as Prompts for Video Object Segmentation
by: Ouyang, Liangyang, et al.
Published: (2024)
by: Ouyang, Liangyang, et al.
Published: (2024)
Can MLLMs Read the Room? A Multimodal Benchmark for Verifying Truthfulness in Multi-Party Social Interactions
by: Kang, Caixin, et al.
Published: (2025)
by: Kang, Caixin, et al.
Published: (2025)
Multi-speaker Attention Alignment for Multimodal Social Interaction
by: Ouyang, Liangyang, et al.
Published: (2025)
by: Ouyang, Liangyang, et al.
Published: (2025)
Single-to-Dual-View Adaptation for Egocentric 3D Hand Pose Estimation
by: Liu, Ruicong, et al.
Published: (2024)
by: Liu, Ruicong, et al.
Published: (2024)
Masked Video and Body-worn IMU Autoencoder for Egocentric Action Recognition
by: Zhang, Mingfang, et al.
Published: (2024)
by: Zhang, Mingfang, et al.
Published: (2024)
Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality?
by: Kang, Caixin, et al.
Published: (2026)
by: Kang, Caixin, et al.
Published: (2026)
AssemblyHands-X: Modeling 3D Hand-Body Coordination for Understanding Bimanual Human Activities
by: Banno, Tatsuro, et al.
Published: (2025)
by: Banno, Tatsuro, et al.
Published: (2025)
Pre-Training for 3D Hand Pose Estimation with Contrastive Learning on Large-Scale Hand Images in the Wild
by: Lin, Nie, et al.
Published: (2024)
by: Lin, Nie, et al.
Published: (2024)
The N-Body Problem: Parallel Execution from Single-Person Egocentric Video
by: Zhu, Zhifan, et al.
Published: (2025)
by: Zhu, Zhifan, et al.
Published: (2025)
Egocentric Gaze Estimation via Neck-Mounted Camera
by: Huang, Haoyu, et al.
Published: (2026)
by: Huang, Haoyu, et al.
Published: (2026)
Leadership Assessment in Pediatric Intensive Care Unit Team Training
by: Ouyang, Liangyang, et al.
Published: (2025)
by: Ouyang, Liangyang, et al.
Published: (2025)
Leveraging RGB Images for Pre-Training of Event-Based Hand Pose Estimation
by: Liu, Ruicong, et al.
Published: (2025)
by: Liu, Ruicong, et al.
Published: (2025)
SiMHand: Mining Similar Hands for Large-Scale 3D Hand Pose Pre-training
by: Lin, Nie, et al.
Published: (2025)
by: Lin, Nie, et al.
Published: (2025)
LORE: Latent Optimization for Precise Semantic Control in Rectified Flow-based Image Editing
by: Ouyang, Liangyang, et al.
Published: (2025)
by: Ouyang, Liangyang, et al.
Published: (2025)
EgoHandICL: Egocentric 3D Hand Reconstruction with In-Context Learning
by: Xie, Binzhu, et al.
Published: (2026)
by: Xie, Binzhu, et al.
Published: (2026)
Uni-Hand: Universal Hand Motion Forecasting in Egocentric Views
by: Ma, Junyi, et al.
Published: (2025)
by: Ma, Junyi, et al.
Published: (2025)
Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision
by: He, Yuping, et al.
Published: (2025)
by: He, Yuping, et al.
Published: (2025)
EggHand: A Multimodal Foundation Model for Egocentric Hand Pose Forecasting
by: Choi, Jaeyoung, et al.
Published: (2026)
by: Choi, Jaeyoung, et al.
Published: (2026)
Affordance-Guided Diffusion Prior for 3D Hand Reconstruction
by: Suzuki, Naru, et al.
Published: (2025)
by: Suzuki, Naru, et al.
Published: (2025)
CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering
by: Zhang, Mingfang, et al.
Published: (2026)
by: Zhang, Mingfang, et al.
Published: (2026)
Modeling Fine-Grained Hand-Object Dynamics for Egocentric Video Representation Learning
by: Pei, Baoqi, et al.
Published: (2025)
by: Pei, Baoqi, et al.
Published: (2025)
EventEgoHands: Event-based Egocentric 3D Hand Mesh Reconstruction
by: Hara, Ryosei, et al.
Published: (2025)
by: Hara, Ryosei, et al.
Published: (2025)
Benchmarks and Challenges in Pose Estimation for Egocentric Hand Interactions with Objects
by: Fan, Zicong, et al.
Published: (2024)
by: Fan, Zicong, et al.
Published: (2024)
EMAG: Ego-motion Aware and Generalizable 2D Hand Forecasting from Egocentric Videos
by: Hatano, Masashi, et al.
Published: (2024)
by: Hatano, Masashi, et al.
Published: (2024)
Hand2World: Autoregressive Egocentric Interaction Generation via Free-Space Hand Gestures
by: Wang, Yuxi, et al.
Published: (2026)
by: Wang, Yuxi, et al.
Published: (2026)
EgoInstruct: An Egocentric Video Dataset of Face-to-face Instructional Interactions with Multi-modal LLM Benchmarking
by: Sakai, Yuki, et al.
Published: (2025)
by: Sakai, Yuki, et al.
Published: (2025)
Introducing HOT3D: An Egocentric Dataset for 3D Hand and Object Tracking
by: Banerjee, Prithviraj, et al.
Published: (2024)
by: Banerjee, Prithviraj, et al.
Published: (2024)
Towards Egocentric 3D Hand Pose Estimation in Unseen Domains
by: Mucha, Wiktor, et al.
Published: (2026)
by: Mucha, Wiktor, et al.
Published: (2026)
EgoSpot:Egocentric Multimodal Control for Hands-Free Mobile Manipulation
by: Zhang, Ganlin, et al.
Published: (2023)
by: Zhang, Ganlin, et al.
Published: (2023)
StreamEQA: Towards Streaming Video Understanding for Embodied Scenarios
by: Wang, Yifei, et al.
Published: (2025)
by: Wang, Yifei, et al.
Published: (2025)
Gazing Into Missteps: Leveraging Eye-Gaze for Unsupervised Mistake Detection in Egocentric Videos of Skilled Human Activities
by: Mazzamuto, Michele, et al.
Published: (2024)
by: Mazzamuto, Michele, et al.
Published: (2024)
Benchmarking 2D Egocentric Hand Pose Datasets
by: Taran, Olga, et al.
Published: (2024)
by: Taran, Olga, et al.
Published: (2024)
In My Perspective, In My Hands: Accurate Egocentric 2D Hand Pose and Action Recognition
by: Mucha, Wiktor, et al.
Published: (2024)
by: Mucha, Wiktor, et al.
Published: (2024)
Gaze Beyond the Frame: Forecasting Egocentric 3D Visual Span
by: Yun, Heeseung, et al.
Published: (2025)
by: Yun, Heeseung, et al.
Published: (2025)
MEgoHand: Multimodal Egocentric Hand-Object Interaction Motion Generation
by: Zhou, Bohan, et al.
Published: (2025)
by: Zhou, Bohan, et al.
Published: (2025)
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model
by: Huang, Yifei, et al.
Published: (2024)
by: Huang, Yifei, et al.
Published: (2024)
Controllable Egocentric Video Generation via Occlusion-Aware Sparse 3D Hand Joints
by: Zhang, Chenyangguang, et al.
Published: (2026)
by: Zhang, Chenyangguang, et al.
Published: (2026)
Similar Items
-
SocialDirector: Training-Free Social Interaction Control for Multi-Person Video Generation
by: Ouyang, Liangyang, et al.
Published: (2026) -
Can MLLMs Read the Room? A Multimodal Benchmark for Assessing Deception in Multi-Party Social Interactions
by: Kang, Caixin, et al.
Published: (2025) -
Egocentric Action-aware Inertial Localization in Point Clouds with Vision-Language Guidance
by: Zhang, Mingfang, et al.
Published: (2025) -
ActionVOS: Actions as Prompts for Video Object Segmentation
by: Ouyang, Liangyang, et al.
Published: (2024) -
Can MLLMs Read the Room? A Multimodal Benchmark for Verifying Truthfulness in Multi-Party Social Interactions
by: Kang, Caixin, et al.
Published: (2025)