Exploiting Spatial-Temporal Context for Interacting Hand Reconstruction on Monocular RGB Video
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Weichao, Hu, Hezhen, Zhou, Wengang, li, Li, Li, Houqiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Self-Supervised Representation Learning with Spatial-Temporal Consistency for Sign Language Recognition
von: Zhao, Weichao, et al.
Veröffentlicht: (2024)
von: Zhao, Weichao, et al.
Veröffentlicht: (2024)
Scaling up Multimodal Pre-training for Sign Language Understanding
von: Zhou, Wengang, et al.
Veröffentlicht: (2024)
von: Zhou, Wengang, et al.
Veröffentlicht: (2024)
Uni-Sign: Toward Unified Sign Language Understanding at Scale
von: Li, Zecheng, et al.
Veröffentlicht: (2025)
von: Li, Zecheng, et al.
Veröffentlicht: (2025)
MASA: Motion-aware Masked Autoencoder with Semantic Alignment for Sign Language Recognition
von: Zhao, Weichao, et al.
Veröffentlicht: (2024)
von: Zhao, Weichao, et al.
Veröffentlicht: (2024)
Cross-Modal Consistency Learning for Sign Language Recognition
von: Wu, Kepeng, et al.
Veröffentlicht: (2025)
von: Wu, Kepeng, et al.
Veröffentlicht: (2025)
LaneTCA: Enhancing Video Lane Detection with Temporal Context Aggregation
von: Zhou, Keyi, et al.
Veröffentlicht: (2024)
von: Zhou, Keyi, et al.
Veröffentlicht: (2024)
Expressive Gaussian Human Avatars from Monocular RGB Video
von: Hu, Hezhen, et al.
Veröffentlicht: (2024)
von: Hu, Hezhen, et al.
Veröffentlicht: (2024)
Video-based Sign Language Recognition without Temporal Segmentation
von: Huang, Jie, et al.
Veröffentlicht: (2018)
von: Huang, Jie, et al.
Veröffentlicht: (2018)
Exploiting GPT-4 Vision for Zero-shot Point Cloud Understanding
von: Sun, Qi, et al.
Veröffentlicht: (2024)
von: Sun, Qi, et al.
Veröffentlicht: (2024)
Mitigating Hallucination in VideoLLMs via Temporal-Aware Activation Engineering
von: Cai, Jianfeng, et al.
Veröffentlicht: (2025)
von: Cai, Jianfeng, et al.
Veröffentlicht: (2025)
Motion-aware 3D Gaussian Splatting for Efficient Dynamic Scene Reconstruction
von: Guo, Zhiyang, et al.
Veröffentlicht: (2024)
von: Guo, Zhiyang, et al.
Veröffentlicht: (2024)
DocR1: Evidence Page-Guided GRPO for Multi-Page Document Understanding
von: Xiong, Junyu, et al.
Veröffentlicht: (2025)
von: Xiong, Junyu, et al.
Veröffentlicht: (2025)
3D Hand Mesh Recovery from Monocular RGB in Camera Space
von: Li, Haonan, et al.
Veröffentlicht: (2024)
von: Li, Haonan, et al.
Veröffentlicht: (2024)
Disambiguating Monocular Reconstruction of 3D Clothed Human with Spatial-Temporal Transformer
von: Deng, Yong, et al.
Veröffentlicht: (2024)
von: Deng, Yong, et al.
Veröffentlicht: (2024)
DeepEraser: Deep Iterative Context Mining for Generic Text Eraser
von: Feng, Hao, et al.
Veröffentlicht: (2024)
von: Feng, Hao, et al.
Veröffentlicht: (2024)
GaussNav: Gaussian Splatting for Visual Navigation
von: Lei, Xiaohan, et al.
Veröffentlicht: (2024)
von: Lei, Xiaohan, et al.
Veröffentlicht: (2024)
MotionRL: Align Text-to-Motion Generation to Human Preferences with Multi-Reward Reinforcement Learning
von: Liu, Xiaoyang, et al.
Veröffentlicht: (2024)
von: Liu, Xiaoyang, et al.
Veröffentlicht: (2024)
Revisiting Shadow Detection from a Vision-Language Perspective
von: Wang, Yonghui, et al.
Veröffentlicht: (2026)
von: Wang, Yonghui, et al.
Veröffentlicht: (2026)
AdaptVision: Dynamic Input Scaling in MLLMs for Versatile Scene Understanding
von: Wang, Yonghui, et al.
Veröffentlicht: (2024)
von: Wang, Yonghui, et al.
Veröffentlicht: (2024)
StepVAR: Structure-Texture Guided Pruning for Visual Autoregressive Models
von: Liu, Keli, et al.
Veröffentlicht: (2026)
von: Liu, Keli, et al.
Veröffentlicht: (2026)
Forest2Seq: Revitalizing Order Prior for Sequential Indoor Scene Synthesis
von: Sun, Qi, et al.
Veröffentlicht: (2024)
von: Sun, Qi, et al.
Veröffentlicht: (2024)
StreetSurfGS: Scalable Urban Street Surface Reconstruction with Planar-based Gaussian Splatting
von: Cui, Xiao, et al.
Veröffentlicht: (2024)
von: Cui, Xiao, et al.
Veröffentlicht: (2024)
Robust Multimodal Large Language Models Against Modality Conflict
von: Zhang, Zongmeng, et al.
Veröffentlicht: (2025)
von: Zhang, Zongmeng, et al.
Veröffentlicht: (2025)
SwinShadow: Shifted Window for Ambiguous Adjacent Shadow Detection
von: Wang, Yonghui, et al.
Veröffentlicht: (2024)
von: Wang, Yonghui, et al.
Veröffentlicht: (2024)
Spatial Decomposition and Temporal Fusion based Inter Prediction for Learned Video Compression
von: Sheng, Xihua, et al.
Veröffentlicht: (2024)
von: Sheng, Xihua, et al.
Veröffentlicht: (2024)
Optimizing Distributional Geometry Alignment with Optimal Transport for Generative Dataset Distillation
von: Cui, Xiao, et al.
Veröffentlicht: (2025)
von: Cui, Xiao, et al.
Veröffentlicht: (2025)
Grasp in Gaussians: Fast Monocular Reconstruction of Dynamic Hand-Object Interactions
von: Aytekin, Ayce Idil, et al.
Veröffentlicht: (2026)
von: Aytekin, Ayce Idil, et al.
Veröffentlicht: (2026)
Instance-aware Exploration-Verification-Exploitation for Instance ImageGoal Navigation
von: Lei, Xiaohan, et al.
Veröffentlicht: (2024)
von: Lei, Xiaohan, et al.
Veröffentlicht: (2024)
GeoHand: Unlocking Prior Geometry Knowledge for Monocular 3D Hand Reconstruction
von: Lin, Weiquan, et al.
Veröffentlicht: (2026)
von: Lin, Weiquan, et al.
Veröffentlicht: (2026)
GHOST: Fast Category-agnostic Hand-Object Interaction Reconstruction from RGB Videos using Gaussian Splatting
von: Aboukhadra, Ahmed Tawfik, et al.
Veröffentlicht: (2026)
von: Aboukhadra, Ahmed Tawfik, et al.
Veröffentlicht: (2026)
Progressive Multi-modal Conditional Prompt Tuning
von: Qiu, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Qiu, Xiaoyu, et al.
Veröffentlicht: (2024)
RoFIR: Robust Fisheye Image Rectification Framework Impervious to Optical Center Deviation
von: Liao, Zhaokang, et al.
Veröffentlicht: (2024)
von: Liao, Zhaokang, et al.
Veröffentlicht: (2024)
Image2Sentence based Asymmetrical Zero-shot Composed Image Retrieval
von: Du, Yongchao, et al.
Veröffentlicht: (2024)
von: Du, Yongchao, et al.
Veröffentlicht: (2024)
Learning Generalizable Human Motion Generator with Reinforcement Learning
von: Mao, Yunyao, et al.
Veröffentlicht: (2024)
von: Mao, Yunyao, et al.
Veröffentlicht: (2024)
ArtHOI: Taming Foundation Models for Monocular 4D Reconstruction of Hand-Articulated-Object Interactions
von: Wang, Zikai, et al.
Veröffentlicht: (2026)
von: Wang, Zikai, et al.
Veröffentlicht: (2026)
Text-Animator: Controllable Visual Text Video Generation
von: Liu, Lin, et al.
Veröffentlicht: (2024)
von: Liu, Lin, et al.
Veröffentlicht: (2024)
MorpheuS: Neural Dynamic 360° Surface Reconstruction from Monocular RGB-D Video
von: Wang, Hengyi, et al.
Veröffentlicht: (2023)
von: Wang, Hengyi, et al.
Veröffentlicht: (2023)
Exploiting Multimodal Spatial-temporal Patterns for Video Object Tracking
von: Hu, Xiantao, et al.
Veröffentlicht: (2024)
von: Hu, Xiantao, et al.
Veröffentlicht: (2024)
Refined Geometry-guided Head Avatar Reconstruction from Monocular RGB Video
von: Park, Pilseo, et al.
Veröffentlicht: (2025)
von: Park, Pilseo, et al.
Veröffentlicht: (2025)
Rethinking Long-tailed Dataset Distillation: A Uni-Level Framework with Unbiased Recovery and Relabeling
von: Cui, Xiao, et al.
Veröffentlicht: (2025)
von: Cui, Xiao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Self-Supervised Representation Learning with Spatial-Temporal Consistency for Sign Language Recognition
von: Zhao, Weichao, et al.
Veröffentlicht: (2024) -
Scaling up Multimodal Pre-training for Sign Language Understanding
von: Zhou, Wengang, et al.
Veröffentlicht: (2024) -
Uni-Sign: Toward Unified Sign Language Understanding at Scale
von: Li, Zecheng, et al.
Veröffentlicht: (2025) -
MASA: Motion-aware Masked Autoencoder with Semantic Alignment for Sign Language Recognition
von: Zhao, Weichao, et al.
Veröffentlicht: (2024) -
Cross-Modal Consistency Learning for Sign Language Recognition
von: Wu, Kepeng, et al.
Veröffentlicht: (2025)