An End-to-End Framework for Video Multi-Person Pose Estimation
Fuente:
arXiv
Saved in:
| Main Author: | Wei, Zhihong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
End-to-End Multi-Person Pose Estimation with Pose-Aware Video Transformer
by: Yu, Yonghui, et al.
Published: (2025)
by: Yu, Yonghui, et al.
Published: (2025)
From Category to Scenery: An End-to-End Framework for Multi-Person Human-Object Interaction Recognition in Videos
by: Qiao, Tanqiu, et al.
Published: (2024)
by: Qiao, Tanqiu, et al.
Published: (2024)
LPSNet: End-to-End Human Pose and Shape Estimation with Lensless Imaging
by: Ge, Haoyang, et al.
Published: (2024)
by: Ge, Haoyang, et al.
Published: (2024)
KGpose: Keypoint-Graph Driven End-to-End Multi-Object 6D Pose Estimation via Point-Wise Pose Voting
by: Jeong, Andrew
Published: (2024)
by: Jeong, Andrew
Published: (2024)
SEMPose: A Single End-to-end Network for Multi-object Pose Estimation
by: Liu, Xin, et al.
Published: (2024)
by: Liu, Xin, et al.
Published: (2024)
End-to-End Probabilistic Geometry-Guided Regression for 6DoF Object Pose Estimation
by: Pöllabauer, Thomas, et al.
Published: (2024)
by: Pöllabauer, Thomas, et al.
Published: (2024)
LMVC: An End-to-End Learned Multiview Video Coding Framework
by: Sheng, Xihua, et al.
Published: (2025)
by: Sheng, Xihua, et al.
Published: (2025)
VividAnimator: An End-to-End Audio and Pose-driven Half-Body Human Animation Framework
by: Huang, Donglin, et al.
Published: (2025)
by: Huang, Donglin, et al.
Published: (2025)
Towards Fully Decoupled End-to-End Person Search
by: Zhang, Pengcheng, et al.
Published: (2023)
by: Zhang, Pengcheng, et al.
Published: (2023)
Leveraging Image Matching Toward End-to-End Relative Camera Pose Regression
by: Khatib, Fadi, et al.
Published: (2022)
by: Khatib, Fadi, et al.
Published: (2022)
WorldJen: An End-to-End Multi-Dimensional Benchmark for Generative Video Models
by: Inbasekar, Karthik, et al.
Published: (2026)
by: Inbasekar, Karthik, et al.
Published: (2026)
SyncAnimation: A Real-Time End-to-End Framework for Audio-Driven Human Pose and Talking Head Animation
by: Liu, Yujian, et al.
Published: (2025)
by: Liu, Yujian, et al.
Published: (2025)
InterMesh: Explicit Interaction-Aware End-to-End Multi-Person Human Mesh Recovery
by: Zheng, Kaili, et al.
Published: (2026)
by: Zheng, Kaili, et al.
Published: (2026)
End-to-End Training for Autoregressive Video Diffusion via Self-Resampling
by: Guo, Yuwei, et al.
Published: (2025)
by: Guo, Yuwei, et al.
Published: (2025)
Tracking by Detection and Query: An Efficient End-to-End Framework for Multi-Object Tracking
by: Jia, Shukun, et al.
Published: (2024)
by: Jia, Shukun, et al.
Published: (2024)
End-to-End Facial Expression Detection in Long Videos
by: Fang, Yini, et al.
Published: (2025)
by: Fang, Yini, et al.
Published: (2025)
STORM: End-to-End Referring Multi-Object Tracking in Videos
by: Lu, Zijia, et al.
Published: (2026)
by: Lu, Zijia, et al.
Published: (2026)
SDT-6D: Fully Sparse Depth-Transformer for Staged End-to-End 6D Pose Estimation in Industrial Multi-View Bin Picking
by: Leuze, Nico, et al.
Published: (2025)
by: Leuze, Nico, et al.
Published: (2025)
Pose-Based Sign Language Spotting via an End-to-End Encoder Architecture
by: Johnny, Samuel Ebimobowei, et al.
Published: (2025)
by: Johnny, Samuel Ebimobowei, et al.
Published: (2025)
GaussianMotion: End-to-End Learning of Animatable Gaussian Avatars with Pose Guidance from Text
by: Shim, Gyumin, et al.
Published: (2025)
by: Shim, Gyumin, et al.
Published: (2025)
SMTrack: End-to-End Trained Spiking Neural Networks for Multi-Object Tracking in RGB Videos
by: Zhong, Pengzhi, et al.
Published: (2025)
by: Zhong, Pengzhi, et al.
Published: (2025)
End-to-End Dense Video Grounding via Parallel Regression
by: Shi, Fengyuan, et al.
Published: (2021)
by: Shi, Fengyuan, et al.
Published: (2021)
Beyond Imperfections: A Conditional Inpainting Approach for End-to-End Artifact Removal in VTON and Pose Transfer
by: Tabatabaei, Aref, et al.
Published: (2024)
by: Tabatabaei, Aref, et al.
Published: (2024)
OneVision: An End-to-End Generative Framework for Multi-view E-commerce Vision Search
by: Zheng, Zexin, et al.
Published: (2025)
by: Zheng, Zexin, et al.
Published: (2025)
TimeLoc: A Unified End-to-End Framework for Precise Timestamp Localization in Long Videos
by: Zhang, Chen-Lin, et al.
Published: (2025)
by: Zhang, Chen-Lin, et al.
Published: (2025)
Weaver: End-to-End Agentic System Training for Video Interleaved Reasoning
by: Shi, Yudi, et al.
Published: (2026)
by: Shi, Yudi, et al.
Published: (2026)
End-to-End Streaming Video Temporal Action Segmentation with Reinforce Learning
by: Zhang, Jinrong, et al.
Published: (2023)
by: Zhang, Jinrong, et al.
Published: (2023)
Multi-Person Pose Estimation Evaluation Using Optimal Transportation and Improved Pose Matching
by: Moriki, Takato, et al.
Published: (2026)
by: Moriki, Takato, et al.
Published: (2026)
Drive&Gen: Co-Evaluating End-to-End Driving and Video Generation Models
by: Wang, Jiahao, et al.
Published: (2025)
by: Wang, Jiahao, et al.
Published: (2025)
EVE: Towards End-to-End Video Subtitle Extraction with Vision-Language Models
by: Yu, Haiyang, et al.
Published: (2025)
by: Yu, Haiyang, et al.
Published: (2025)
MoGA: Mixture-of-Groups Attention for End-to-End Long Video Generation
by: Jia, Weinan, et al.
Published: (2025)
by: Jia, Weinan, et al.
Published: (2025)
Unleashing Generalization of End-to-End Autonomous Driving with Controllable Long Video Generation
by: Ma, Enhui, et al.
Published: (2024)
by: Ma, Enhui, et al.
Published: (2024)
MoCha:End-to-End Video Character Replacement without Structural Guidance
by: Xu, Zhengbo, et al.
Published: (2026)
by: Xu, Zhengbo, et al.
Published: (2026)
VSD-MOT: End-to-End Multi-Object Tracking in Low-Quality Video Scenes Guided by Visual Semantic Distillation
by: Du, Jun
Published: (2026)
by: Du, Jun
Published: (2026)
End-to-End Shared Attention Estimation via Group Detection with Feedback Refinement
by: Nakatani, Chihiro, et al.
Published: (2026)
by: Nakatani, Chihiro, et al.
Published: (2026)
Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation
by: Chen, Yuheng, et al.
Published: (2026)
by: Chen, Yuheng, et al.
Published: (2026)
REMM:Rotation-Equivariant Framework for End-to-End Multimodal Image Matching
by: Nie, Han, et al.
Published: (2024)
by: Nie, Han, et al.
Published: (2024)
MASSM: An End-to-End Deep Learning Framework for Multi-Anatomy Statistical Shape Modeling Directly From Images
by: Ukey, Janmesh, et al.
Published: (2024)
by: Ukey, Janmesh, et al.
Published: (2024)
End to End Face Reconstruction via Differentiable PnP
by: Lu, Yiren, et al.
Published: (2024)
by: Lu, Yiren, et al.
Published: (2024)
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs
by: Cheng, Dabing, et al.
Published: (2025)
by: Cheng, Dabing, et al.
Published: (2025)
Similar Items
-
End-to-End Multi-Person Pose Estimation with Pose-Aware Video Transformer
by: Yu, Yonghui, et al.
Published: (2025) -
From Category to Scenery: An End-to-End Framework for Multi-Person Human-Object Interaction Recognition in Videos
by: Qiao, Tanqiu, et al.
Published: (2024) -
LPSNet: End-to-End Human Pose and Shape Estimation with Lensless Imaging
by: Ge, Haoyang, et al.
Published: (2024) -
KGpose: Keypoint-Graph Driven End-to-End Multi-Object 6D Pose Estimation via Point-Wise Pose Voting
by: Jeong, Andrew
Published: (2024) -
SEMPose: A Single End-to-end Network for Multi-object Pose Estimation
by: Liu, Xin, et al.
Published: (2024)