MamKPD: A Simple Mamba Baseline for Real-Time 2D Keypoint Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Dang, Yonghao, Liu, Liyuan, Kang, Hui, Ye, Ping, Yin, Jianqin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Generically Contrastive Spatiotemporal Representation Enhancement for 3D Skeleton Action Recognition
by: Zhang, Shaojie, et al.
Published: (2023)
by: Zhang, Shaojie, et al.
Published: (2023)
SiT-MLP: A Simple MLP with Point-wise Topology Feature Learning for Skeleton-based Action Recognition
by: Zhang, Shaojie, et al.
Published: (2023)
by: Zhang, Shaojie, et al.
Published: (2023)
DHRNet: A Dual-Path Hierarchical Relation Network for Multi-Person Pose Estimation
by: Dang, Yonghao, et al.
Published: (2024)
by: Dang, Yonghao, et al.
Published: (2024)
MaskSem: Semantic-Guided Masking for Learning 3D Hybrid High-Order Motion Representation
by: Wei, Wei, et al.
Published: (2025)
by: Wei, Wei, et al.
Published: (2025)
Leveraging the Video-level Semantic Consistency of Event for Audio-visual Event Localization
by: Jiang, Yuanyuan, et al.
Published: (2022)
by: Jiang, Yuanyuan, et al.
Published: (2022)
ActivityCLIP: Enhancing Group Activity Recognition by Mining Complementary Information from Text to Supplement Image Modality
by: Xu, Guoliang, et al.
Published: (2024)
by: Xu, Guoliang, et al.
Published: (2024)
Physics-constrained Attack against Convolution-based Human Motion Prediction
by: Duan, Chengxu, et al.
Published: (2023)
by: Duan, Chengxu, et al.
Published: (2023)
Kinematics Modeling Network for Video-based Human Pose Estimation
by: Dang, Yonghao, et al.
Published: (2022)
by: Dang, Yonghao, et al.
Published: (2022)
ESG-Net: Event-Aware Semantic Guided Network for Dense Audio-Visual Event Localization
by: Li, Huilai, et al.
Published: (2025)
by: Li, Huilai, et al.
Published: (2025)
L2HCount:Generalizing Crowd Counting from Low to High Crowd Density via Density Simulation
by: Xu, Guoliang, et al.
Published: (2025)
by: Xu, Guoliang, et al.
Published: (2025)
ResDiT: Evoking the Intrinsic Resolution Scalability in Diffusion Transformers
by: Ma, Yiyang, et al.
Published: (2025)
by: Ma, Yiyang, et al.
Published: (2025)
3DGAA: Realistic and Robust 3D Gaussian-based Adversarial Attack for Autonomous Driving
by: Zhang, Yixun, et al.
Published: (2025)
by: Zhang, Yixun, et al.
Published: (2025)
Mamba YOLO: A Simple Baseline for Object Detection with State Space Model
by: Wang, Zeyu, et al.
Published: (2024)
by: Wang, Zeyu, et al.
Published: (2024)
EAR: Enhancing Uni-Modal Representations for Weakly Supervised Audio-Visual Video Parsing
by: Li, Huilai, et al.
Published: (2026)
by: Li, Huilai, et al.
Published: (2026)
StyMam: A Mamba-Based Generator for Artistic Style Transfer
by: Hong, Zhou, et al.
Published: (2026)
by: Hong, Zhou, et al.
Published: (2026)
Towards Physically Realizable Adversarial Attacks in Embodied Vision Navigation
by: Chen, Meng, et al.
Published: (2024)
by: Chen, Meng, et al.
Published: (2024)
MamFusion: Multi-Mamba with Temporal Fusion for Partially Relevant Video Retrieval
by: Ying, Xinru, et al.
Published: (2025)
by: Ying, Xinru, et al.
Published: (2025)
BiTAA: A Bi-Task Adversarial Attack for Object Detection and Depth Estimation via 3D Gaussian Splatting
by: Zhang, Yixun, et al.
Published: (2025)
by: Zhang, Yixun, et al.
Published: (2025)
MambaIR: A Simple Baseline for Image Restoration with State-Space Model
by: Guo, Hang, et al.
Published: (2024)
by: Guo, Hang, et al.
Published: (2024)
MambaTrack: A Simple Baseline for Multiple Object Tracking with State Space Model
by: Xiao, Changcheng, et al.
Published: (2024)
by: Xiao, Changcheng, et al.
Published: (2024)
Mam-App: A Novel Parameter-Efficient Mamba Model for Apple Leaf Disease Classification
by: Mahamood, Md Nadim, et al.
Published: (2026)
by: Mahamood, Md Nadim, et al.
Published: (2026)
MedVLThinker: Simple Baselines for Multimodal Medical Reasoning
by: Huang, Xiaoke, et al.
Published: (2025)
by: Huang, Xiaoke, et al.
Published: (2025)
RT-DETRv2: Improved Baseline with Bag-of-Freebies for Real-Time Detection Transformer
by: Lv, Wenyu, et al.
Published: (2024)
by: Lv, Wenyu, et al.
Published: (2024)
Revisiting Simple Baselines for In-The-Wild Deepfake Detection
by: Castaneda, Orlando, et al.
Published: (2025)
by: Castaneda, Orlando, et al.
Published: (2025)
Towards more realistic human motion prediction with attention to motion coordination
by: Ding, Pengxiang, et al.
Published: (2024)
by: Ding, Pengxiang, et al.
Published: (2024)
CLIP-Powered TASS: Target-Aware Single-Stream Network for Audio-Visual Question Answering
by: Jiang, Yuanyuan, et al.
Published: (2024)
by: Jiang, Yuanyuan, et al.
Published: (2024)
A Two-stream Hybrid CNN-Transformer Network for Skeleton-based Human Interaction Recognition
by: Yin, Ruoqi, et al.
Published: (2023)
by: Yin, Ruoqi, et al.
Published: (2023)
SaMam: Style-aware State Space Model for Arbitrary Image Style Transfer
by: Liu, Hongda, et al.
Published: (2025)
by: Liu, Hongda, et al.
Published: (2025)
ER-Pose: Rethinking Keypoint-Driven Representation Learning for Real-Time Human Pose Estimation
by: Li, Nanjun, et al.
Published: (2026)
by: Li, Nanjun, et al.
Published: (2026)
CT3D++: Improving 3D Object Detection with Keypoint-induced Channel-wise Transformer
by: Sheng, Hualian, et al.
Published: (2024)
by: Sheng, Hualian, et al.
Published: (2024)
GLane3D : Detecting Lanes with Graph of 3D Keypoints
by: Öztürk, Halil İbrahim, et al.
Published: (2025)
by: Öztürk, Halil İbrahim, et al.
Published: (2025)
Open-Vocabulary Animal Keypoint Detection with Semantic-feature Matching
by: Zhang, Hao, et al.
Published: (2023)
by: Zhang, Hao, et al.
Published: (2023)
Geometry-Aware 3D Salient Object Detection Network
by: Wang, Chen, et al.
Published: (2025)
by: Wang, Chen, et al.
Published: (2025)
DaD: Distilled Reinforcement Learning for Diverse Keypoint Detection
by: Edstedt, Johan, et al.
Published: (2025)
by: Edstedt, Johan, et al.
Published: (2025)
MiM-ISTD: Mamba-in-Mamba for Efficient Infrared Small Target Detection
by: Chen, Tianxiang, et al.
Published: (2024)
by: Chen, Tianxiang, et al.
Published: (2024)
From Pairs to Sequences: Track-Aware Policy Gradients for Keypoint Detection
by: Liu, Yepeng, et al.
Published: (2026)
by: Liu, Yepeng, et al.
Published: (2026)
YOLOPoint Joint Keypoint and Object Detection
by: Backhaus, Anton, et al.
Published: (2024)
by: Backhaus, Anton, et al.
Published: (2024)
X-Pose: Detecting Any Keypoints
by: Yang, Jie, et al.
Published: (2023)
by: Yang, Jie, et al.
Published: (2023)
ConvMambaNet: A Hybrid CNN-Mamba State Space Architecture for Accurate and Real-Time EEG Seizure Detection
by: Khan, Md. Nishan, et al.
Published: (2026)
by: Khan, Md. Nishan, et al.
Published: (2026)
FIND: A Simple yet Effective Baseline for Diffusion-Generated Image Detection
by: Li, Jie, et al.
Published: (2026)
by: Li, Jie, et al.
Published: (2026)
Similar Items
-
A Generically Contrastive Spatiotemporal Representation Enhancement for 3D Skeleton Action Recognition
by: Zhang, Shaojie, et al.
Published: (2023) -
SiT-MLP: A Simple MLP with Point-wise Topology Feature Learning for Skeleton-based Action Recognition
by: Zhang, Shaojie, et al.
Published: (2023) -
DHRNet: A Dual-Path Hierarchical Relation Network for Multi-Person Pose Estimation
by: Dang, Yonghao, et al.
Published: (2024) -
MaskSem: Semantic-Guided Masking for Learning 3D Hybrid High-Order Motion Representation
by: Wei, Wei, et al.
Published: (2025) -
Leveraging the Video-level Semantic Consistency of Event for Audio-visual Event Localization
by: Jiang, Yuanyuan, et al.
Published: (2022)