MamKPD: A Simple Mamba Baseline for Real-Time 2D Keypoint Detection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dang, Yonghao, Liu, Liyuan, Kang, Hui, Ye, Ping, Yin, Jianqin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Generically Contrastive Spatiotemporal Representation Enhancement for 3D Skeleton Action Recognition
von: Zhang, Shaojie, et al.
Veröffentlicht: (2023)
von: Zhang, Shaojie, et al.
Veröffentlicht: (2023)
SiT-MLP: A Simple MLP with Point-wise Topology Feature Learning for Skeleton-based Action Recognition
von: Zhang, Shaojie, et al.
Veröffentlicht: (2023)
von: Zhang, Shaojie, et al.
Veröffentlicht: (2023)
DHRNet: A Dual-Path Hierarchical Relation Network for Multi-Person Pose Estimation
von: Dang, Yonghao, et al.
Veröffentlicht: (2024)
von: Dang, Yonghao, et al.
Veröffentlicht: (2024)
MaskSem: Semantic-Guided Masking for Learning 3D Hybrid High-Order Motion Representation
von: Wei, Wei, et al.
Veröffentlicht: (2025)
von: Wei, Wei, et al.
Veröffentlicht: (2025)
Leveraging the Video-level Semantic Consistency of Event for Audio-visual Event Localization
von: Jiang, Yuanyuan, et al.
Veröffentlicht: (2022)
von: Jiang, Yuanyuan, et al.
Veröffentlicht: (2022)
ActivityCLIP: Enhancing Group Activity Recognition by Mining Complementary Information from Text to Supplement Image Modality
von: Xu, Guoliang, et al.
Veröffentlicht: (2024)
von: Xu, Guoliang, et al.
Veröffentlicht: (2024)
Physics-constrained Attack against Convolution-based Human Motion Prediction
von: Duan, Chengxu, et al.
Veröffentlicht: (2023)
von: Duan, Chengxu, et al.
Veröffentlicht: (2023)
Kinematics Modeling Network for Video-based Human Pose Estimation
von: Dang, Yonghao, et al.
Veröffentlicht: (2022)
von: Dang, Yonghao, et al.
Veröffentlicht: (2022)
ESG-Net: Event-Aware Semantic Guided Network for Dense Audio-Visual Event Localization
von: Li, Huilai, et al.
Veröffentlicht: (2025)
von: Li, Huilai, et al.
Veröffentlicht: (2025)
L2HCount:Generalizing Crowd Counting from Low to High Crowd Density via Density Simulation
von: Xu, Guoliang, et al.
Veröffentlicht: (2025)
von: Xu, Guoliang, et al.
Veröffentlicht: (2025)
ResDiT: Evoking the Intrinsic Resolution Scalability in Diffusion Transformers
von: Ma, Yiyang, et al.
Veröffentlicht: (2025)
von: Ma, Yiyang, et al.
Veröffentlicht: (2025)
3DGAA: Realistic and Robust 3D Gaussian-based Adversarial Attack for Autonomous Driving
von: Zhang, Yixun, et al.
Veröffentlicht: (2025)
von: Zhang, Yixun, et al.
Veröffentlicht: (2025)
Mamba YOLO: A Simple Baseline for Object Detection with State Space Model
von: Wang, Zeyu, et al.
Veröffentlicht: (2024)
von: Wang, Zeyu, et al.
Veröffentlicht: (2024)
EAR: Enhancing Uni-Modal Representations for Weakly Supervised Audio-Visual Video Parsing
von: Li, Huilai, et al.
Veröffentlicht: (2026)
von: Li, Huilai, et al.
Veröffentlicht: (2026)
StyMam: A Mamba-Based Generator for Artistic Style Transfer
von: Hong, Zhou, et al.
Veröffentlicht: (2026)
von: Hong, Zhou, et al.
Veröffentlicht: (2026)
Towards Physically Realizable Adversarial Attacks in Embodied Vision Navigation
von: Chen, Meng, et al.
Veröffentlicht: (2024)
von: Chen, Meng, et al.
Veröffentlicht: (2024)
MamFusion: Multi-Mamba with Temporal Fusion for Partially Relevant Video Retrieval
von: Ying, Xinru, et al.
Veröffentlicht: (2025)
von: Ying, Xinru, et al.
Veröffentlicht: (2025)
BiTAA: A Bi-Task Adversarial Attack for Object Detection and Depth Estimation via 3D Gaussian Splatting
von: Zhang, Yixun, et al.
Veröffentlicht: (2025)
von: Zhang, Yixun, et al.
Veröffentlicht: (2025)
MambaIR: A Simple Baseline for Image Restoration with State-Space Model
von: Guo, Hang, et al.
Veröffentlicht: (2024)
von: Guo, Hang, et al.
Veröffentlicht: (2024)
MambaTrack: A Simple Baseline for Multiple Object Tracking with State Space Model
von: Xiao, Changcheng, et al.
Veröffentlicht: (2024)
von: Xiao, Changcheng, et al.
Veröffentlicht: (2024)
Mam-App: A Novel Parameter-Efficient Mamba Model for Apple Leaf Disease Classification
von: Mahamood, Md Nadim, et al.
Veröffentlicht: (2026)
von: Mahamood, Md Nadim, et al.
Veröffentlicht: (2026)
MedVLThinker: Simple Baselines for Multimodal Medical Reasoning
von: Huang, Xiaoke, et al.
Veröffentlicht: (2025)
von: Huang, Xiaoke, et al.
Veröffentlicht: (2025)
RT-DETRv2: Improved Baseline with Bag-of-Freebies for Real-Time Detection Transformer
von: Lv, Wenyu, et al.
Veröffentlicht: (2024)
von: Lv, Wenyu, et al.
Veröffentlicht: (2024)
Revisiting Simple Baselines for In-The-Wild Deepfake Detection
von: Castaneda, Orlando, et al.
Veröffentlicht: (2025)
von: Castaneda, Orlando, et al.
Veröffentlicht: (2025)
Towards more realistic human motion prediction with attention to motion coordination
von: Ding, Pengxiang, et al.
Veröffentlicht: (2024)
von: Ding, Pengxiang, et al.
Veröffentlicht: (2024)
CLIP-Powered TASS: Target-Aware Single-Stream Network for Audio-Visual Question Answering
von: Jiang, Yuanyuan, et al.
Veröffentlicht: (2024)
von: Jiang, Yuanyuan, et al.
Veröffentlicht: (2024)
A Two-stream Hybrid CNN-Transformer Network for Skeleton-based Human Interaction Recognition
von: Yin, Ruoqi, et al.
Veröffentlicht: (2023)
von: Yin, Ruoqi, et al.
Veröffentlicht: (2023)
SaMam: Style-aware State Space Model for Arbitrary Image Style Transfer
von: Liu, Hongda, et al.
Veröffentlicht: (2025)
von: Liu, Hongda, et al.
Veröffentlicht: (2025)
ER-Pose: Rethinking Keypoint-Driven Representation Learning for Real-Time Human Pose Estimation
von: Li, Nanjun, et al.
Veröffentlicht: (2026)
von: Li, Nanjun, et al.
Veröffentlicht: (2026)
CT3D++: Improving 3D Object Detection with Keypoint-induced Channel-wise Transformer
von: Sheng, Hualian, et al.
Veröffentlicht: (2024)
von: Sheng, Hualian, et al.
Veröffentlicht: (2024)
GLane3D : Detecting Lanes with Graph of 3D Keypoints
von: Öztürk, Halil İbrahim, et al.
Veröffentlicht: (2025)
von: Öztürk, Halil İbrahim, et al.
Veröffentlicht: (2025)
Open-Vocabulary Animal Keypoint Detection with Semantic-feature Matching
von: Zhang, Hao, et al.
Veröffentlicht: (2023)
von: Zhang, Hao, et al.
Veröffentlicht: (2023)
Geometry-Aware 3D Salient Object Detection Network
von: Wang, Chen, et al.
Veröffentlicht: (2025)
von: Wang, Chen, et al.
Veröffentlicht: (2025)
DaD: Distilled Reinforcement Learning for Diverse Keypoint Detection
von: Edstedt, Johan, et al.
Veröffentlicht: (2025)
von: Edstedt, Johan, et al.
Veröffentlicht: (2025)
MiM-ISTD: Mamba-in-Mamba for Efficient Infrared Small Target Detection
von: Chen, Tianxiang, et al.
Veröffentlicht: (2024)
von: Chen, Tianxiang, et al.
Veröffentlicht: (2024)
From Pairs to Sequences: Track-Aware Policy Gradients for Keypoint Detection
von: Liu, Yepeng, et al.
Veröffentlicht: (2026)
von: Liu, Yepeng, et al.
Veröffentlicht: (2026)
YOLOPoint Joint Keypoint and Object Detection
von: Backhaus, Anton, et al.
Veröffentlicht: (2024)
von: Backhaus, Anton, et al.
Veröffentlicht: (2024)
X-Pose: Detecting Any Keypoints
von: Yang, Jie, et al.
Veröffentlicht: (2023)
von: Yang, Jie, et al.
Veröffentlicht: (2023)
ConvMambaNet: A Hybrid CNN-Mamba State Space Architecture for Accurate and Real-Time EEG Seizure Detection
von: Khan, Md. Nishan, et al.
Veröffentlicht: (2026)
von: Khan, Md. Nishan, et al.
Veröffentlicht: (2026)
FIND: A Simple yet Effective Baseline for Diffusion-Generated Image Detection
von: Li, Jie, et al.
Veröffentlicht: (2026)
von: Li, Jie, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
A Generically Contrastive Spatiotemporal Representation Enhancement for 3D Skeleton Action Recognition
von: Zhang, Shaojie, et al.
Veröffentlicht: (2023) -
SiT-MLP: A Simple MLP with Point-wise Topology Feature Learning for Skeleton-based Action Recognition
von: Zhang, Shaojie, et al.
Veröffentlicht: (2023) -
DHRNet: A Dual-Path Hierarchical Relation Network for Multi-Person Pose Estimation
von: Dang, Yonghao, et al.
Veröffentlicht: (2024) -
MaskSem: Semantic-Guided Masking for Learning 3D Hybrid High-Order Motion Representation
von: Wei, Wei, et al.
Veröffentlicht: (2025) -
Leveraging the Video-level Semantic Consistency of Event for Audio-visual Event Localization
von: Jiang, Yuanyuan, et al.
Veröffentlicht: (2022)