Efficient High-Resolution Visual Representation Learning with State Space Model for Human Pose Estimation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Hao, Ma, Yongqiang, Shao, Wenqi, Luo, Ping, Zheng, Nanning, Zhang, Kaipeng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
B-AVIBench: Towards Evaluating the Robustness of Large Vision-Language Model on Black-box Adversarial Visual-Instructions
by: Zhang, Hao, et al.
Published: (2024)
by: Zhang, Hao, et al.
Published: (2024)
Open-Vocabulary Animal Keypoint Detection with Semantic-feature Matching
by: Zhang, Hao, et al.
Published: (2023)
by: Zhang, Hao, et al.
Published: (2023)
UniHOI: Unified Human-Object Interaction Understanding via Unified Token Space
by: Yang, Panqi, et al.
Published: (2025)
by: Yang, Panqi, et al.
Published: (2025)
MLLMs-Augmented Visual-Language Representation Learning
by: Liu, Yanqing, et al.
Published: (2023)
by: Liu, Yanqing, et al.
Published: (2023)
EVA: Mixture-of-Experts Semantic Variant Alignment for Compositional Zero-Shot Learning
by: Zhang, Xiao, et al.
Published: (2025)
by: Zhang, Xiao, et al.
Published: (2025)
High-Resolution Spatiotemporal Modeling with Global-Local State Space Models for Video-Based Human Pose Estimation
by: Feng, Runyang, et al.
Published: (2025)
by: Feng, Runyang, et al.
Published: (2025)
SHaRPose: Sparse High-Resolution Representation for Human Pose Estimation
by: An, Xiaoqi, et al.
Published: (2023)
by: An, Xiaoqi, et al.
Published: (2023)
See Through Their Minds: Learning Transferable Neural Representation from Cross-Subject fMRI
by: Liu, Yulong, et al.
Published: (2024)
by: Liu, Yulong, et al.
Published: (2024)
SAMRefiner: Taming Segment Anything Model for Universal Mask Refinement
by: Lin, Yuqi, et al.
Published: (2025)
by: Lin, Yuqi, et al.
Published: (2025)
Dite-HRNet: Dynamic Lightweight High-Resolution Network for Human Pose Estimation
by: Li, Qun, et al.
Published: (2022)
by: Li, Qun, et al.
Published: (2022)
Focus on Low-Resolution Information: Multi-Granular Information-Lossless Model for Low-Resolution Human Pose Estimation
by: Gu, Zejun, et al.
Published: (2024)
by: Gu, Zejun, et al.
Published: (2024)
ChartAssisstant: A Universal Chart Multimodal Language Model via Chart-to-Table Pre-training and Multitask Instruction Tuning
by: Meng, Fanqing, et al.
Published: (2024)
by: Meng, Fanqing, et al.
Published: (2024)
A Compact Hybrid Convolution--Frequency State Space Network for Learned Image Compression
by: Pan, Haodong, et al.
Published: (2025)
by: Pan, Haodong, et al.
Published: (2025)
Rethinking Human Evaluation Protocol for Text-to-Video Models: Enhancing Reliability,Reproducibility, and Practicality
by: Zhang, Tianle, et al.
Published: (2024)
by: Zhang, Tianle, et al.
Published: (2024)
Learning Brain Tumor Representation in 3D High-Resolution MR Images via Interpretable State Space Models
by: Hu, Qingqiao, et al.
Published: (2024)
by: Hu, Qingqiao, et al.
Published: (2024)
Dynamic Multimodal Evaluation with Flexible Complexity by Vision-Language Bootstrapping
by: Yang, Yue, et al.
Published: (2024)
by: Yang, Yue, et al.
Published: (2024)
Cross-Domain Knowledge Distillation for Low-Resolution Human Pose Estimation
by: Gu, Zejun, et al.
Published: (2024)
by: Gu, Zejun, et al.
Published: (2024)
Diffree: Text-Guided Shape Free Object Inpainting with Diffusion Model
by: Zhao, Lirui, et al.
Published: (2024)
by: Zhao, Lirui, et al.
Published: (2024)
Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model
by: Zhu, Lianghui, et al.
Published: (2024)
by: Zhu, Lianghui, et al.
Published: (2024)
STAR-Pose: Efficient Low-Resolution Video Human Pose Estimation via Spatial-Temporal Adaptive Super-Resolution
by: Jin, Yucheng, et al.
Published: (2025)
by: Jin, Yucheng, et al.
Published: (2025)
TREND: Unsupervised 3D Representation Learning via Temporal Forecasting for LiDAR Perception
by: Chen, Runjian, et al.
Published: (2024)
by: Chen, Runjian, et al.
Published: (2024)
Voxel or Pillar: Exploring Efficient Point Cloud Representation for 3D Object Detection
by: Huang, Yuhao, et al.
Published: (2023)
by: Huang, Yuhao, et al.
Published: (2023)
LGM-Pose: A Lightweight Global Modeling Network for Real-time Human Pose Estimation
by: Guo, Biao, et al.
Published: (2025)
by: Guo, Biao, et al.
Published: (2025)
Greit-HRNet: Grouped Lightweight High-Resolution Network for Human Pose Estimation
by: Han, Junjia
Published: (2024)
by: Han, Junjia
Published: (2024)
ZipVL: Efficient Large Vision-Language Models with Dynamic Token Sparsification
by: He, Yefei, et al.
Published: (2024)
by: He, Yefei, et al.
Published: (2024)
VMambaCC: A Visual State Space Model for Crowd Counting
by: Ma, Hao-Yuan, et al.
Published: (2024)
by: Ma, Hao-Yuan, et al.
Published: (2024)
CLAP: Unsupervised 3D Representation Learning for Fusion 3D Perception via Curvature Sampling and Prototype Learning
by: Chen, Runjian, et al.
Published: (2024)
by: Chen, Runjian, et al.
Published: (2024)
TP-Eval: Tap Multimodal LLMs' Potential in Evaluation by Customizing Prompts
by: Xie, Yuxuan, et al.
Published: (2024)
by: Xie, Yuxuan, et al.
Published: (2024)
PO-MSCKF: An Efficient Visual-Inertial Odometry by Reconstructing the Multi-State Constrained Kalman Filter with the Pose-only Theory
by: Du, Xueyu, et al.
Published: (2024)
by: Du, Xueyu, et al.
Published: (2024)
Learning Efficient and Generalizable Human Representation with Human Gaussian Model
by: Liu, Yifan, et al.
Published: (2025)
by: Liu, Yifan, et al.
Published: (2025)
Exploiting Aggregation and Segregation of Representations for Domain Adaptive Human Pose Estimation
by: Peng, Qucheng, et al.
Published: (2024)
by: Peng, Qucheng, et al.
Published: (2024)
Efficient Sparse-to-Dense Visual Localization via Compact Gaussian Scene Representation and Accelerated Dense Pose Estimation
by: Li, Zizhuo, et al.
Published: (2026)
by: Li, Zizhuo, et al.
Published: (2026)
ER-Pose: Rethinking Keypoint-Driven Representation Learning for Real-Time Human Pose Estimation
by: Li, Nanjun, et al.
Published: (2026)
by: Li, Nanjun, et al.
Published: (2026)
PoseMamba: Monocular 3D Human Pose Estimation with Bidirectional Global-Local Spatio-Temporal State Space Model
by: Huang, Yunlong, et al.
Published: (2024)
by: Huang, Yunlong, et al.
Published: (2024)
TinyLVLM-eHub: Towards Comprehensive and Efficient Evaluation for Large Vision-Language Models
by: Shao, Wenqi, et al.
Published: (2023)
by: Shao, Wenqi, et al.
Published: (2023)
Advanced Object Detection and Pose Estimation with Hybrid Task Cascade and High-Resolution Networks
by: Jin, Yuhui, et al.
Published: (2025)
by: Jin, Yuhui, et al.
Published: (2025)
MV-SSM: Multi-View State Space Modeling for 3D Human Pose Estimation
by: Chharia, Aviral, et al.
Published: (2025)
by: Chharia, Aviral, et al.
Published: (2025)
PhyBench: A Physical Commonsense Benchmark for Evaluating Text-to-Image Models
by: Meng, Fanqing, et al.
Published: (2024)
by: Meng, Fanqing, et al.
Published: (2024)
SearchLVLMs: A Plug-and-Play Framework for Augmenting Large Vision-Language Models by Searching Up-to-Date Internet Knowledge
by: Li, Chuanhao, et al.
Published: (2024)
by: Li, Chuanhao, et al.
Published: (2024)
XPose: eXplainable Human Pose Estimation
by: Qiu, Luyu, et al.
Published: (2024)
by: Qiu, Luyu, et al.
Published: (2024)
Similar Items
-
B-AVIBench: Towards Evaluating the Robustness of Large Vision-Language Model on Black-box Adversarial Visual-Instructions
by: Zhang, Hao, et al.
Published: (2024) -
Open-Vocabulary Animal Keypoint Detection with Semantic-feature Matching
by: Zhang, Hao, et al.
Published: (2023) -
UniHOI: Unified Human-Object Interaction Understanding via Unified Token Space
by: Yang, Panqi, et al.
Published: (2025) -
MLLMs-Augmented Visual-Language Representation Learning
by: Liu, Yanqing, et al.
Published: (2023) -
EVA: Mixture-of-Experts Semantic Variant Alignment for Compositional Zero-Shot Learning
by: Zhang, Xiao, et al.
Published: (2025)