MaskSem: Semantic-Guided Masking for Learning 3D Hybrid High-Order Motion Representation
Fuente:
arXiv
Saved in:
| Main Authors: | Wei, Wei, Zhang, Shaojie, Dang, Yonghao, Yin, Jianqin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Generically Contrastive Spatiotemporal Representation Enhancement for 3D Skeleton Action Recognition
by: Zhang, Shaojie, et al.
Published: (2023)
by: Zhang, Shaojie, et al.
Published: (2023)
SiT-MLP: A Simple MLP with Point-wise Topology Feature Learning for Skeleton-based Action Recognition
by: Zhang, Shaojie, et al.
Published: (2023)
by: Zhang, Shaojie, et al.
Published: (2023)
Kinematics Modeling Network for Video-based Human Pose Estimation
by: Dang, Yonghao, et al.
Published: (2022)
by: Dang, Yonghao, et al.
Published: (2022)
Leveraging the Video-level Semantic Consistency of Event for Audio-visual Event Localization
by: Jiang, Yuanyuan, et al.
Published: (2022)
by: Jiang, Yuanyuan, et al.
Published: (2022)
ESG-Net: Event-Aware Semantic Guided Network for Dense Audio-Visual Event Localization
by: Li, Huilai, et al.
Published: (2025)
by: Li, Huilai, et al.
Published: (2025)
Physics-constrained Attack against Convolution-based Human Motion Prediction
by: Duan, Chengxu, et al.
Published: (2023)
by: Duan, Chengxu, et al.
Published: (2023)
ActivityCLIP: Enhancing Group Activity Recognition by Mining Complementary Information from Text to Supplement Image Modality
by: Xu, Guoliang, et al.
Published: (2024)
by: Xu, Guoliang, et al.
Published: (2024)
MamKPD: A Simple Mamba Baseline for Real-Time 2D Keypoint Detection
by: Dang, Yonghao, et al.
Published: (2024)
by: Dang, Yonghao, et al.
Published: (2024)
3DGAA: Realistic and Robust 3D Gaussian-based Adversarial Attack for Autonomous Driving
by: Zhang, Yixun, et al.
Published: (2025)
by: Zhang, Yixun, et al.
Published: (2025)
L2HCount:Generalizing Crowd Counting from Low to High Crowd Density via Density Simulation
by: Xu, Guoliang, et al.
Published: (2025)
by: Xu, Guoliang, et al.
Published: (2025)
EAR: Enhancing Uni-Modal Representations for Weakly Supervised Audio-Visual Video Parsing
by: Li, Huilai, et al.
Published: (2026)
by: Li, Huilai, et al.
Published: (2026)
Towards Physically Realizable Adversarial Attacks in Embodied Vision Navigation
by: Chen, Meng, et al.
Published: (2024)
by: Chen, Meng, et al.
Published: (2024)
Learning Generalizable 3D Medical Image Representations from Mask-Guided Self-Supervision
by: Gao, Yunhe, et al.
Published: (2026)
by: Gao, Yunhe, et al.
Published: (2026)
SMILE: Infusing Spatial and Motion Semantics in Masked Video Learning
by: Thoker, Fida Mohammad, et al.
Published: (2025)
by: Thoker, Fida Mohammad, et al.
Published: (2025)
ResDiT: Evoking the Intrinsic Resolution Scalability in Diffusion Transformers
by: Ma, Yiyang, et al.
Published: (2025)
by: Ma, Yiyang, et al.
Published: (2025)
MaskGaussian: Adaptive 3D Gaussian Representation from Probabilistic Masks
by: Liu, Yifei, et al.
Published: (2024)
by: Liu, Yifei, et al.
Published: (2024)
ReMoMask: Retrieval-Augmented Masked Motion Generation
by: Li, Zhengdao, et al.
Published: (2025)
by: Li, Zhengdao, et al.
Published: (2025)
DHRNet: A Dual-Path Hierarchical Relation Network for Multi-Person Pose Estimation
by: Dang, Yonghao, et al.
Published: (2024)
by: Dang, Yonghao, et al.
Published: (2024)
Semantics-Guided Multimodal Masked Autoencoder Pretraining for 3D BEV Object Detection
by: Wariyapperuma, Prabuddhi, et al.
Published: (2026)
by: Wariyapperuma, Prabuddhi, et al.
Published: (2026)
Learning Multiple Representations with Inconsistency-Guided Detail Regularization for Mask-Guided Matting
by: Jiang, Weihao, et al.
Published: (2024)
by: Jiang, Weihao, et al.
Published: (2024)
T-MAE: Temporal Masked Autoencoders for Point Cloud Representation Learning
by: Wei, Weijie, et al.
Published: (2023)
by: Wei, Weijie, et al.
Published: (2023)
Social-MAE: Social Masked Autoencoder for Multi-person Motion Representation Learning
by: Ehsanpour, Mahsa, et al.
Published: (2024)
by: Ehsanpour, Mahsa, et al.
Published: (2024)
MMGT: Motion Mask Guided Two-Stage Network for Co-Speech Gesture Video Generation
by: Wang, Siyuan, et al.
Published: (2025)
by: Wang, Siyuan, et al.
Published: (2025)
SemTalk: Holistic Co-speech Motion Generation with Frame-level Semantic Emphasis
by: Zhang, Xiangyue, et al.
Published: (2024)
by: Zhang, Xiangyue, et al.
Published: (2024)
SARMAE: Masked Autoencoder for SAR Representation Learning
by: Liu, Danxu, et al.
Published: (2025)
by: Liu, Danxu, et al.
Published: (2025)
Attention-Guided Masked Autoencoders For Learning Image Representations
by: Sick, Leon, et al.
Published: (2024)
by: Sick, Leon, et al.
Published: (2024)
MaskControl: Spatio-Temporal Control for Masked Motion Synthesis
by: Pinyoanuntapong, Ekkasit, et al.
Published: (2024)
by: Pinyoanuntapong, Ekkasit, et al.
Published: (2024)
Frequency-Guided Masking for Enhanced Vision Self-Supervised Learning
by: Monsefi, Amin Karimi, et al.
Published: (2024)
by: Monsefi, Amin Karimi, et al.
Published: (2024)
Clean-GS: Semantic Mask-Guided Pruning for 3D Gaussian Splatting
by: Mishra, Subhankar
Published: (2026)
by: Mishra, Subhankar
Published: (2026)
Suppressing Non-Semantic Noise in Masked Image Modeling Representations
by: Hjelkrem-Tan, Martine, et al.
Published: (2026)
by: Hjelkrem-Tan, Martine, et al.
Published: (2026)
Learning Context-Adaptive Motion Priors for Masked Motion Diffusion Models with Efficient Kinematic Attention Aggregation
by: Jiang, Junkun, et al.
Published: (2026)
by: Jiang, Junkun, et al.
Published: (2026)
Improving Masked Autoencoders by Learning Where to Mask
by: Chen, Haijian, et al.
Published: (2023)
by: Chen, Haijian, et al.
Published: (2023)
Exclusivity-Guided Mask Learning for Semi-Supervised Crowd Instance Segmentation and Counting
by: Huang, Jiyang, et al.
Published: (2026)
by: Huang, Jiyang, et al.
Published: (2026)
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs
by: Zhang, Shaojie, et al.
Published: (2025)
by: Zhang, Shaojie, et al.
Published: (2025)
Jigsaw3D: Disentangled 3D Style Transfer via Patch Shuffling and Masking
by: Ye, Yuteng, et al.
Published: (2025)
by: Ye, Yuteng, et al.
Published: (2025)
MASA: Motion-aware Masked Autoencoder with Semantic Alignment for Sign Language Recognition
by: Zhao, Weichao, et al.
Published: (2024)
by: Zhao, Weichao, et al.
Published: (2024)
Masked Diffusion Captioning for Visual Feature Learning
by: Feng, Chao, et al.
Published: (2025)
by: Feng, Chao, et al.
Published: (2025)
ProtoMask: Segmentation-Guided Prototype Learning
by: Meinert, Steffen, et al.
Published: (2025)
by: Meinert, Steffen, et al.
Published: (2025)
HU-based Foreground Masking for 3D Medical Masked Image Modeling
by: Lee, Jin, et al.
Published: (2025)
by: Lee, Jin, et al.
Published: (2025)
SemanticMIM: Marring Masked Image Modeling with Semantics Compression for General Visual Representation
by: Yuan, Yike, et al.
Published: (2024)
by: Yuan, Yike, et al.
Published: (2024)
Similar Items
-
A Generically Contrastive Spatiotemporal Representation Enhancement for 3D Skeleton Action Recognition
by: Zhang, Shaojie, et al.
Published: (2023) -
SiT-MLP: A Simple MLP with Point-wise Topology Feature Learning for Skeleton-based Action Recognition
by: Zhang, Shaojie, et al.
Published: (2023) -
Kinematics Modeling Network for Video-based Human Pose Estimation
by: Dang, Yonghao, et al.
Published: (2022) -
Leveraging the Video-level Semantic Consistency of Event for Audio-visual Event Localization
by: Jiang, Yuanyuan, et al.
Published: (2022) -
ESG-Net: Event-Aware Semantic Guided Network for Dense Audio-Visual Event Localization
by: Li, Huilai, et al.
Published: (2025)