MaskSem: Semantic-Guided Masking for Learning 3D Hybrid High-Order Motion Representation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wei, Wei, Zhang, Shaojie, Dang, Yonghao, Yin, Jianqin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Generically Contrastive Spatiotemporal Representation Enhancement for 3D Skeleton Action Recognition
von: Zhang, Shaojie, et al.
Veröffentlicht: (2023)
von: Zhang, Shaojie, et al.
Veröffentlicht: (2023)
SiT-MLP: A Simple MLP with Point-wise Topology Feature Learning for Skeleton-based Action Recognition
von: Zhang, Shaojie, et al.
Veröffentlicht: (2023)
von: Zhang, Shaojie, et al.
Veröffentlicht: (2023)
Kinematics Modeling Network for Video-based Human Pose Estimation
von: Dang, Yonghao, et al.
Veröffentlicht: (2022)
von: Dang, Yonghao, et al.
Veröffentlicht: (2022)
Leveraging the Video-level Semantic Consistency of Event for Audio-visual Event Localization
von: Jiang, Yuanyuan, et al.
Veröffentlicht: (2022)
von: Jiang, Yuanyuan, et al.
Veröffentlicht: (2022)
ESG-Net: Event-Aware Semantic Guided Network for Dense Audio-Visual Event Localization
von: Li, Huilai, et al.
Veröffentlicht: (2025)
von: Li, Huilai, et al.
Veröffentlicht: (2025)
Physics-constrained Attack against Convolution-based Human Motion Prediction
von: Duan, Chengxu, et al.
Veröffentlicht: (2023)
von: Duan, Chengxu, et al.
Veröffentlicht: (2023)
ActivityCLIP: Enhancing Group Activity Recognition by Mining Complementary Information from Text to Supplement Image Modality
von: Xu, Guoliang, et al.
Veröffentlicht: (2024)
von: Xu, Guoliang, et al.
Veröffentlicht: (2024)
MamKPD: A Simple Mamba Baseline for Real-Time 2D Keypoint Detection
von: Dang, Yonghao, et al.
Veröffentlicht: (2024)
von: Dang, Yonghao, et al.
Veröffentlicht: (2024)
3DGAA: Realistic and Robust 3D Gaussian-based Adversarial Attack for Autonomous Driving
von: Zhang, Yixun, et al.
Veröffentlicht: (2025)
von: Zhang, Yixun, et al.
Veröffentlicht: (2025)
L2HCount:Generalizing Crowd Counting from Low to High Crowd Density via Density Simulation
von: Xu, Guoliang, et al.
Veröffentlicht: (2025)
von: Xu, Guoliang, et al.
Veröffentlicht: (2025)
EAR: Enhancing Uni-Modal Representations for Weakly Supervised Audio-Visual Video Parsing
von: Li, Huilai, et al.
Veröffentlicht: (2026)
von: Li, Huilai, et al.
Veröffentlicht: (2026)
Towards Physically Realizable Adversarial Attacks in Embodied Vision Navigation
von: Chen, Meng, et al.
Veröffentlicht: (2024)
von: Chen, Meng, et al.
Veröffentlicht: (2024)
Learning Generalizable 3D Medical Image Representations from Mask-Guided Self-Supervision
von: Gao, Yunhe, et al.
Veröffentlicht: (2026)
von: Gao, Yunhe, et al.
Veröffentlicht: (2026)
SMILE: Infusing Spatial and Motion Semantics in Masked Video Learning
von: Thoker, Fida Mohammad, et al.
Veröffentlicht: (2025)
von: Thoker, Fida Mohammad, et al.
Veröffentlicht: (2025)
ResDiT: Evoking the Intrinsic Resolution Scalability in Diffusion Transformers
von: Ma, Yiyang, et al.
Veröffentlicht: (2025)
von: Ma, Yiyang, et al.
Veröffentlicht: (2025)
MaskGaussian: Adaptive 3D Gaussian Representation from Probabilistic Masks
von: Liu, Yifei, et al.
Veröffentlicht: (2024)
von: Liu, Yifei, et al.
Veröffentlicht: (2024)
ReMoMask: Retrieval-Augmented Masked Motion Generation
von: Li, Zhengdao, et al.
Veröffentlicht: (2025)
von: Li, Zhengdao, et al.
Veröffentlicht: (2025)
DHRNet: A Dual-Path Hierarchical Relation Network for Multi-Person Pose Estimation
von: Dang, Yonghao, et al.
Veröffentlicht: (2024)
von: Dang, Yonghao, et al.
Veröffentlicht: (2024)
Semantics-Guided Multimodal Masked Autoencoder Pretraining for 3D BEV Object Detection
von: Wariyapperuma, Prabuddhi, et al.
Veröffentlicht: (2026)
von: Wariyapperuma, Prabuddhi, et al.
Veröffentlicht: (2026)
Learning Multiple Representations with Inconsistency-Guided Detail Regularization for Mask-Guided Matting
von: Jiang, Weihao, et al.
Veröffentlicht: (2024)
von: Jiang, Weihao, et al.
Veröffentlicht: (2024)
T-MAE: Temporal Masked Autoencoders for Point Cloud Representation Learning
von: Wei, Weijie, et al.
Veröffentlicht: (2023)
von: Wei, Weijie, et al.
Veröffentlicht: (2023)
Social-MAE: Social Masked Autoencoder for Multi-person Motion Representation Learning
von: Ehsanpour, Mahsa, et al.
Veröffentlicht: (2024)
von: Ehsanpour, Mahsa, et al.
Veröffentlicht: (2024)
MMGT: Motion Mask Guided Two-Stage Network for Co-Speech Gesture Video Generation
von: Wang, Siyuan, et al.
Veröffentlicht: (2025)
von: Wang, Siyuan, et al.
Veröffentlicht: (2025)
SemTalk: Holistic Co-speech Motion Generation with Frame-level Semantic Emphasis
von: Zhang, Xiangyue, et al.
Veröffentlicht: (2024)
von: Zhang, Xiangyue, et al.
Veröffentlicht: (2024)
SARMAE: Masked Autoencoder for SAR Representation Learning
von: Liu, Danxu, et al.
Veröffentlicht: (2025)
von: Liu, Danxu, et al.
Veröffentlicht: (2025)
Attention-Guided Masked Autoencoders For Learning Image Representations
von: Sick, Leon, et al.
Veröffentlicht: (2024)
von: Sick, Leon, et al.
Veröffentlicht: (2024)
MaskControl: Spatio-Temporal Control for Masked Motion Synthesis
von: Pinyoanuntapong, Ekkasit, et al.
Veröffentlicht: (2024)
von: Pinyoanuntapong, Ekkasit, et al.
Veröffentlicht: (2024)
Frequency-Guided Masking for Enhanced Vision Self-Supervised Learning
von: Monsefi, Amin Karimi, et al.
Veröffentlicht: (2024)
von: Monsefi, Amin Karimi, et al.
Veröffentlicht: (2024)
Clean-GS: Semantic Mask-Guided Pruning for 3D Gaussian Splatting
von: Mishra, Subhankar
Veröffentlicht: (2026)
von: Mishra, Subhankar
Veröffentlicht: (2026)
Suppressing Non-Semantic Noise in Masked Image Modeling Representations
von: Hjelkrem-Tan, Martine, et al.
Veröffentlicht: (2026)
von: Hjelkrem-Tan, Martine, et al.
Veröffentlicht: (2026)
Learning Context-Adaptive Motion Priors for Masked Motion Diffusion Models with Efficient Kinematic Attention Aggregation
von: Jiang, Junkun, et al.
Veröffentlicht: (2026)
von: Jiang, Junkun, et al.
Veröffentlicht: (2026)
Improving Masked Autoencoders by Learning Where to Mask
von: Chen, Haijian, et al.
Veröffentlicht: (2023)
von: Chen, Haijian, et al.
Veröffentlicht: (2023)
Exclusivity-Guided Mask Learning for Semi-Supervised Crowd Instance Segmentation and Counting
von: Huang, Jiyang, et al.
Veröffentlicht: (2026)
von: Huang, Jiyang, et al.
Veröffentlicht: (2026)
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs
von: Zhang, Shaojie, et al.
Veröffentlicht: (2025)
von: Zhang, Shaojie, et al.
Veröffentlicht: (2025)
Jigsaw3D: Disentangled 3D Style Transfer via Patch Shuffling and Masking
von: Ye, Yuteng, et al.
Veröffentlicht: (2025)
von: Ye, Yuteng, et al.
Veröffentlicht: (2025)
MASA: Motion-aware Masked Autoencoder with Semantic Alignment for Sign Language Recognition
von: Zhao, Weichao, et al.
Veröffentlicht: (2024)
von: Zhao, Weichao, et al.
Veröffentlicht: (2024)
Masked Diffusion Captioning for Visual Feature Learning
von: Feng, Chao, et al.
Veröffentlicht: (2025)
von: Feng, Chao, et al.
Veröffentlicht: (2025)
ProtoMask: Segmentation-Guided Prototype Learning
von: Meinert, Steffen, et al.
Veröffentlicht: (2025)
von: Meinert, Steffen, et al.
Veröffentlicht: (2025)
HU-based Foreground Masking for 3D Medical Masked Image Modeling
von: Lee, Jin, et al.
Veröffentlicht: (2025)
von: Lee, Jin, et al.
Veröffentlicht: (2025)
SemanticMIM: Marring Masked Image Modeling with Semantics Compression for General Visual Representation
von: Yuan, Yike, et al.
Veröffentlicht: (2024)
von: Yuan, Yike, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
A Generically Contrastive Spatiotemporal Representation Enhancement for 3D Skeleton Action Recognition
von: Zhang, Shaojie, et al.
Veröffentlicht: (2023) -
SiT-MLP: A Simple MLP with Point-wise Topology Feature Learning for Skeleton-based Action Recognition
von: Zhang, Shaojie, et al.
Veröffentlicht: (2023) -
Kinematics Modeling Network for Video-based Human Pose Estimation
von: Dang, Yonghao, et al.
Veröffentlicht: (2022) -
Leveraging the Video-level Semantic Consistency of Event for Audio-visual Event Localization
von: Jiang, Yuanyuan, et al.
Veröffentlicht: (2022) -
ESG-Net: Event-Aware Semantic Guided Network for Dense Audio-Visual Event Localization
von: Li, Huilai, et al.
Veröffentlicht: (2025)