SMILE: Infusing Spatial and Motion Semantics in Masked Video Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Thoker, Fida Mohammad, Jiang, Letian, Zhao, Chen, Ghanem, Bernard |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TrackMAE: Video Representation Learning via Track Mask and Predict
by: Vandeghen, Renaud, et al.
Published: (2026)
by: Vandeghen, Renaud, et al.
Published: (2026)
SEVERE++: Evaluating Benchmark Sensitivity in Generalization of Video Representation Learning
by: Thoker, Fida Mohammad, et al.
Published: (2025)
by: Thoker, Fida Mohammad, et al.
Published: (2025)
LocoMotion: Learning Motion-Focused Video-Language Representations
by: Doughty, Hazel, et al.
Published: (2024)
by: Doughty, Hazel, et al.
Published: (2024)
SIGMA: Sinkhorn-Guided Masked Video Modeling
by: Salehi, Mohammadreza, et al.
Published: (2024)
by: Salehi, Mohammadreza, et al.
Published: (2024)
Video Self-Stitching Graph Network for Temporal Action Localization
by: Zhao, Chen, et al.
Published: (2020)
by: Zhao, Chen, et al.
Published: (2020)
FLORO: A Multimodal Geospatial Foundation Model for Ecological Remote Sensing Across Sensors and Scales
by: Rodriguez, Jorge L., et al.
Published: (2026)
by: Rodriguez, Jorge L., et al.
Published: (2026)
BOLT: Boost Large Vision-Language Model Without Training for Long-form Video Understanding
by: Liu, Shuming, et al.
Published: (2025)
by: Liu, Shuming, et al.
Published: (2025)
Hybrid Structure-from-Motion and Camera Relocalization for Enhanced Egocentric Localization
by: Mai, Jinjie, et al.
Published: (2024)
by: Mai, Jinjie, et al.
Published: (2024)
MoCA-Video: Motion-Aware Concept Alignment for Consistent Video Editing
by: Zhang, Tong, et al.
Published: (2025)
by: Zhang, Tong, et al.
Published: (2025)
$β$-CLIP: Text-Conditioned Contrastive Learning for Multi-Granular Vision-Language Alignment
by: Zohra, Fatimah, et al.
Published: (2025)
by: Zohra, Fatimah, et al.
Published: (2025)
Test-Time Adaptation for Combating Missing Modalities in Egocentric Videos
by: Ramazanova, Merey, et al.
Published: (2024)
by: Ramazanova, Merey, et al.
Published: (2024)
Audio-Infused Automatic Image Colorization by Exploiting Audio Scene Semantics
by: Zhao, Pengcheng, et al.
Published: (2024)
by: Zhao, Pengcheng, et al.
Published: (2024)
MaskSem: Semantic-Guided Masking for Learning 3D Hybrid High-Order Motion Representation
by: Wei, Wei, et al.
Published: (2025)
by: Wei, Wei, et al.
Published: (2025)
Learning Context-Adaptive Motion Priors for Masked Motion Diffusion Models with Efficient Kinematic Attention Aggregation
by: Jiang, Junkun, et al.
Published: (2026)
by: Jiang, Junkun, et al.
Published: (2026)
MASA: Motion-aware Masked Autoencoder with Semantic Alignment for Sign Language Recognition
by: Zhao, Weichao, et al.
Published: (2024)
by: Zhao, Weichao, et al.
Published: (2024)
End-to-End Temporal Action Detection with 1B Parameters Across 1000 Frames
by: Liu, Shuming, et al.
Published: (2023)
by: Liu, Shuming, et al.
Published: (2023)
SMC++: Masked Learning of Unsupervised Video Semantic Compression
by: Tian, Yuan, et al.
Published: (2024)
by: Tian, Yuan, et al.
Published: (2024)
Dynamically Masked Discriminator for Generative Adversarial Networks
by: Zhang, Wentian, et al.
Published: (2023)
by: Zhang, Wentian, et al.
Published: (2023)
SMILE: A Composite Lexical-Semantic Metric for Question-Answering Evaluation
by: Kendre, Shrikant, et al.
Published: (2025)
by: Kendre, Shrikant, et al.
Published: (2025)
Semantic Segmentation on VSPW Dataset through Masked Video Consistency
by: Liang, Chen, et al.
Published: (2024)
by: Liang, Chen, et al.
Published: (2024)
SynMotion: Semantic-Visual Adaptation for Motion Customized Video Generation
by: Tan, Shuai, et al.
Published: (2025)
by: Tan, Shuai, et al.
Published: (2025)
Learning Semantic Segmentation with Query Points Supervision on Aerial Images
by: Rivier, Santiago, et al.
Published: (2023)
by: Rivier, Santiago, et al.
Published: (2023)
Infused Suppression Of Magnification Artefacts For Micro-AU Detection
by: Khor, Huai-Qian, et al.
Published: (2025)
by: Khor, Huai-Qian, et al.
Published: (2025)
SMILE: A Super-resolution Guided Multi-task Learning Method for Hyperspectral Unmixing
by: Li, Ruiying, et al.
Published: (2025)
by: Li, Ruiying, et al.
Published: (2025)
Investigating Event-Based Cameras for Video Frame Interpolation in Sports
by: Deckyvere, Antoine, et al.
Published: (2024)
by: Deckyvere, Antoine, et al.
Published: (2024)
SoccerNet-Caption: Dense Video Captioning for Soccer Broadcasts Commentaries
by: Mkhallati, Hassan, et al.
Published: (2023)
by: Mkhallati, Hassan, et al.
Published: (2023)
Towards Active Learning for Action Spotting in Association Football Videos
by: Giancola, Silvio, et al.
Published: (2023)
by: Giancola, Silvio, et al.
Published: (2023)
SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving
by: Li, Peizheng, et al.
Published: (2025)
by: Li, Peizheng, et al.
Published: (2025)
MaskControl: Spatio-Temporal Control for Masked Motion Synthesis
by: Pinyoanuntapong, Ekkasit, et al.
Published: (2024)
by: Pinyoanuntapong, Ekkasit, et al.
Published: (2024)
Text-driven Human Motion Generation with Motion Masked Diffusion Model
by: Chen, Xingyu
Published: (2024)
by: Chen, Xingyu
Published: (2024)
SAVeS: Steering Safety Judgments in Vision-Language Models via Semantic Cues
by: Hinojosa, Carlos, et al.
Published: (2026)
by: Hinojosa, Carlos, et al.
Published: (2026)
VideoMAC: Video Masked Autoencoders Meet ConvNets
by: Pei, Gensheng, et al.
Published: (2024)
by: Pei, Gensheng, et al.
Published: (2024)
Vision-EKIPL: External Knowledge-Infused Policy Learning for Visual Reasoning
by: Wang, Chaoyang, et al.
Published: (2025)
by: Wang, Chaoyang, et al.
Published: (2025)
ColorMAE: Exploring data-independent masking strategies in Masked AutoEncoders
by: Hinojosa, Carlos, et al.
Published: (2024)
by: Hinojosa, Carlos, et al.
Published: (2024)
Follow-Your-Motion: Video Motion Transfer via Efficient Spatial-Temporal Decoupled Finetuning
by: Ma, Yue, et al.
Published: (2025)
by: Ma, Yue, et al.
Published: (2025)
Think Before You Diffuse: Infusing Physical Rules into Video Diffusion
by: Zhang, Ke, et al.
Published: (2025)
by: Zhang, Ke, et al.
Published: (2025)
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation
by: Cuttano, Claudia, et al.
Published: (2024)
by: Cuttano, Claudia, et al.
Published: (2024)
HybridFlow: Infusing Continuity into Masked Codebook for Extreme Low-Bitrate Image Compression
by: Lu, Lei, et al.
Published: (2024)
by: Lu, Lei, et al.
Published: (2024)
OSL-ActionSpotting: A Unified Library for Action Spotting in Sports Videos
by: Benzakour, Yassine, et al.
Published: (2024)
by: Benzakour, Yassine, et al.
Published: (2024)
NearID: Identity Representation Learning via Near-identity Distractors
by: Cvejic, Aleksandar, et al.
Published: (2026)
by: Cvejic, Aleksandar, et al.
Published: (2026)
Similar Items
-
TrackMAE: Video Representation Learning via Track Mask and Predict
by: Vandeghen, Renaud, et al.
Published: (2026) -
SEVERE++: Evaluating Benchmark Sensitivity in Generalization of Video Representation Learning
by: Thoker, Fida Mohammad, et al.
Published: (2025) -
LocoMotion: Learning Motion-Focused Video-Language Representations
by: Doughty, Hazel, et al.
Published: (2024) -
SIGMA: Sinkhorn-Guided Masked Video Modeling
by: Salehi, Mohammadreza, et al.
Published: (2024) -
Video Self-Stitching Graph Network for Temporal Action Localization
by: Zhao, Chen, et al.
Published: (2020)