EAST: Early Action Prediction Sampling Strategy with Token Masking
Fuente:
arXiv
Saved in:
| Main Authors: | Sović, Iva, Martinović, Ivan, Oršić, Marin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DEARLi: Decoupled Enhancement of Recognition and Localization for Semi-supervised Panoptic Segmentation
by: Martinović, Ivan, et al.
Published: (2025)
by: Martinović, Ivan, et al.
Published: (2025)
Dense outlier detection and open-set recognition based on training with noisy negative images
by: Bevandić, Petra, et al.
Published: (2021)
by: Bevandić, Petra, et al.
Published: (2021)
MC-PanDA: Mask Confidence for Panoptic Domain Adaptation
by: Martinović, Ivan, et al.
Published: (2024)
by: Martinović, Ivan, et al.
Published: (2024)
Weakly supervised training of universal visual concepts for multi-domain semantic segmentation
by: Bevandić, Petra, et al.
Published: (2022)
by: Bevandić, Petra, et al.
Published: (2022)
A Survey on Training-free Open-Vocabulary Semantic Segmentation
by: Kombol, Naomi, et al.
Published: (2025)
by: Kombol, Naomi, et al.
Published: (2025)
Computational Tradeoffs in Image Synthesis: Diffusion, Masked-Token, and Next-Token Prediction
by: Kilian, Maciej, et al.
Published: (2024)
by: Kilian, Maciej, et al.
Published: (2024)
Token-Space Mask Prediction for Efficient Vision Transformer Segmentation
by: Galagain, Calvin, et al.
Published: (2026)
by: Galagain, Calvin, et al.
Published: (2026)
Medical Referring Image Segmentation via Next-Token Mask Prediction
by: Chen, Xinyu, et al.
Published: (2025)
by: Chen, Xinyu, et al.
Published: (2025)
SAM-Guided Masked Token Prediction for 3D Scene Understanding
by: Chen, Zhimin, et al.
Published: (2024)
by: Chen, Zhimin, et al.
Published: (2024)
SPAR: Single-Pass Any-Resolution ViT for Open-vocabulary Segmentation
by: Kombol, Naomi, et al.
Published: (2026)
by: Kombol, Naomi, et al.
Published: (2026)
What Holds Back Open-Vocabulary Segmentation?
by: Šarić, Josip, et al.
Published: (2025)
by: Šarić, Josip, et al.
Published: (2025)
WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens
by: Wang, Xiaofeng, et al.
Published: (2024)
by: Wang, Xiaofeng, et al.
Published: (2024)
ALTo: Adaptive-Length Tokenizer for Autoregressive Mask Generation
by: Wang, Lingfeng, et al.
Published: (2025)
by: Wang, Lingfeng, et al.
Published: (2025)
Emerging Property of Masked Token for Effective Pre-training
by: Choi, Hyesong, et al.
Published: (2024)
by: Choi, Hyesong, et al.
Published: (2024)
Morphing Tokens Draw Strong Masked Image Models
by: Kim, Taekyung, et al.
Published: (2023)
by: Kim, Taekyung, et al.
Published: (2023)
MaskFuser: Masked Fusion of Joint Multi-Modal Tokenization for End-to-End Autonomous Driving
by: Duan, Yiqun, et al.
Published: (2024)
by: Duan, Yiqun, et al.
Published: (2024)
TEMPURA: Temporal Event Masked Prediction and Understanding for Reasoning in Action
by: Cheng, Jen-Hao, et al.
Published: (2025)
by: Cheng, Jen-Hao, et al.
Published: (2025)
Beyond ViT Tokens: Masked-Diffusion Pretrained Convolutional Pathology Foundation Model for Cell-Level Dense Prediction
by: Chen, Weiming, et al.
Published: (2026)
by: Chen, Weiming, et al.
Published: (2026)
T4P: Test-Time Training of Trajectory Prediction via Masked Autoencoder and Actor-specific Token Memory
by: Park, Daehee, et al.
Published: (2024)
by: Park, Daehee, et al.
Published: (2024)
Improved Masked Image Generation with Knowledge-Augmented Token Representations
by: Liang, Guotao, et al.
Published: (2025)
by: Liang, Guotao, et al.
Published: (2025)
SeiT++: Masked Token Modeling Improves Storage-efficient Training
by: Lee, Minhyun, et al.
Published: (2023)
by: Lee, Minhyun, et al.
Published: (2023)
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction
by: Yang, Shu-wen, et al.
Published: (2025)
by: Yang, Shu-wen, et al.
Published: (2025)
EarlyTom: Early Token Compression Completes Fast Video Understanding
by: Wang, Hesong, et al.
Published: (2026)
by: Wang, Hesong, et al.
Published: (2026)
Vision Transformer with Super Token Sampling
by: Huang, Huaibo, et al.
Published: (2022)
by: Huang, Huaibo, et al.
Published: (2022)
Masked Diffusion Vision-Language Models for Temporal Action Localization
by: Wang, Fengshun, et al.
Published: (2026)
by: Wang, Fengshun, et al.
Published: (2026)
Kronecker Mask and Interpretive Prompts are Language-Action Video Learners
by: Yang, Jingyi, et al.
Published: (2025)
by: Yang, Jingyi, et al.
Published: (2025)
MaskMed: Decoupled Mask and Class Prediction for Medical Image Segmentation
by: Xie, Bin, et al.
Published: (2025)
by: Xie, Bin, et al.
Published: (2025)
Salience-Based Adaptive Masking: Revisiting Token Dynamics for Enhanced Pre-training
by: Choi, Hyesong, et al.
Published: (2024)
by: Choi, Hyesong, et al.
Published: (2024)
Hierarchical Masked Autoregressive Models with Low-Resolution Token Pivots
by: Zheng, Guangting, et al.
Published: (2025)
by: Zheng, Guangting, et al.
Published: (2025)
Masked Video and Body-worn IMU Autoencoder for Egocentric Action Recognition
by: Zhang, Mingfang, et al.
Published: (2024)
by: Zhang, Mingfang, et al.
Published: (2024)
HiMTok: Learning Hierarchical Mask Tokens for Image Segmentation with Large Multimodal Model
by: Wang, Tao, et al.
Published: (2025)
by: Wang, Tao, et al.
Published: (2025)
Reinforcement Learning meets Masked Video Modeling : Trajectory-Guided Adaptive Token Selection
by: Rai, Ayush K., et al.
Published: (2025)
by: Rai, Ayush K., et al.
Published: (2025)
RedVTP: Training-Free Acceleration of Diffusion Vision-Language Models Inference via Masked Token-Guided Visual Token Pruning
by: Xu, Jingqi, et al.
Published: (2025)
by: Xu, Jingqi, et al.
Published: (2025)
Unsupervised Anomaly Detection via Masked Diffusion Posterior Sampling
by: Wu, Di, et al.
Published: (2024)
by: Wu, Di, et al.
Published: (2024)
Cross-Modal Action Recognition in Egocentric Video Using Mamba: Integrating RGB and Hand Skeleton Streams via CLS Token Fusion Strategies
by: Gorostegui, Juan Ignacio Bustos, et al.
Published: (2026)
by: Gorostegui, Juan Ignacio Bustos, et al.
Published: (2026)
Event Masked Autoencoder: Point-wise Action Recognition with Event-Based Cameras
by: Sun, Jingkai, et al.
Published: (2025)
by: Sun, Jingkai, et al.
Published: (2025)
Token Painter: Training-Free Text-Guided Image Inpainting via Mask Autoregressive Models
by: Jiang, Longtao, et al.
Published: (2025)
by: Jiang, Longtao, et al.
Published: (2025)
Know Your Attention Maps: Class-specific Token Masking for Weakly Supervised Semantic Segmentation
by: Hanna, Joelle, et al.
Published: (2025)
by: Hanna, Joelle, et al.
Published: (2025)
Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens
by: Kim, Dongwon, et al.
Published: (2025)
by: Kim, Dongwon, et al.
Published: (2025)
Trajectory-aligned Space-time Tokens for Few-shot Action Recognition
by: Kumar, Pulkit, et al.
Published: (2024)
by: Kumar, Pulkit, et al.
Published: (2024)
Similar Items
-
DEARLi: Decoupled Enhancement of Recognition and Localization for Semi-supervised Panoptic Segmentation
by: Martinović, Ivan, et al.
Published: (2025) -
Dense outlier detection and open-set recognition based on training with noisy negative images
by: Bevandić, Petra, et al.
Published: (2021) -
MC-PanDA: Mask Confidence for Panoptic Domain Adaptation
by: Martinović, Ivan, et al.
Published: (2024) -
Weakly supervised training of universal visual concepts for multi-domain semantic segmentation
by: Bevandić, Petra, et al.
Published: (2022) -
A Survey on Training-free Open-Vocabulary Semantic Segmentation
by: Kombol, Naomi, et al.
Published: (2025)