Motion Guided Token Compression for Efficient Masked Video Modeling
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Feng, Yukun, Shi, Yangming, Liu, Fengze, Yan, Tan |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Resource-Efficient Motion Control for Video Generation via Dynamic Mask Guidance
par: Feng, Sicong, et autres
Publié: (2025)
par: Feng, Sicong, et autres
Publié: (2025)
Fre-Res: Frequency-Residual Video Token Compression for Efficient Video MLLMs
par: Feng, Yigui, et autres
Publié: (2026)
par: Feng, Yigui, et autres
Publié: (2026)
DC-AR: Efficient Masked Autoregressive Image Generation with Deep Compression Hybrid Tokenizer
par: Wu, Yecheng, et autres
Publié: (2025)
par: Wu, Yecheng, et autres
Publié: (2025)
StreamingTOM: Streaming Token Compression for Efficient Video Understanding
par: Chen, Xueyi, et autres
Publié: (2025)
par: Chen, Xueyi, et autres
Publié: (2025)
QMoP: Query Guided Mixture-of-Projector for Efficient Visual Token Compression
par: Li, Zhongyang, et autres
Publié: (2026)
par: Li, Zhongyang, et autres
Publié: (2026)
Contribution-aware Token Compression for Efficient Video Understanding via Reinforcement Learning
par: Ma, Yinchao, et autres
Publié: (2026)
par: Ma, Yinchao, et autres
Publié: (2026)
Factorized-Dreamer: Training A High-Quality Video Generator with Limited and Low-Quality Data
par: Yang, Tao, et autres
Publié: (2024)
par: Yang, Tao, et autres
Publié: (2024)
TokenMotion: Decoupled Motion Control via Token Disentanglement for Human-centric Video Generation
par: Li, Ruineng, et autres
Publié: (2025)
par: Li, Ruineng, et autres
Publié: (2025)
FCoT-VL:Advancing Text-oriented Large Vision-Language Models with Efficient Visual Token Compression
par: Li, Jianjian, et autres
Publié: (2025)
par: Li, Jianjian, et autres
Publié: (2025)
PhyVLLM: Physics-Guided Video Language Model with Motion-Appearance Disentanglement
par: Zhan, Yu-Wei, et autres
Publié: (2025)
par: Zhan, Yu-Wei, et autres
Publié: (2025)
Geometry-Guided Camera Motion Understanding in VideoLLMs
par: Feng, Haoan, et autres
Publié: (2026)
par: Feng, Haoan, et autres
Publié: (2026)
An Invisible Backdoor Attack Based On Semantic Feature
par: Chen, Yangming
Publié: (2024)
par: Chen, Yangming
Publié: (2024)
FLoC: Facility Location-Based Efficient Visual Token Compression for Long Video Understanding
par: Cho, Janghoon, et autres
Publié: (2025)
par: Cho, Janghoon, et autres
Publié: (2025)
Physics-Guided Motion Loss for Video Generation Model
par: Xue, Bowen, et autres
Publié: (2025)
par: Xue, Bowen, et autres
Publié: (2025)
Motion-Aware Caching for Efficient Autoregressive Video Generation
par: Xu, Jing, et autres
Publié: (2026)
par: Xu, Jing, et autres
Publié: (2026)
OTT-Vid: Optimal Transport Temporal Token Compression for Video Large Language Models
par: Kang, Minseok, et autres
Publié: (2026)
par: Kang, Minseok, et autres
Publié: (2026)
InfoTok: Adaptive Discrete Video Tokenizer via Information-Theoretic Compression
par: Ye, Haotian, et autres
Publié: (2025)
par: Ye, Haotian, et autres
Publié: (2025)
TTF: Temporal Token Fusion for Efficient Video-Language Model
par: Huo, Simin, et autres
Publié: (2026)
par: Huo, Simin, et autres
Publié: (2026)
VideoMaMa: Mask-Guided Video Matting via Generative Prior
par: Lim, Sangbeom, et autres
Publié: (2026)
par: Lim, Sangbeom, et autres
Publié: (2026)
Locatability-Guided Adaptive Reasoning for Image Geo-Localization with Vision-Language Models
par: Yu, Bo, et autres
Publié: (2026)
par: Yu, Bo, et autres
Publié: (2026)
Hybrid-Level Instruction Injection for Video Token Compression in Multi-modal Large Language Models
par: Liu, Zhihang, et autres
Publié: (2025)
par: Liu, Zhihang, et autres
Publié: (2025)
HybridToken-VLM: Hybrid Token Compression for Vision-Language Models
par: Zhang, Jusheng, et autres
Publié: (2025)
par: Zhang, Jusheng, et autres
Publié: (2025)
Mask and Compress: Efficient Skeleton-based Action Recognition in Continual Learning
par: Mosconi, Matteo, et autres
Publié: (2024)
par: Mosconi, Matteo, et autres
Publié: (2024)
ImgCoT: Compressing Long Chain of Thought into Compact Visual Tokens for Efficient Reasoning of Large Language Model
par: Chen, Xiaoshu, et autres
Publié: (2026)
par: Chen, Xiaoshu, et autres
Publié: (2026)
SNED: Superposition Network Architecture Search for Efficient Video Diffusion Model
par: Li, Zhengang, et autres
Publié: (2024)
par: Li, Zhengang, et autres
Publié: (2024)
Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation
par: Yariv, Guy, et autres
Publié: (2025)
par: Yariv, Guy, et autres
Publié: (2025)
Token Sequence Compression for Efficient Multimodal Computing
par: Omri, Yasmine, et autres
Publié: (2025)
par: Omri, Yasmine, et autres
Publié: (2025)
Advanced Sign Language Video Generation with Compressed and Quantized Multi-Condition Tokenization
par: Wang, Cong, et autres
Publié: (2025)
par: Wang, Cong, et autres
Publié: (2025)
Efficient Video Diffusion with Sparse Information Transmission for Video Compression
par: Zhou, Mingde, et autres
Publié: (2026)
par: Zhou, Mingde, et autres
Publié: (2026)
LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
par: Zhang, Shaolei, et autres
Publié: (2025)
par: Zhang, Shaolei, et autres
Publié: (2025)
StreamingAssistant: Efficient Visual Token Pruning for Accelerating Online Video Understanding
par: Jin, Xinqi, et autres
Publié: (2025)
par: Jin, Xinqi, et autres
Publié: (2025)
Sharp Eyes and Memory for VideoLLMs: Information-Aware Visual Token Pruning for Efficient and Reliable VideoLLM Reasoning
par: Qin, Jialong, et autres
Publié: (2025)
par: Qin, Jialong, et autres
Publié: (2025)
Efficient Training for Human Video Generation with Entropy-Guided Prioritized Progressive Learning
par: Li, Changlin, et autres
Publié: (2025)
par: Li, Changlin, et autres
Publié: (2025)
DC-VideoGen: Efficient Video Generation with Deep Compression Video Autoencoder
par: Chen, Junyu, et autres
Publié: (2025)
par: Chen, Junyu, et autres
Publié: (2025)
AnyMo: Scaling Any-Modality Conditional Motion Generation with Masked Modeling
par: Li, Yiheng, et autres
Publié: (2026)
par: Li, Yiheng, et autres
Publié: (2026)
Physics-Grounded Motion Forecasting via Equation Discovery for Trajectory-Guided Image-to-Video Generation
par: Feng, Tao, et autres
Publié: (2025)
par: Feng, Tao, et autres
Publié: (2025)
Masked Autoencoders Are Effective Tokenizers for Diffusion Models
par: Chen, Hao, et autres
Publié: (2025)
par: Chen, Hao, et autres
Publié: (2025)
VideoMAR: Autoregressive Video Generatio with Continuous Tokens
par: Yu, Hu, et autres
Publié: (2025)
par: Yu, Hu, et autres
Publié: (2025)
FlowCoMotion: Text-to-Motion Generation via Token-Latent Flow Modeling
par: Guan, Dawei, et autres
Publié: (2026)
par: Guan, Dawei, et autres
Publié: (2026)
AMG: Avatar Motion Guided Video Generation
par: Yang, Zhangsihao, et autres
Publié: (2024)
par: Yang, Zhangsihao, et autres
Publié: (2024)
Documents similaires
-
Resource-Efficient Motion Control for Video Generation via Dynamic Mask Guidance
par: Feng, Sicong, et autres
Publié: (2025) -
Fre-Res: Frequency-Residual Video Token Compression for Efficient Video MLLMs
par: Feng, Yigui, et autres
Publié: (2026) -
DC-AR: Efficient Masked Autoregressive Image Generation with Deep Compression Hybrid Tokenizer
par: Wu, Yecheng, et autres
Publié: (2025) -
StreamingTOM: Streaming Token Compression for Efficient Video Understanding
par: Chen, Xueyi, et autres
Publié: (2025) -
QMoP: Query Guided Mixture-of-Projector for Efficient Visual Token Compression
par: Li, Zhongyang, et autres
Publié: (2026)