Motion Guided Token Compression for Efficient Masked Video Modeling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Feng, Yukun, Shi, Yangming, Liu, Fengze, Yan, Tan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Resource-Efficient Motion Control for Video Generation via Dynamic Mask Guidance
von: Feng, Sicong, et al.
Veröffentlicht: (2025)
von: Feng, Sicong, et al.
Veröffentlicht: (2025)
Fre-Res: Frequency-Residual Video Token Compression for Efficient Video MLLMs
von: Feng, Yigui, et al.
Veröffentlicht: (2026)
von: Feng, Yigui, et al.
Veröffentlicht: (2026)
DC-AR: Efficient Masked Autoregressive Image Generation with Deep Compression Hybrid Tokenizer
von: Wu, Yecheng, et al.
Veröffentlicht: (2025)
von: Wu, Yecheng, et al.
Veröffentlicht: (2025)
StreamingTOM: Streaming Token Compression for Efficient Video Understanding
von: Chen, Xueyi, et al.
Veröffentlicht: (2025)
von: Chen, Xueyi, et al.
Veröffentlicht: (2025)
QMoP: Query Guided Mixture-of-Projector for Efficient Visual Token Compression
von: Li, Zhongyang, et al.
Veröffentlicht: (2026)
von: Li, Zhongyang, et al.
Veröffentlicht: (2026)
Contribution-aware Token Compression for Efficient Video Understanding via Reinforcement Learning
von: Ma, Yinchao, et al.
Veröffentlicht: (2026)
von: Ma, Yinchao, et al.
Veröffentlicht: (2026)
Factorized-Dreamer: Training A High-Quality Video Generator with Limited and Low-Quality Data
von: Yang, Tao, et al.
Veröffentlicht: (2024)
von: Yang, Tao, et al.
Veröffentlicht: (2024)
TokenMotion: Decoupled Motion Control via Token Disentanglement for Human-centric Video Generation
von: Li, Ruineng, et al.
Veröffentlicht: (2025)
von: Li, Ruineng, et al.
Veröffentlicht: (2025)
FCoT-VL:Advancing Text-oriented Large Vision-Language Models with Efficient Visual Token Compression
von: Li, Jianjian, et al.
Veröffentlicht: (2025)
von: Li, Jianjian, et al.
Veröffentlicht: (2025)
PhyVLLM: Physics-Guided Video Language Model with Motion-Appearance Disentanglement
von: Zhan, Yu-Wei, et al.
Veröffentlicht: (2025)
von: Zhan, Yu-Wei, et al.
Veröffentlicht: (2025)
Geometry-Guided Camera Motion Understanding in VideoLLMs
von: Feng, Haoan, et al.
Veröffentlicht: (2026)
von: Feng, Haoan, et al.
Veröffentlicht: (2026)
An Invisible Backdoor Attack Based On Semantic Feature
von: Chen, Yangming
Veröffentlicht: (2024)
von: Chen, Yangming
Veröffentlicht: (2024)
FLoC: Facility Location-Based Efficient Visual Token Compression for Long Video Understanding
von: Cho, Janghoon, et al.
Veröffentlicht: (2025)
von: Cho, Janghoon, et al.
Veröffentlicht: (2025)
Physics-Guided Motion Loss for Video Generation Model
von: Xue, Bowen, et al.
Veröffentlicht: (2025)
von: Xue, Bowen, et al.
Veröffentlicht: (2025)
Motion-Aware Caching for Efficient Autoregressive Video Generation
von: Xu, Jing, et al.
Veröffentlicht: (2026)
von: Xu, Jing, et al.
Veröffentlicht: (2026)
OTT-Vid: Optimal Transport Temporal Token Compression for Video Large Language Models
von: Kang, Minseok, et al.
Veröffentlicht: (2026)
von: Kang, Minseok, et al.
Veröffentlicht: (2026)
InfoTok: Adaptive Discrete Video Tokenizer via Information-Theoretic Compression
von: Ye, Haotian, et al.
Veröffentlicht: (2025)
von: Ye, Haotian, et al.
Veröffentlicht: (2025)
TTF: Temporal Token Fusion for Efficient Video-Language Model
von: Huo, Simin, et al.
Veröffentlicht: (2026)
von: Huo, Simin, et al.
Veröffentlicht: (2026)
VideoMaMa: Mask-Guided Video Matting via Generative Prior
von: Lim, Sangbeom, et al.
Veröffentlicht: (2026)
von: Lim, Sangbeom, et al.
Veröffentlicht: (2026)
Locatability-Guided Adaptive Reasoning for Image Geo-Localization with Vision-Language Models
von: Yu, Bo, et al.
Veröffentlicht: (2026)
von: Yu, Bo, et al.
Veröffentlicht: (2026)
Hybrid-Level Instruction Injection for Video Token Compression in Multi-modal Large Language Models
von: Liu, Zhihang, et al.
Veröffentlicht: (2025)
von: Liu, Zhihang, et al.
Veröffentlicht: (2025)
HybridToken-VLM: Hybrid Token Compression for Vision-Language Models
von: Zhang, Jusheng, et al.
Veröffentlicht: (2025)
von: Zhang, Jusheng, et al.
Veröffentlicht: (2025)
Mask and Compress: Efficient Skeleton-based Action Recognition in Continual Learning
von: Mosconi, Matteo, et al.
Veröffentlicht: (2024)
von: Mosconi, Matteo, et al.
Veröffentlicht: (2024)
ImgCoT: Compressing Long Chain of Thought into Compact Visual Tokens for Efficient Reasoning of Large Language Model
von: Chen, Xiaoshu, et al.
Veröffentlicht: (2026)
von: Chen, Xiaoshu, et al.
Veröffentlicht: (2026)
SNED: Superposition Network Architecture Search for Efficient Video Diffusion Model
von: Li, Zhengang, et al.
Veröffentlicht: (2024)
von: Li, Zhengang, et al.
Veröffentlicht: (2024)
Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation
von: Yariv, Guy, et al.
Veröffentlicht: (2025)
von: Yariv, Guy, et al.
Veröffentlicht: (2025)
Token Sequence Compression for Efficient Multimodal Computing
von: Omri, Yasmine, et al.
Veröffentlicht: (2025)
von: Omri, Yasmine, et al.
Veröffentlicht: (2025)
Advanced Sign Language Video Generation with Compressed and Quantized Multi-Condition Tokenization
von: Wang, Cong, et al.
Veröffentlicht: (2025)
von: Wang, Cong, et al.
Veröffentlicht: (2025)
Efficient Video Diffusion with Sparse Information Transmission for Video Compression
von: Zhou, Mingde, et al.
Veröffentlicht: (2026)
von: Zhou, Mingde, et al.
Veröffentlicht: (2026)
LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
von: Zhang, Shaolei, et al.
Veröffentlicht: (2025)
von: Zhang, Shaolei, et al.
Veröffentlicht: (2025)
StreamingAssistant: Efficient Visual Token Pruning for Accelerating Online Video Understanding
von: Jin, Xinqi, et al.
Veröffentlicht: (2025)
von: Jin, Xinqi, et al.
Veröffentlicht: (2025)
Sharp Eyes and Memory for VideoLLMs: Information-Aware Visual Token Pruning for Efficient and Reliable VideoLLM Reasoning
von: Qin, Jialong, et al.
Veröffentlicht: (2025)
von: Qin, Jialong, et al.
Veröffentlicht: (2025)
Efficient Training for Human Video Generation with Entropy-Guided Prioritized Progressive Learning
von: Li, Changlin, et al.
Veröffentlicht: (2025)
von: Li, Changlin, et al.
Veröffentlicht: (2025)
DC-VideoGen: Efficient Video Generation with Deep Compression Video Autoencoder
von: Chen, Junyu, et al.
Veröffentlicht: (2025)
von: Chen, Junyu, et al.
Veröffentlicht: (2025)
AnyMo: Scaling Any-Modality Conditional Motion Generation with Masked Modeling
von: Li, Yiheng, et al.
Veröffentlicht: (2026)
von: Li, Yiheng, et al.
Veröffentlicht: (2026)
Physics-Grounded Motion Forecasting via Equation Discovery for Trajectory-Guided Image-to-Video Generation
von: Feng, Tao, et al.
Veröffentlicht: (2025)
von: Feng, Tao, et al.
Veröffentlicht: (2025)
Masked Autoencoders Are Effective Tokenizers for Diffusion Models
von: Chen, Hao, et al.
Veröffentlicht: (2025)
von: Chen, Hao, et al.
Veröffentlicht: (2025)
VideoMAR: Autoregressive Video Generatio with Continuous Tokens
von: Yu, Hu, et al.
Veröffentlicht: (2025)
von: Yu, Hu, et al.
Veröffentlicht: (2025)
FlowCoMotion: Text-to-Motion Generation via Token-Latent Flow Modeling
von: Guan, Dawei, et al.
Veröffentlicht: (2026)
von: Guan, Dawei, et al.
Veröffentlicht: (2026)
AMG: Avatar Motion Guided Video Generation
von: Yang, Zhangsihao, et al.
Veröffentlicht: (2024)
von: Yang, Zhangsihao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Resource-Efficient Motion Control for Video Generation via Dynamic Mask Guidance
von: Feng, Sicong, et al.
Veröffentlicht: (2025) -
Fre-Res: Frequency-Residual Video Token Compression for Efficient Video MLLMs
von: Feng, Yigui, et al.
Veröffentlicht: (2026) -
DC-AR: Efficient Masked Autoregressive Image Generation with Deep Compression Hybrid Tokenizer
von: Wu, Yecheng, et al.
Veröffentlicht: (2025) -
StreamingTOM: Streaming Token Compression for Efficient Video Understanding
von: Chen, Xueyi, et al.
Veröffentlicht: (2025) -
QMoP: Query Guided Mixture-of-Projector for Efficient Visual Token Compression
von: Li, Zhongyang, et al.
Veröffentlicht: (2026)