No Time to Waste: Squeeze Time into Channel for Mobile Video Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Zhai, Yingjie, Li, Wenshuo, Tang, Yehui, Chen, Xinghao, Wang, Yunhe |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Context-Guided Spatial Feature Reconstruction for Efficient Semantic Segmentation
by: Ni, Zhenliang, et al.
Published: (2024)
by: Ni, Zhenliang, et al.
Published: (2024)
SLAB: Efficient Transformers with Simplified Linear Attention and Progressive Re-parameterized Batch Normalization
by: Guo, Jialong, et al.
Published: (2024)
by: Guo, Jialong, et al.
Published: (2024)
TinySAM: Pushing the Envelope for Efficient Segment Anything Model
by: Shu, Han, et al.
Published: (2023)
by: Shu, Han, et al.
Published: (2023)
SJD-PAC: Accelerating Speculative Jacobi Decoding via Proactive Drafting and Adaptive Continuation
by: Kang, Jialiang, et al.
Published: (2026)
by: Kang, Jialiang, et al.
Published: (2026)
ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative Decoding
by: Kang, Jialiang, et al.
Published: (2025)
by: Kang, Jialiang, et al.
Published: (2025)
GenVidBench: A 6-Million Benchmark for AI-Generated Video Detection
by: Ni, Zhenliang, et al.
Published: (2025)
by: Ni, Zhenliang, et al.
Published: (2025)
DATE: Dynamic Absolute Time Enhancement for Long Video Understanding
by: Yuan, Chao, et al.
Published: (2025)
by: Yuan, Chao, et al.
Published: (2025)
Align-KD: Distilling Cross-Modal Alignment Knowledge for Mobile Vision-Language Model
by: Feng, Qianhan, et al.
Published: (2024)
by: Feng, Qianhan, et al.
Published: (2024)
DECO: Unleashing the Potential of ConvNets for Query-based Detection and Segmentation
by: Chen, Xinghao, et al.
Published: (2023)
by: Chen, Xinghao, et al.
Published: (2023)
Free Video-LLM: Prompt-guided Visual Perception for Efficient Training-free Video LLMs
by: Han, Kai, et al.
Published: (2024)
by: Han, Kai, et al.
Published: (2024)
Full-Stage Pseudo Label Quality Enhancement for Weakly-supervised Temporal Action Localization
by: Feng, Qianhan, et al.
Published: (2024)
by: Feng, Qianhan, et al.
Published: (2024)
GhostNetV3: Exploring the Training Strategies for Compact Models
by: Liu, Zhenhua, et al.
Published: (2024)
by: Liu, Zhenhua, et al.
Published: (2024)
GPT4Image: Large Pre-trained Models Help Vision Models Learn Better on Perception Task
by: Ding, Ning, et al.
Published: (2023)
by: Ding, Ning, et al.
Published: (2023)
Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts
by: Rang, Miao, et al.
Published: (2025)
by: Rang, Miao, et al.
Published: (2025)
Token Compensator: Altering Inference Cost of Vision Transformer without Re-Tuning
by: Jie, Shibo, et al.
Published: (2024)
by: Jie, Shibo, et al.
Published: (2024)
SAM-DiffSR: Structure-Modulated Diffusion Model for Image Super-Resolution
by: Wang, Chengcheng, et al.
Published: (2024)
by: Wang, Chengcheng, et al.
Published: (2024)
Q-Bench-Video: Benchmarking the Video Quality Understanding of LMMs
by: Zhang, Zicheng, et al.
Published: (2024)
by: Zhang, Zicheng, et al.
Published: (2024)
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams
by: Zhang, Haoji, et al.
Published: (2025)
by: Zhang, Haoji, et al.
Published: (2025)
Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams
by: Zhang, Haoji, et al.
Published: (2024)
by: Zhang, Haoji, et al.
Published: (2024)
Test-Time Temporal Sampling for Efficient MLLM Video Understanding
by: Wang, Kaibin, et al.
Published: (2025)
by: Wang, Kaibin, et al.
Published: (2025)
ScaleNet: Scaling up Pretrained Neural Networks with Incremental Parameters
by: Hao, Zhiwei, et al.
Published: (2025)
by: Hao, Zhiwei, et al.
Published: (2025)
Mobile-VideoGPT: Fast and Accurate Model for Mobile Video Understanding
by: Shaker, Abdelrahman, et al.
Published: (2025)
by: Shaker, Abdelrahman, et al.
Published: (2025)
AURA: Always-On Understanding and Real-Time Assistance via Video Streams
by: Lu, Xudong, et al.
Published: (2026)
by: Lu, Xudong, et al.
Published: (2026)
SeaFormer++: Squeeze-enhanced Axial Transformer for Mobile Visual Recognition
by: Wan, Qiang, et al.
Published: (2023)
by: Wan, Qiang, et al.
Published: (2023)
Memory-Space Visual Prompting for Efficient Vision-Language Fine-Tuning
by: Jie, Shibo, et al.
Published: (2024)
by: Jie, Shibo, et al.
Published: (2024)
Harnessing Object Grounding for Time-Sensitive Video Understanding
by: Wu, Tz-Ying, et al.
Published: (2025)
by: Wu, Tz-Ying, et al.
Published: (2025)
ANYPORTAL: Zero-Shot Consistent Video Background Replacement
by: Gao, Wenshuo, et al.
Published: (2025)
by: Gao, Wenshuo, et al.
Published: (2025)
A Survey on Transformer Compression
by: Tang, Yehui, et al.
Published: (2024)
by: Tang, Yehui, et al.
Published: (2024)
An Empirical Study of World Model Quantization
by: Fu, Zhongqian, et al.
Published: (2026)
by: Fu, Zhongqian, et al.
Published: (2026)
SqueezeMe: Mobile-Ready Distillation of Gaussian Full-Body Avatars
by: Iandola, Forrest, et al.
Published: (2024)
by: Iandola, Forrest, et al.
Published: (2024)
BulletTime: Decoupled Control of Time and Camera Pose for Video Generation
by: Wang, Yiming, et al.
Published: (2025)
by: Wang, Yiming, et al.
Published: (2025)
TEM^3-Learning: Time-Efficient Multimodal Multi-Task Learning for Advanced Assistive Driving
by: Liu, Wenzhuo, et al.
Published: (2025)
by: Liu, Wenzhuo, et al.
Published: (2025)
Revealing the Power of Post-Training for Small Language Models via Knowledge Distillation
by: Rang, Miao, et al.
Published: (2025)
by: Rang, Miao, et al.
Published: (2025)
Simple 3D Pose Features Support Human and Machine Social Scene Understanding
by: Qin, Wenshuo, et al.
Published: (2025)
by: Qin, Wenshuo, et al.
Published: (2025)
RTV-Bench: Benchmarking MLLM Continuous Perception, Understanding and Reasoning through Real-Time Video
by: Xun, Shuhang, et al.
Published: (2025)
by: Xun, Shuhang, et al.
Published: (2025)
MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling
by: Zhang, Yue, et al.
Published: (2024)
by: Zhang, Yue, et al.
Published: (2024)
Real-Time Vehicle Detection and Urban Traffic Behavior Analysis Based on UAV Traffic Videos on Mobile Devices
by: Zhu, Yuan, et al.
Published: (2024)
by: Zhu, Yuan, et al.
Published: (2024)
How Can Objects Help Video-Language Understanding?
by: Tang, Zitian, et al.
Published: (2025)
by: Tang, Zitian, et al.
Published: (2025)
Hyperspectral Adapter for Object Tracking based on Hyperspectral Video
by: Gao, Long, et al.
Published: (2025)
by: Gao, Long, et al.
Published: (2025)
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models
by: Ma, Martin Q., et al.
Published: (2026)
by: Ma, Martin Q., et al.
Published: (2026)
Similar Items
-
Context-Guided Spatial Feature Reconstruction for Efficient Semantic Segmentation
by: Ni, Zhenliang, et al.
Published: (2024) -
SLAB: Efficient Transformers with Simplified Linear Attention and Progressive Re-parameterized Batch Normalization
by: Guo, Jialong, et al.
Published: (2024) -
TinySAM: Pushing the Envelope for Efficient Segment Anything Model
by: Shu, Han, et al.
Published: (2023) -
SJD-PAC: Accelerating Speculative Jacobi Decoding via Proactive Drafting and Adaptive Continuation
by: Kang, Jialiang, et al.
Published: (2026) -
ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative Decoding
by: Kang, Jialiang, et al.
Published: (2025)