View while Moving: Efficient Video Recognition in Long-untrimmed Videos
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tian, Ye, Yang, Mengyu, Zhang, Lanshan, Zhang, Zhizhen, Liu, Yang, Xie, Xiaohui, Que, Xirong, Wang, Wendong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AdaViPro: Region-based Adaptive Visual Prompt for Large-Scale Models Adapting
von: Yang, Mengyu, et al.
Veröffentlicht: (2024)
von: Yang, Mengyu, et al.
Veröffentlicht: (2024)
LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding
von: Qiu, Jihao, et al.
Veröffentlicht: (2026)
von: Qiu, Jihao, et al.
Veröffentlicht: (2026)
MoCrop: Training Free Motion Guided Cropping for Efficient Video Action Recognition
von: Huang, Binhua, et al.
Veröffentlicht: (2025)
von: Huang, Binhua, et al.
Veröffentlicht: (2025)
Multi-model learning by sequential reading of untrimmed videos for action recognition
von: Kamiya, Kodai, et al.
Veröffentlicht: (2024)
von: Kamiya, Kodai, et al.
Veröffentlicht: (2024)
VideoWeaver: Multimodal Multi-View Video-to-Video Transfer for Embodied Agents
von: Eskandar, George, et al.
Veröffentlicht: (2026)
von: Eskandar, George, et al.
Veröffentlicht: (2026)
Learning Compact Video Representations for Efficient Long-form Video Understanding in Large Multimodal Models
von: Chen, Yuxiao, et al.
Veröffentlicht: (2026)
von: Chen, Yuxiao, et al.
Veröffentlicht: (2026)
AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video Understanding
von: Xue, Zhucun, et al.
Veröffentlicht: (2025)
von: Xue, Zhucun, et al.
Veröffentlicht: (2025)
SAVE: Speech-Aware Video Representation Learning for Video-Text Retrieval
von: Zhao, Ruixiang, et al.
Veröffentlicht: (2026)
von: Zhao, Ruixiang, et al.
Veröffentlicht: (2026)
Motion Focus Recognition in Fast-Moving Egocentric Video
von: Hong, Si-En, et al.
Veröffentlicht: (2026)
von: Hong, Si-En, et al.
Veröffentlicht: (2026)
Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025)
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025)
VideoMerge: Towards Training-free Long Video Generation
von: Zhang, Siyang, et al.
Veröffentlicht: (2025)
von: Zhang, Siyang, et al.
Veröffentlicht: (2025)
Hybrid-Tower: Fine-grained Pseudo-query Interaction and Generation for Text-to-Video Retrieval
von: Lan, Bangxiang, et al.
Veröffentlicht: (2025)
von: Lan, Bangxiang, et al.
Veröffentlicht: (2025)
PRVR: Partially Relevant Video Retrieval
von: Chen, Xianke, et al.
Veröffentlicht: (2022)
von: Chen, Xianke, et al.
Veröffentlicht: (2022)
VideoMem: Enhancing Ultra-Long Video Understanding via Adaptive Memory Management
von: Jin, Hongbo, et al.
Veröffentlicht: (2025)
von: Jin, Hongbo, et al.
Veröffentlicht: (2025)
OmniRoam: World Wandering via Long-Horizon Panoramic Video Generation
von: Liu, Yuheng, et al.
Veröffentlicht: (2026)
von: Liu, Yuheng, et al.
Veröffentlicht: (2026)
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding
von: Zhang, Yunzhu, et al.
Veröffentlicht: (2025)
von: Zhang, Yunzhu, et al.
Veröffentlicht: (2025)
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding
von: Xu, Weili, et al.
Veröffentlicht: (2025)
von: Xu, Weili, et al.
Veröffentlicht: (2025)
SVFormer: A Direct Training Spiking Transformer for Efficient Video Action Recognition
von: Yu, Liutao, et al.
Veröffentlicht: (2024)
von: Yu, Liutao, et al.
Veröffentlicht: (2024)
VideoLucy: Deep Memory Backtracking for Long Video Understanding
von: Zuo, Jialong, et al.
Veröffentlicht: (2025)
von: Zuo, Jialong, et al.
Veröffentlicht: (2025)
Video-to-Task Learning via Motion-Guided Attention for Few-Shot Action Recognition
von: Guo, Hanyu, et al.
Veröffentlicht: (2024)
von: Guo, Hanyu, et al.
Veröffentlicht: (2024)
Wan-Move: Motion-controllable Video Generation via Latent Trajectory Guidance
von: Chu, Ruihang, et al.
Veröffentlicht: (2025)
von: Chu, Ruihang, et al.
Veröffentlicht: (2025)
Hallucination Mitigation Prompts Long-term Video Understanding
von: Sun, Yiwei, et al.
Veröffentlicht: (2024)
von: Sun, Yiwei, et al.
Veröffentlicht: (2024)
VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
VideoZoomer: Reinforcement-Learned Temporal Focusing for Long Video Reasoning
von: Ding, Yang, et al.
Veröffentlicht: (2025)
von: Ding, Yang, et al.
Veröffentlicht: (2025)
Shot-Aware Frame Sampling for Video Understanding
von: Zhao, Mengyu, et al.
Veröffentlicht: (2026)
von: Zhao, Mengyu, et al.
Veröffentlicht: (2026)
VideoTIR: Accurate Understanding for Long Videos with Efficient Tool-Integrated Reasoning
von: Gao, Zhe, et al.
Veröffentlicht: (2026)
von: Gao, Zhe, et al.
Veröffentlicht: (2026)
SWIFT: Prompt-Adaptive Memory for Efficient Interactive Long Video Generation
von: Tan, Shanwen, et al.
Veröffentlicht: (2026)
von: Tan, Shanwen, et al.
Veröffentlicht: (2026)
Highly Efficient and Unsupervised Framework for Moving Object Detection in Satellite Videos
von: Xiao, C., et al.
Veröffentlicht: (2024)
von: Xiao, C., et al.
Veröffentlicht: (2024)
OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding
von: Zhao, Ruixiang, et al.
Veröffentlicht: (2026)
von: Zhao, Ruixiang, et al.
Veröffentlicht: (2026)
Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs
von: Ventura, Lucas, et al.
Veröffentlicht: (2025)
von: Ventura, Lucas, et al.
Veröffentlicht: (2025)
FiVE: A Fine-grained Video Editing Benchmark for Evaluating Emerging Diffusion and Rectified Flow Models
von: Li, Minghan, et al.
Veröffentlicht: (2025)
von: Li, Minghan, et al.
Veröffentlicht: (2025)
Video-Infinity: Distributed Long Video Generation
von: Tan, Zhenxiong, et al.
Veröffentlicht: (2024)
von: Tan, Zhenxiong, et al.
Veröffentlicht: (2024)
VideoTetris: Towards Compositional Text-to-Video Generation
von: Tian, Ye, et al.
Veröffentlicht: (2024)
von: Tian, Ye, et al.
Veröffentlicht: (2024)
LoViT: Long Video Transformer for Surgical Phase Recognition
von: Liu, Yang, et al.
Veröffentlicht: (2023)
von: Liu, Yang, et al.
Veröffentlicht: (2023)
FiLA-Video: Spatio-Temporal Compression for Fine-Grained Long Video Understanding
von: Guo, Yanan, et al.
Veröffentlicht: (2025)
von: Guo, Yanan, et al.
Veröffentlicht: (2025)
Long Context Tuning for Video Generation
von: Guo, Yuwei, et al.
Veröffentlicht: (2025)
von: Guo, Yuwei, et al.
Veröffentlicht: (2025)
Video-in-the-Loop: Span-Grounded Long Video QA with Interleaved Reasoning
von: Wang, Chendong, et al.
Veröffentlicht: (2025)
von: Wang, Chendong, et al.
Veröffentlicht: (2025)
FOCUS: Efficient Keyframe Selection for Long Video Understanding
von: Zhu, Zirui, et al.
Veröffentlicht: (2025)
von: Zhu, Zirui, et al.
Veröffentlicht: (2025)
Learning Partially-Decorrelated Common Spaces for Ad-hoc Video Search
von: Hu, Fan, et al.
Veröffentlicht: (2025)
von: Hu, Fan, et al.
Veröffentlicht: (2025)
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams
von: Zhang, Haoji, et al.
Veröffentlicht: (2025)
von: Zhang, Haoji, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
AdaViPro: Region-based Adaptive Visual Prompt for Large-Scale Models Adapting
von: Yang, Mengyu, et al.
Veröffentlicht: (2024) -
LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding
von: Qiu, Jihao, et al.
Veröffentlicht: (2026) -
MoCrop: Training Free Motion Guided Cropping for Efficient Video Action Recognition
von: Huang, Binhua, et al.
Veröffentlicht: (2025) -
Multi-model learning by sequential reading of untrimmed videos for action recognition
von: Kamiya, Kodai, et al.
Veröffentlicht: (2024) -
VideoWeaver: Multimodal Multi-View Video-to-Video Transfer for Embodied Agents
von: Eskandar, George, et al.
Veröffentlicht: (2026)