AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Man, Yuanbin, Huang, Ying, Zhang, Chengming, Li, Bingzhe, Niu, Wei, Yin, Miao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AdaRing: Towards Ultra-Light Vision-Language Adaptation via Cross-Layer Tensor Ring Decomposition
von: Huang, Ying, et al.
Veröffentlicht: (2025)
von: Huang, Ying, et al.
Veröffentlicht: (2025)
Event-VStream: Event-Driven Real-Time Understanding for Long Video Streams
von: Guo, Zhenghui, et al.
Veröffentlicht: (2026)
von: Guo, Zhenghui, et al.
Veröffentlicht: (2026)
AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding
von: Li, Handong, et al.
Veröffentlicht: (2026)
von: Li, Handong, et al.
Veröffentlicht: (2026)
Turbo4DGen: Ultra-Fast Acceleration for 4D Generation
von: Man, Yuanbin, et al.
Veröffentlicht: (2026)
von: Man, Yuanbin, et al.
Veröffentlicht: (2026)
AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video Understanding
von: Xue, Zhucun, et al.
Veröffentlicht: (2025)
von: Xue, Zhucun, et al.
Veröffentlicht: (2025)
AdaReTaKe: Adaptive Redundancy Reduction to Perceive Longer for Video-language Understanding
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
Question-guided Visual Compression with Memory Feedback for Long-Term Video Understanding
von: Yamao, Sosuke, et al.
Veröffentlicht: (2026)
von: Yamao, Sosuke, et al.
Veröffentlicht: (2026)
SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration
von: Yang, Zhongyu, et al.
Veröffentlicht: (2026)
von: Yang, Zhongyu, et al.
Veröffentlicht: (2026)
LVBench: An Extreme Long Video Understanding Benchmark
von: Wang, Weihan, et al.
Veröffentlicht: (2024)
von: Wang, Weihan, et al.
Veröffentlicht: (2024)
VideoMem: Enhancing Ultra-Long Video Understanding via Adaptive Memory Management
von: Jin, Hongbo, et al.
Veröffentlicht: (2025)
von: Jin, Hongbo, et al.
Veröffentlicht: (2025)
MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding
von: He, Bo, et al.
Veröffentlicht: (2024)
von: He, Bo, et al.
Veröffentlicht: (2024)
AdaFocus: Adaptive Relevance-Diversity Sampling with Zero-Cache Look-back for Efficient Long Video Understanding
von: Yang, Xiao, et al.
Veröffentlicht: (2026)
von: Yang, Xiao, et al.
Veröffentlicht: (2026)
X-LeBench: A Benchmark for Extremely Long Egocentric Video Understanding
von: Zhou, Wenqi, et al.
Veröffentlicht: (2025)
von: Zhou, Wenqi, et al.
Veröffentlicht: (2025)
StreamMeCo: Long-Term Agent Memory Compression for Efficient Streaming Video Understanding
von: Wang, Junxi, et al.
Veröffentlicht: (2026)
von: Wang, Junxi, et al.
Veröffentlicht: (2026)
VideoPro: Adaptive Program Reasoning for Long Video Understanding
von: Li, Chenglin, et al.
Veröffentlicht: (2025)
von: Li, Chenglin, et al.
Veröffentlicht: (2025)
GCAgent: Long-Video Understanding via Schematic and Narrative Episodic Memory
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
AdaRD-key: Adaptive Relevance-Diversity Keyframe Sampling for Long-form Video understanding
von: Zhang, Xian, et al.
Veröffentlicht: (2025)
von: Zhang, Xian, et al.
Veröffentlicht: (2025)
AdaFlow: Efficient Long Video Editing via Adaptive Attention Slimming And Keyframe Selection
von: Zhang, Shuheng, et al.
Veröffentlicht: (2025)
von: Zhang, Shuheng, et al.
Veröffentlicht: (2025)
Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025)
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025)
VideoLucy: Deep Memory Backtracking for Long Video Understanding
von: Zuo, Jialong, et al.
Veröffentlicht: (2025)
von: Zuo, Jialong, et al.
Veröffentlicht: (2025)
Cross-Modal Dual-Causal Learning for Long-Term Action Recognition
von: Shaowu, Xu, et al.
Veröffentlicht: (2025)
von: Shaowu, Xu, et al.
Veröffentlicht: (2025)
EchoingPixels: Cross-Modal Adaptive Token Reduction for Efficient Audio-Visual LLMs
von: Gong, Chao, et al.
Veröffentlicht: (2025)
von: Gong, Chao, et al.
Veröffentlicht: (2025)
AdaVid: Adaptive Video-Language Pretraining
von: Patel, Chaitanya, et al.
Veröffentlicht: (2025)
von: Patel, Chaitanya, et al.
Veröffentlicht: (2025)
VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video Understanding
von: Yin, Yufei, et al.
Veröffentlicht: (2025)
von: Yin, Yufei, et al.
Veröffentlicht: (2025)
Memory Consolidation Enables Long-Context Video Understanding
von: Balažević, Ivana, et al.
Veröffentlicht: (2024)
von: Balažević, Ivana, et al.
Veröffentlicht: (2024)
AdaTooler-V: Adaptive Tool-Use for Images and Videos
von: Wang, Chaoyang, et al.
Veröffentlicht: (2025)
von: Wang, Chaoyang, et al.
Veröffentlicht: (2025)
AdaFuse-Det: Adaptive Cross-Modal Fusion of Event Cameras for Robust Object Detection in Low-Light RGB Imagery
von: Imandi, Raju, et al.
Veröffentlicht: (2026)
von: Imandi, Raju, et al.
Veröffentlicht: (2026)
Adaptive Greedy Frame Selection for Long Video Understanding
von: Huang, Yuning, et al.
Veröffentlicht: (2026)
von: Huang, Yuning, et al.
Veröffentlicht: (2026)
HiCM$^2$: Hierarchical Compact Memory Modeling for Dense Video Captioning
von: Kim, Minkuk, et al.
Veröffentlicht: (2024)
von: Kim, Minkuk, et al.
Veröffentlicht: (2024)
Memory-enhanced Retrieval Augmentation for Long Video Understanding
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
CM-MaskSD: Cross-Modality Masked Self-Distillation for Referring Image Segmentation
von: Wang, Wenxuan, et al.
Veröffentlicht: (2023)
von: Wang, Wenxuan, et al.
Veröffentlicht: (2023)
Long-VMNet: Accelerating Long-Form Video Understanding via Fixed Memory
von: Gurukar, Saket, et al.
Veröffentlicht: (2025)
von: Gurukar, Saket, et al.
Veröffentlicht: (2025)
VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding
von: He, Haichen, et al.
Veröffentlicht: (2026)
von: He, Haichen, et al.
Veröffentlicht: (2026)
VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos
von: Ren, Xubin, et al.
Veröffentlicht: (2025)
von: Ren, Xubin, et al.
Veröffentlicht: (2025)
LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
von: Shen, Xiaoqian, et al.
Veröffentlicht: (2024)
von: Shen, Xiaoqian, et al.
Veröffentlicht: (2024)
CL2CM: Improving Cross-Lingual Cross-Modal Retrieval via Cross-Lingual Knowledge Transfer
von: Wang, Yabing, et al.
Veröffentlicht: (2023)
von: Wang, Yabing, et al.
Veröffentlicht: (2023)
Do You Remember? Dense Video Captioning with Cross-Modal Memory Retrieval
von: Kim, Minkuk, et al.
Veröffentlicht: (2024)
von: Kim, Minkuk, et al.
Veröffentlicht: (2024)
AdaOcc: Adaptive-Resolution Occupancy Prediction
von: Chen, Chao, et al.
Veröffentlicht: (2024)
von: Chen, Chao, et al.
Veröffentlicht: (2024)
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval
von: Huang, Jinghao, et al.
Veröffentlicht: (2025)
von: Huang, Jinghao, et al.
Veröffentlicht: (2025)
Encoder-Decoder Based Long Short-Term Memory (LSTM) Model for Video Captioning
von: Adewale, Sikiru, et al.
Veröffentlicht: (2023)
von: Adewale, Sikiru, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
AdaRing: Towards Ultra-Light Vision-Language Adaptation via Cross-Layer Tensor Ring Decomposition
von: Huang, Ying, et al.
Veröffentlicht: (2025) -
Event-VStream: Event-Driven Real-Time Understanding for Long Video Streams
von: Guo, Zhenghui, et al.
Veröffentlicht: (2026) -
AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding
von: Li, Handong, et al.
Veröffentlicht: (2026) -
Turbo4DGen: Ultra-Fast Acceleration for 4D Generation
von: Man, Yuanbin, et al.
Veröffentlicht: (2026) -
AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video Understanding
von: Xue, Zhucun, et al.
Veröffentlicht: (2025)