ReWind: Understanding Long Videos with Instructed Learnable Memory
Fuente:
arXiv
Saved in:
| Main Authors: | Diko, Anxhelo, Wang, Tinghuai, Swaileh, Wassim, Sun, Shiyan, Patras, Ioannis |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Semantically Guided Action Anticipation
by: Diko, Anxhelo, et al.
Published: (2024)
by: Diko, Anxhelo, et al.
Published: (2024)
Semantically Guided Representation Learning For Action Anticipation
by: Diko, Anxhelo, et al.
Published: (2024)
by: Diko, Anxhelo, et al.
Published: (2024)
ReViT: Enhancing Vision Transformers Feature Diversity with Attention Residual Connections
by: Diko, Anxhelo, et al.
Published: (2024)
by: Diko, Anxhelo, et al.
Published: (2024)
VidCtx: Context-aware Video Question Answering with Image Models
by: Goulas, Andreas, et al.
Published: (2024)
by: Goulas, Andreas, et al.
Published: (2024)
VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
by: Wang, Shihao, et al.
Published: (2025)
by: Wang, Shihao, et al.
Published: (2025)
FashionSD-X: Multimodal Fashion Garment Synthesis using Latent Diffusion
by: Singh, Abhishek Kumar, et al.
Published: (2024)
by: Singh, Abhishek Kumar, et al.
Published: (2024)
LAFS: Landmark-based Facial Self-supervised Learning for Face Recognition
by: Sun, Zhonglin, et al.
Published: (2024)
by: Sun, Zhonglin, et al.
Published: (2024)
P-TAME: Explain Any Image Classifier with Trained Perturbations
by: Ntrougkas, Mariano V., et al.
Published: (2025)
by: Ntrougkas, Mariano V., et al.
Published: (2025)
Revisiting Deepfake Detection: Chronological Continual Learning and the Limits of Generalization
by: Fontana, Federico, et al.
Published: (2025)
by: Fontana, Federico, et al.
Published: (2025)
MOAB: Multi-Modal Outer Arithmetic Block For Fusion Of Histopathological Images And Genetic Data For Brain Tumor Grading
by: Alwazzan, Omnia, et al.
Published: (2024)
by: Alwazzan, Omnia, et al.
Published: (2024)
Enhancing Long Video Understanding via Hierarchical Event-Based Memory
by: Cheng, Dingxin, et al.
Published: (2024)
by: Cheng, Dingxin, et al.
Published: (2024)
GCAgent: Long-Video Understanding via Schematic and Narrative Episodic Memory
by: Yeo, Jeong Hun, et al.
Published: (2025)
by: Yeo, Jeong Hun, et al.
Published: (2025)
Scaling the Long Video Understanding of Multimodal Large Language Models via Visual Memory Mechanism
by: Chen, Tao, et al.
Published: (2026)
by: Chen, Tao, et al.
Published: (2026)
Video-EM: Event-Centric Episodic Memory for Long-Form Video Understanding
by: Wang, Yun, et al.
Published: (2025)
by: Wang, Yun, et al.
Published: (2025)
Context Propagation from Proposals for Semantic Video Object Segmentation
by: Wang, Tinghuai
Published: (2024)
by: Wang, Tinghuai
Published: (2024)
ReL-SAR: Representation Learning for Skeleton Action Recognition with Convolutional Transformers and BYOL
by: Naimi, Safwen, et al.
Published: (2024)
by: Naimi, Safwen, et al.
Published: (2024)
Video Panels for Long Video Understanding
by: Doorenbos, Lars, et al.
Published: (2025)
by: Doorenbos, Lars, et al.
Published: (2025)
DiffusionAct: Controllable Diffusion Autoencoder for One-shot Face Reenactment
by: Bounareli, Stella, et al.
Published: (2024)
by: Bounareli, Stella, et al.
Published: (2024)
FairCoT: Enhancing Fairness in Text-to-Image Generation via Chain of Thought Reasoning with Multimodal Large Language Models
by: Sahili, Zahraa Al, et al.
Published: (2024)
by: Sahili, Zahraa Al, et al.
Published: (2024)
Visual Agentic Memory: Enabling Online Long Video Understanding via Online Indexing, Hierarchical Memory, and Agentic Retrieval
by: Li, Aiden Yiliu, et al.
Published: (2026)
by: Li, Aiden Yiliu, et al.
Published: (2026)
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction
by: Man, Yuanbin, et al.
Published: (2024)
by: Man, Yuanbin, et al.
Published: (2024)
Plan-X: Instruct Video Generation via Semantic Planning
by: Huang, Lun, et al.
Published: (2025)
by: Huang, Lun, et al.
Published: (2025)
VCA: Video Curious Agent for Long Video Understanding
by: Yang, Zeyuan, et al.
Published: (2024)
by: Yang, Zeyuan, et al.
Published: (2024)
LVBench: An Extreme Long Video Understanding Benchmark
by: Wang, Weihan, et al.
Published: (2024)
by: Wang, Weihan, et al.
Published: (2024)
Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding
by: Wu, Hang, et al.
Published: (2026)
by: Wu, Hang, et al.
Published: (2026)
FluxMem: Adaptive Hierarchical Memory for Streaming Video Understanding
by: Xie, Yiweng, et al.
Published: (2026)
by: Xie, Yiweng, et al.
Published: (2026)
One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos
by: Bai, Zechen, et al.
Published: (2024)
by: Bai, Zechen, et al.
Published: (2024)
Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos
by: Fei, Jiajun, et al.
Published: (2024)
by: Fei, Jiajun, et al.
Published: (2024)
Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video Understanding
by: Chen, Houlun, et al.
Published: (2026)
by: Chen, Houlun, et al.
Published: (2026)
Relax Forcing: Relaxed KV-Memory for Consistent Long Video Generation
by: Zhao, Zengqun, et al.
Published: (2026)
by: Zhao, Zengqun, et al.
Published: (2026)
EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding
by: Wang, Ziyang, et al.
Published: (2026)
by: Wang, Ziyang, et al.
Published: (2026)
HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding
by: Zou, Heqing, et al.
Published: (2025)
by: Zou, Heqing, et al.
Published: (2025)
ReXSonoVQA: A Video QA Benchmark for Procedure-Centric Ultrasound Understanding
by: Wang, Xucheng, et al.
Published: (2026)
by: Wang, Xucheng, et al.
Published: (2026)
ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding
by: Lu, Hao, et al.
Published: (2025)
by: Lu, Hao, et al.
Published: (2025)
ReCA: Multi-Shot Long Video Extrapolation via Recursive Context Allocation
by: Liu, Akide, et al.
Published: (2026)
by: Liu, Akide, et al.
Published: (2026)
VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG
by: Fu, Honghao, et al.
Published: (2026)
by: Fu, Honghao, et al.
Published: (2026)
Memory-Efficient Continual Learning Object Segmentation for Long Video
by: Nazemi, Amir, et al.
Published: (2023)
by: Nazemi, Amir, et al.
Published: (2023)
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding
by: Huang, De-An, et al.
Published: (2025)
by: Huang, De-An, et al.
Published: (2025)
VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding
by: Lin, Kuanwei, et al.
Published: (2026)
by: Lin, Kuanwei, et al.
Published: (2026)
VideoExplorer: Think With Videos For Agentic Long-Video Understanding
by: Yuan, Huaying, et al.
Published: (2025)
by: Yuan, Huaying, et al.
Published: (2025)
Similar Items
-
Semantically Guided Action Anticipation
by: Diko, Anxhelo, et al.
Published: (2024) -
Semantically Guided Representation Learning For Action Anticipation
by: Diko, Anxhelo, et al.
Published: (2024) -
ReViT: Enhancing Vision Transformers Feature Diversity with Attention Residual Connections
by: Diko, Anxhelo, et al.
Published: (2024) -
VidCtx: Context-aware Video Question Answering with Image Models
by: Goulas, Andreas, et al.
Published: (2024) -
VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
by: Wang, Shihao, et al.
Published: (2025)