LifelongMemory: Leveraging LLMs for Answering Queries in Long-form Egocentric Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Ying, Yang, Yanlai, Ren, Mengye |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Memory Storyboard: Leveraging Temporal Segmentation for Streaming Self-Supervised Learning from Egocentric Videos
by: Yang, Yanlai, et al.
Published: (2025)
by: Yang, Yanlai, et al.
Published: (2025)
MA-EgoQA: Question Answering over Egocentric Videos from Multiple Embodied Agents
by: Kim, Kangsan, et al.
Published: (2026)
by: Kim, Kangsan, et al.
Published: (2026)
Seeking the Unfamiliar but Memorable: Conceptual Creativity as Meta-Learning
by: Ren, Mengye
Published: (2026)
by: Ren, Mengye
Published: (2026)
StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
by: Yang, Yanlai, et al.
Published: (2025)
by: Yang, Yanlai, et al.
Published: (2025)
Midway Network: Learning Representations for Recognition and Motion from Latent Dynamics
by: Hoang, Christopher, et al.
Published: (2025)
by: Hoang, Christopher, et al.
Published: (2025)
Advancing Egocentric Video Question Answering with Multimodal Large Language Models
by: Patel, Alkesh, et al.
Published: (2025)
by: Patel, Alkesh, et al.
Published: (2025)
Ego-VPA: Egocentric Video Understanding with Parameter-efficient Adaptation
by: Wu, Tz-Ying, et al.
Published: (2024)
by: Wu, Tz-Ying, et al.
Published: (2024)
Integrating Present and Past in Unsupervised Continual Learning
by: Zhang, Yipeng, et al.
Published: (2024)
by: Zhang, Yipeng, et al.
Published: (2024)
MM-Ego: Towards Building Egocentric Multimodal LLMs for Video QA
by: Ye, Hanrong, et al.
Published: (2024)
by: Ye, Hanrong, et al.
Published: (2024)
Grounded Question-Answering in Long Egocentric Videos
by: Di, Shangzhe, et al.
Published: (2023)
by: Di, Shangzhe, et al.
Published: (2023)
Opinion: Learning Intuitive Physics May Require More than Visual Data
by: Su, Ellen, et al.
Published: (2025)
by: Su, Ellen, et al.
Published: (2025)
Lifelong Learning of Video Diffusion Models From a Single Video Stream
by: Yoo, Jason, et al.
Published: (2024)
by: Yoo, Jason, et al.
Published: (2024)
Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding
by: Nagrani, Arsha, et al.
Published: (2026)
by: Nagrani, Arsha, et al.
Published: (2026)
PooDLe: Pooled and dense self-supervised learning from naturalistic videos
by: Wang, Alex N., et al.
Published: (2024)
by: Wang, Alex N., et al.
Published: (2024)
CinePile: A Long Video Question Answering Dataset and Benchmark
by: Rawal, Ruchit, et al.
Published: (2024)
by: Rawal, Ruchit, et al.
Published: (2024)
A Backpack Full of Skills: Egocentric Video Understanding with Diverse Task Perspectives
by: Peirone, Simone Alberto, et al.
Published: (2024)
by: Peirone, Simone Alberto, et al.
Published: (2024)
Towards Lifelong Few-Shot Customization of Text-to-Image Diffusion
by: Song, Nan, et al.
Published: (2024)
by: Song, Nan, et al.
Published: (2024)
$\infty$-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation
by: Santos, Saul, et al.
Published: (2025)
by: Santos, Saul, et al.
Published: (2025)
Ego4OOD: Rethinking Egocentric Video Domain Generalization via Covariate Shift Scoring
by: Vaseqi, Zahra, et al.
Published: (2026)
by: Vaseqi, Zahra, et al.
Published: (2026)
Understanding-Enhanced Model Collaboration for Long-Tailed Egocentric Mistake Detection
by: Han, Boyu, et al.
Published: (2026)
by: Han, Boyu, et al.
Published: (2026)
RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language
by: Biswas, Subrata, et al.
Published: (2025)
by: Biswas, Subrata, et al.
Published: (2025)
Query-Conditioned Evidential Keyframe Sampling for MLLM-Based Long-Form Video Understanding
by: Wang, Yiheng, et al.
Published: (2026)
by: Wang, Yiheng, et al.
Published: (2026)
Dual-Stage Reweighted MoE for Long-Tailed Egocentric Mistake Detection
by: Han, Boyu, et al.
Published: (2025)
by: Han, Boyu, et al.
Published: (2025)
EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video
by: Hoque, Ryan, et al.
Published: (2025)
by: Hoque, Ryan, et al.
Published: (2025)
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
by: Li, Xinhao, et al.
Published: (2024)
by: Li, Xinhao, et al.
Published: (2024)
Task-Aware Multi-Expert Architecture For Lifelong Deep Learning
by: Wang, Jianyu, et al.
Published: (2025)
by: Wang, Jianyu, et al.
Published: (2025)
Towards Lifelong Aerial Autonomy: Geometric Memory Management for Continual Visual Place Recognition in Dynamic Environments
by: Shao, Xingyu, et al.
Published: (2026)
by: Shao, Xingyu, et al.
Published: (2026)
LightCache: Memory-Efficient, Training-Free Acceleration for Video Generation
by: Xiao, Yang, et al.
Published: (2025)
by: Xiao, Yang, et al.
Published: (2025)
LongViTU: Instruction Tuning for Long-Form Video Understanding
by: Wu, Rujie, et al.
Published: (2025)
by: Wu, Rujie, et al.
Published: (2025)
3D-Aware Instance Segmentation and Tracking in Egocentric Videos
by: Bhalgat, Yash, et al.
Published: (2024)
by: Bhalgat, Yash, et al.
Published: (2024)
Improving Video Question Answering through query-based frame selection
by: Patil, Himanshu, et al.
Published: (2026)
by: Patil, Himanshu, et al.
Published: (2026)
Context-Aware Aerial Object Detection: Leveraging Inter-Object and Background Relationships
by: Ren, Botao, et al.
Published: (2024)
by: Ren, Botao, et al.
Published: (2024)
Perception First: A Frontier Native-Video Model with Self-Consistency for Implicit Video Question Answering
by: Alavi, Ali
Published: (2026)
by: Alavi, Ali
Published: (2026)
Graph2Video: Leveraging Video Models to Model Dynamic Graph Evolution
by: Liu, Hua, et al.
Published: (2026)
by: Liu, Hua, et al.
Published: (2026)
Divide, then Ground: Adapting Frame Selection to Query Types for Long-Form Video Understanding
by: Li, Jialuo, et al.
Published: (2025)
by: Li, Jialuo, et al.
Published: (2025)
Leveraging Data to Say No: Memory Augmented Plug-and-Play Selective Prediction
by: Sarkar, Aditya, et al.
Published: (2026)
by: Sarkar, Aditya, et al.
Published: (2026)
NoisyEQA: Benchmarking Embodied Question Answering Against Noisy Queries
by: Wu, Tao, et al.
Published: (2024)
by: Wu, Tao, et al.
Published: (2024)
Efficient Lifelong Model Evaluation in an Era of Rapid Progress
by: Prabhu, Ameya, et al.
Published: (2024)
by: Prabhu, Ameya, et al.
Published: (2024)
Being-H0.7: A Latent World-Action Model from Egocentric Videos
by: Luo, Hao, et al.
Published: (2026)
by: Luo, Hao, et al.
Published: (2026)
What to Do Next? Memorizing skills from Egocentric Instructional Video
by: Bi, Jing, et al.
Published: (2025)
by: Bi, Jing, et al.
Published: (2025)
Similar Items
-
Memory Storyboard: Leveraging Temporal Segmentation for Streaming Self-Supervised Learning from Egocentric Videos
by: Yang, Yanlai, et al.
Published: (2025) -
MA-EgoQA: Question Answering over Egocentric Videos from Multiple Embodied Agents
by: Kim, Kangsan, et al.
Published: (2026) -
Seeking the Unfamiliar but Memorable: Conceptual Creativity as Meta-Learning
by: Ren, Mengye
Published: (2026) -
StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
by: Yang, Yanlai, et al.
Published: (2025) -
Midway Network: Learning Representations for Recognition and Motion from Latent Dynamics
by: Hoang, Christopher, et al.
Published: (2025)