Venus: An Efficient Edge Memory-and-Retrieval System for VLM-based Online Video Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Ye, Shengyuan, Ouyang, Bei, Qian, Tianyi, Zeng, Liekang, Yuan, Mu, Chu, Xiaowen, Hong, Weijie, Chen, Xu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Jupiter: Fast and Resource-Efficient Collaborative Inference of Generative LLMs on Edge Devices
by: Ye, Shengyuan, et al.
Published: (2025)
by: Ye, Shengyuan, et al.
Published: (2025)
Resource-Efficient Personal Large Language Models Fine-Tuning with Collaborative Edge Computing
by: Ye, Shengyuan, et al.
Published: (2024)
by: Ye, Shengyuan, et al.
Published: (2024)
Online Optimization of DNN Inference Network Utility in Collaborative Edge Computing
by: Li, Rui, et al.
Published: (2024)
by: Li, Rui, et al.
Published: (2024)
Asteroid: Resource-Efficient Hybrid Pipeline Parallelism for Collaborative DNN Training on Heterogeneous Edge Devices
by: Ye, Shengyuan, et al.
Published: (2024)
by: Ye, Shengyuan, et al.
Published: (2024)
Galaxy: A Resource-Efficient Collaborative Edge AI System for In-situ Transformer Inference
by: Ye, Shengyuan, et al.
Published: (2024)
by: Ye, Shengyuan, et al.
Published: (2024)
Implementation of Big AI Models for Wireless Networks with Collaborative Edge Computing
by: Zeng, Liekang, et al.
Published: (2024)
by: Zeng, Liekang, et al.
Published: (2024)
CoEdge-RAG: Optimizing Hierarchical Scheduling for Retrieval-Augmented LLMs in Collaborative Edge Computing
by: Hong, Guihang, et al.
Published: (2025)
by: Hong, Guihang, et al.
Published: (2025)
PerCache: Predictive Hierarchical Cache for RAG Applications on Mobile Devices
by: Liu, Kaiwei, et al.
Published: (2025)
by: Liu, Kaiwei, et al.
Published: (2025)
Edge Graph Intelligence: Reciprocally Empowering Edge Networks with Graph Intelligence
by: Zeng, Liekang, et al.
Published: (2024)
by: Zeng, Liekang, et al.
Published: (2024)
MAP-UOT: A Memory-Efficient Approach to Unbalanced Optimal Transport Implementation
by: Sun, Chengyu, et al.
Published: (2024)
by: Sun, Chengyu, et al.
Published: (2024)
IM-PIR: In-Memory Private Information Retrieval
by: Mwaisela, Mpoki, et al.
Published: (2025)
by: Mwaisela, Mpoki, et al.
Published: (2025)
Understanding the Performance and Power of LLM Inferencing on Edge Accelerators
by: Arya, Mayank, et al.
Published: (2025)
by: Arya, Mayank, et al.
Published: (2025)
EdgeShard: Efficient LLM Inference via Collaborative Edge Computing
by: Zhang, Mingjin, et al.
Published: (2024)
by: Zhang, Mingjin, et al.
Published: (2024)
Understanding the Landscape of Ampere GPU Memory Errors
by: Zhu, Zhu, et al.
Published: (2025)
by: Zhu, Zhu, et al.
Published: (2025)
Efficient Column-Wise N:M Pruning on RISC-V CPU
by: Chu, Chi-Wei, et al.
Published: (2025)
by: Chu, Chi-Wei, et al.
Published: (2025)
Ponder: Online Prediction of Task Memory Requirements for Scientific Workflows
by: Lehmann, Fabian, et al.
Published: (2024)
by: Lehmann, Fabian, et al.
Published: (2024)
LIME:Accelerating Collaborative Lossless LLM Inference on Memory-Constrained Edge Devices
by: Sun, Mingyu, et al.
Published: (2025)
by: Sun, Mingyu, et al.
Published: (2025)
In-Vehicle Edge System for Real-Time Dashcam Video Analysis
by: Lee, Seyul, et al.
Published: (2024)
by: Lee, Seyul, et al.
Published: (2024)
OCTOPINF: Workload-Aware Inference Serving for Edge Video Analytics
by: Nguyen, Thanh-Tung, et al.
Published: (2025)
by: Nguyen, Thanh-Tung, et al.
Published: (2025)
Stochastic Modeling for Energy-Efficient Edge Infrastructure
by: Rossi, Fabio Diniz
Published: (2025)
by: Rossi, Fabio Diniz
Published: (2025)
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management
by: Qianli, Liu, et al.
Published: (2025)
by: Qianli, Liu, et al.
Published: (2025)
GoldFish: Serverless Actors with Short-Term Memory State for the Edge-Cloud Continuum
by: Marcelino, Cynthia, et al.
Published: (2024)
by: Marcelino, Cynthia, et al.
Published: (2024)
SkyMemory: A LEO Edge Cache for Transformer Inference Optimization and Scale Out
by: Sandholm, Thomas, et al.
Published: (2025)
by: Sandholm, Thomas, et al.
Published: (2025)
Accelerating AIGC Services with Latent Action Diffusion Scheduling in Edge Networks
by: Xu, Changfu, et al.
Published: (2024)
by: Xu, Changfu, et al.
Published: (2024)
ZipCCL: Efficient Lossless Data Compression of Communication Collectives for Accelerating LLM Training
by: Lin, Wenxiang, et al.
Published: (2026)
by: Lin, Wenxiang, et al.
Published: (2026)
PVU: Design and Implementation of a Posit Vector Arithmetic Unit (PVU) for Enhanced Floating-Point Computing in Edge and AI Applications
by: Wu, Xinyu, et al.
Published: (2025)
by: Wu, Xinyu, et al.
Published: (2025)
Sizey: Memory-Efficient Execution of Scientific Workflow Tasks
by: Bader, Jonathan, et al.
Published: (2024)
by: Bader, Jonathan, et al.
Published: (2024)
AME: An Efficient Heterogeneous Agentic Memory Engine for Smartphones
by: Zhao, Xinkui, et al.
Published: (2025)
by: Zhao, Xinkui, et al.
Published: (2025)
An Efficient and Adaptive Watermark Detection System with Tile-based Error Correction
by: Zhong, Xinrui, et al.
Published: (2025)
by: Zhong, Xinrui, et al.
Published: (2025)
POSEIDON : Efficient Function Placement at the Edge using Deep Reinforcement Learning
by: Jain, Prakhar, et al.
Published: (2024)
by: Jain, Prakhar, et al.
Published: (2024)
Parm: Efficient Training of Large Sparsely-Activated Models with Dedicated Schedules
by: Pan, Xinglin, et al.
Published: (2024)
by: Pan, Xinglin, et al.
Published: (2024)
GMLake: Efficient and Transparent GPU Memory Defragmentation for Large-scale DNN Training with Virtual Memory Stitching
by: Guo, Cong, et al.
Published: (2024)
by: Guo, Cong, et al.
Published: (2024)
Understanding Read-Write Wait-Free Coverings in the Fully-Anonymous Shared-Memory Model
by: Losa, Giuliano, et al.
Published: (2024)
by: Losa, Giuliano, et al.
Published: (2024)
FCPO: Federated Continual Policy Optimization for Real-Time High-Throughput Edge Video Analytics
by: Liebe, Lucas, et al.
Published: (2025)
by: Liebe, Lucas, et al.
Published: (2025)
ML-ECS: A Collaborative Multimodal Learning Framework for Edge-Cloud Synergies
by: Liu, Yuze, et al.
Published: (2026)
by: Liu, Yuze, et al.
Published: (2026)
Efficient and Portable Support for Overdecomposition on Distributed Memory GPGPU Platforms
by: Bhosale, Aditya, et al.
Published: (2026)
by: Bhosale, Aditya, et al.
Published: (2026)
Flash-KMeans: Fast and Memory-Efficient Exact K-Means
by: Yang, Shuo, et al.
Published: (2026)
by: Yang, Shuo, et al.
Published: (2026)
Efficient Scheduling of Vehicular Tasks on Edge Systems with Green Energy and Battery Storage
by: Sarkar, Suvarthi, et al.
Published: (2024)
by: Sarkar, Suvarthi, et al.
Published: (2024)
Failure-Resilient and Carbon-Efficient Deployment of Microservices over the Cloud-Edge Continuum
by: Ponce, Francisco, et al.
Published: (2026)
by: Ponce, Francisco, et al.
Published: (2026)
Efficient Training Approaches for Performance Anomaly Detection Models in Edge Computing Environments
by: Fernando, Duneesha, et al.
Published: (2024)
by: Fernando, Duneesha, et al.
Published: (2024)
Similar Items
-
Jupiter: Fast and Resource-Efficient Collaborative Inference of Generative LLMs on Edge Devices
by: Ye, Shengyuan, et al.
Published: (2025) -
Resource-Efficient Personal Large Language Models Fine-Tuning with Collaborative Edge Computing
by: Ye, Shengyuan, et al.
Published: (2024) -
Online Optimization of DNN Inference Network Utility in Collaborative Edge Computing
by: Li, Rui, et al.
Published: (2024) -
Asteroid: Resource-Efficient Hybrid Pipeline Parallelism for Collaborative DNN Training on Heterogeneous Edge Devices
by: Ye, Shengyuan, et al.
Published: (2024) -
Galaxy: A Resource-Efficient Collaborative Edge AI System for In-situ Transformer Inference
by: Ye, Shengyuan, et al.
Published: (2024)