MIRAGE: Runtime Scheduling for Multi-Vector Image Retrieval with Hierarchical Decomposition
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Li, Maoliang, Li, Ke, Liu, Yaoyang, Chen, Jiayu, Zheng, Zihao, Wu, Yinjun, Liu, Chenchen, Chen, Xiang |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
HeRo: Adaptive Orchestration of Agentic RAG on Heterogeneous Mobile SoC
par: Li, Maoliang, et autres
Publié: (2026)
par: Li, Maoliang, et autres
Publié: (2026)
RAPID: Redundancy-Aware and Compatibility-Optimal Edge-Cloud Partitioned Inference for Diverse VLA Models
par: Zheng, Zihao, et autres
Publié: (2026)
par: Zheng, Zihao, et autres
Publié: (2026)
Agent.xpu: Efficient Scheduling of Agentic LLM Workloads on Heterogeneous SoC
par: Wei, Xinming, et autres
Publié: (2025)
par: Wei, Xinming, et autres
Publié: (2025)
C-FedRAG: A Confidential Federated Retrieval-Augmented Generation System
par: Addison, Parker, et autres
Publié: (2024)
par: Addison, Parker, et autres
Publié: (2024)
Characterizing the Dilemma of Performance and Index Size in Billion-Scale Vector Search and Breaking It with Second-Tier Memory
par: Cheng, Rongxin, et autres
Publié: (2024)
par: Cheng, Rongxin, et autres
Publié: (2024)
RoboECC: Multi-Factor-Aware Edge-Cloud Collaborative Deployment for VLA Models
par: Zheng, Zihao, et autres
Publié: (2026)
par: Zheng, Zihao, et autres
Publié: (2026)
GPU-accelerated Multi-relational Parallel Graph Retrieval for Web-scale Recommendations
par: Guo, Zhuoning, et autres
Publié: (2025)
par: Guo, Zhuoning, et autres
Publié: (2025)
CoEdge-RAG: Optimizing Hierarchical Scheduling for Retrieval-Augmented LLMs in Collaborative Edge Computing
par: Hong, Guihang, et autres
Publié: (2025)
par: Hong, Guihang, et autres
Publié: (2025)
INSPIRIT: Optimizing Heterogeneous Task Scheduling through Adaptive Priority in Task-based Runtime Systems
par: Wang, Yiqing, et autres
Publié: (2024)
par: Wang, Yiqing, et autres
Publié: (2024)
SIVF: GPU-Resident IVF Index for Streaming Vector Search
par: Zhao, Dongfang
Publié: (2026)
par: Zhao, Dongfang
Publié: (2026)
Efficient Distributed Retrieval-Augmented Generation for Enhancing Language Model Performance
par: Liu, Shangyu, et autres
Publié: (2025)
par: Liu, Shangyu, et autres
Publié: (2025)
SaberLDA: Sparsity-Aware Learning of Topic Models on GPUs
par: Li, Kaiwei, et autres
Publié: (2016)
par: Li, Kaiwei, et autres
Publié: (2016)
Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services
par: Chen, Haoyu, et autres
Publié: (2025)
par: Chen, Haoyu, et autres
Publié: (2025)
Hyperion: Hierarchical Scheduling for Parallel LLM Acceleration in Multi-tier Networks
par: Ma, Mulei, et autres
Publié: (2025)
par: Ma, Mulei, et autres
Publié: (2025)
HydraInfer: Hybrid Disaggregated Scheduling for Multimodal Large Language Model Serving
par: Dong, Xianzhe, et autres
Publié: (2025)
par: Dong, Xianzhe, et autres
Publié: (2025)
A Model-agnostic Strategy to Mitigate Embedding Degradation in Personalized Federated Recommendation
par: Shen, Jiakui, et autres
Publié: (2025)
par: Shen, Jiakui, et autres
Publié: (2025)
Exact Nearest-Neighbor Search on Energy-Efficient FPGA Devices
par: Dazzi, Patrizio, et autres
Publié: (2025)
par: Dazzi, Patrizio, et autres
Publié: (2025)
A Big Data Architecture for Early Identification and Categorization of Dark Web Sites
par: Pastor-Galindo, Javier, et autres
Publié: (2024)
par: Pastor-Galindo, Javier, et autres
Publié: (2024)
An OPC UA-based industrial Big Data architecture
par: Hirsch, Eduard, et autres
Publié: (2023)
par: Hirsch, Eduard, et autres
Publié: (2023)
Fair Kernel-Lock-Free Claim/Release Protocol for Shared Object Access in Cooperatively Scheduled Runtimes
par: Chalmers, Kevin, et autres
Publié: (2025)
par: Chalmers, Kevin, et autres
Publié: (2025)
Federated Cross-Domain Click-Through Rate Prediction With Large Language Model Augmentation
par: Qin, Jiangcheng, et autres
Publié: (2025)
par: Qin, Jiangcheng, et autres
Publié: (2025)
Fantasy: Efficient Large-scale Vector Search on GPU Clusters with GPUDirect Async
par: Liu, Yi, et autres
Publié: (2025)
par: Liu, Yi, et autres
Publié: (2025)
SageSched: Efficient LLM Scheduling Confronting Demand Uncertainty and Hybridity
par: Gan, Zhenghao, et autres
Publié: (2026)
par: Gan, Zhenghao, et autres
Publié: (2026)
Arrow: Adaptive Scheduling Mechanisms for Disaggregated LLM Inference Architecture
par: Wu, Yu, et autres
Publié: (2025)
par: Wu, Yu, et autres
Publié: (2025)
DIET: Customized Slimming for Incompatible Networks in Sequential Recommendation
par: Fu, Kairui, et autres
Publié: (2024)
par: Fu, Kairui, et autres
Publié: (2024)
Efficient Federated Search for Retrieval-Augmented Generation using Lightweight Routing
par: Dhasade, Akash, et autres
Publié: (2025)
par: Dhasade, Akash, et autres
Publié: (2025)
Scalable Distributed Vector Search via Accuracy Preserving Index Construction
par: Xu, Yuming, et autres
Publié: (2025)
par: Xu, Yuming, et autres
Publié: (2025)
HexAGenT: Efficient Agentic LLM Serving via Workflow- and Heterogeneity-Aware Scheduling
par: Peng, You, et autres
Publié: (2026)
par: Peng, You, et autres
Publié: (2026)
PROSERVE: Unified Multi-Priority Request Scheduling for LLM Serving
par: Huang, Weizhe, et autres
Publié: (2025)
par: Huang, Weizhe, et autres
Publié: (2025)
FATE: Future-State-Aware Scheduling for Heterogeneous LLM Workflows
par: Huang, Zirui, et autres
Publié: (2026)
par: Huang, Zirui, et autres
Publié: (2026)
SLO-Aware Scheduling for Large Language Model Inferences
par: Huang, Jinqi, et autres
Publié: (2025)
par: Huang, Jinqi, et autres
Publié: (2025)
RingAda: Pipelining Large Model Fine-Tuning on Edge Devices with Scheduled Layer Unfreezing
par: Li, Liang, et autres
Publié: (2025)
par: Li, Liang, et autres
Publié: (2025)
Slice-Level Scheduling for High Throughput and Load Balanced LLM Serving
par: Cheng, Ke, et autres
Publié: (2024)
par: Cheng, Ke, et autres
Publié: (2024)
Curator: Efficient Indexing for Multi-Tenant Vector Databases
par: Jin, Yicheng, et autres
Publié: (2024)
par: Jin, Yicheng, et autres
Publié: (2024)
HybridFlow: Resource-Adaptive Subtask Routing for Efficient Edge-Cloud LLM Inference
par: Dong, Jiangwen, et autres
Publié: (2025)
par: Dong, Jiangwen, et autres
Publié: (2025)
Robust Implementation of Retrieval-Augmented Generation on Edge-based Computing-in-Memory Architectures
par: Qin, Ruiyang, et autres
Publié: (2024)
par: Qin, Ruiyang, et autres
Publié: (2024)
Towards Efficient and Scalable Distributed Vector Search with RDMA
par: Zhi, Xiangyu, et autres
Publié: (2025)
par: Zhi, Xiangyu, et autres
Publié: (2025)
HP-MDR: High-performance and Portable Data Refactoring and Progressive Retrieval with Advanced GPUs
par: Li, Yanliang, et autres
Publié: (2025)
par: Li, Yanliang, et autres
Publié: (2025)
A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability
par: Liu, Ruitao, et autres
Publié: (2026)
par: Liu, Ruitao, et autres
Publié: (2026)
Schedule-Level Shared-Prefix Reuse for LLM RL Training
par: Li, Pengbo, et autres
Publié: (2026)
par: Li, Pengbo, et autres
Publié: (2026)
Documents similaires
-
HeRo: Adaptive Orchestration of Agentic RAG on Heterogeneous Mobile SoC
par: Li, Maoliang, et autres
Publié: (2026) -
RAPID: Redundancy-Aware and Compatibility-Optimal Edge-Cloud Partitioned Inference for Diverse VLA Models
par: Zheng, Zihao, et autres
Publié: (2026) -
Agent.xpu: Efficient Scheduling of Agentic LLM Workloads on Heterogeneous SoC
par: Wei, Xinming, et autres
Publié: (2025) -
C-FedRAG: A Confidential Federated Retrieval-Augmented Generation System
par: Addison, Parker, et autres
Publié: (2024) -
Characterizing the Dilemma of Performance and Index Size in Billion-Scale Vector Search and Breaking It with Second-Tier Memory
par: Cheng, Rongxin, et autres
Publié: (2024)