Hybrid-RACA: Hybrid Retrieval-Augmented Composition Assistance for Real-time Text Prediction
Fuente:
arXiv
Saved in:
| Main Authors: | Xia, Menglin, Zhang, Xuchao, Couturier, Camille, Zheng, Guoqing, Rajmohan, Saravan, Ruhle, Victor |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SageServe: Optimizing LLM Serving on Cloud Data Centers with Forecast Aware Auto-Scaling
by: Jaiswal, Shashwat, et al.
Published: (2025)
by: Jaiswal, Shashwat, et al.
Published: (2025)
Distributed Retrieval-Augmented Generation
by: Xu, Chenhao, et al.
Published: (2025)
by: Xu, Chenhao, et al.
Published: (2025)
Intelligent Router for LLM Workloads: Improving Performance Through Workload-Aware Load Balancing
by: Jain, Kunal, et al.
Published: (2024)
by: Jain, Kunal, et al.
Published: (2024)
HADIS: Hybrid Adaptive Diffusion Model Serving for Efficient Text-to-Image Generation
by: Yang, Qizheng, et al.
Published: (2025)
by: Yang, Qizheng, et al.
Published: (2025)
Dependency Aware Incident Linking in Large Cloud Systems
by: Ghosh, Supriyo, et al.
Published: (2024)
by: Ghosh, Supriyo, et al.
Published: (2024)
An Empirical Study of Production Incidents in Generative AI Cloud Services
by: Yan, Haoran, et al.
Published: (2025)
by: Yan, Haoran, et al.
Published: (2025)
HybridFlow: Resource-Adaptive Subtask Routing for Efficient Edge-Cloud LLM Inference
by: Dong, Jiangwen, et al.
Published: (2025)
by: Dong, Jiangwen, et al.
Published: (2025)
Optimization of Hybrid Quantum-Classical Algorithms
by: Remme, Lian, et al.
Published: (2025)
by: Remme, Lian, et al.
Published: (2025)
Zero-Execution Retrieval-Augmented Configuration Tuning of Spark Applications
by: Suri, Raunaq, et al.
Published: (2025)
by: Suri, Raunaq, et al.
Published: (2025)
A Cloud-Based Spatio-Temporal GNN-Transformer Hybrid Model for Traffic Flow Forecasting with External Feature Integration
by: Zheng, Zhuo, et al.
Published: (2025)
by: Zheng, Zhuo, et al.
Published: (2025)
Orchestrating the Execution of Serverless Functions in Hybrid Clouds
by: Peri, Aristotelis, et al.
Published: (2024)
by: Peri, Aristotelis, et al.
Published: (2024)
Reliable Communication in Hybrid Authentication and Trust Models
by: Chotkan, Rowdy, et al.
Published: (2024)
by: Chotkan, Rowdy, et al.
Published: (2024)
ElasWave: An Elastic-Native System for Scalable Hybrid-Parallel Training
by: Kang, Xueze, et al.
Published: (2025)
by: Kang, Xueze, et al.
Published: (2025)
Workload Intelligence: Punching Holes Through the Cloud Abstraction
by: Huang, Lexiang, et al.
Published: (2024)
by: Huang, Lexiang, et al.
Published: (2024)
Advancing Hybrid Defense for Byzantine Attacks in Federated Learning
by: Yue, Kai, et al.
Published: (2024)
by: Yue, Kai, et al.
Published: (2024)
In Situ In Transit Hybrid Analysis with Catalyst-ADIOS2
by: Mazen, François, et al.
Published: (2024)
by: Mazen, François, et al.
Published: (2024)
A Fault Tolerance Mechanism for Hybrid Scientific Workflows
by: Mulone, Alberto, et al.
Published: (2024)
by: Mulone, Alberto, et al.
Published: (2024)
HyProv: Hybrid Provenance Management for Scientific Workflows
by: Bountris, Vasilis, et al.
Published: (2025)
by: Bountris, Vasilis, et al.
Published: (2025)
Elastic Data Transfer Optimization with Hybrid Reinforcement Learning
by: Swargo, Rasman Mubtasim, et al.
Published: (2025)
by: Swargo, Rasman Mubtasim, et al.
Published: (2025)
SLURM Heterogeneous Jobs for Hybrid Classical-Quantum Workflows
by: Esposito, Aniello, et al.
Published: (2025)
by: Esposito, Aniello, et al.
Published: (2025)
Efficient LLM Inference with Activation Checkpointing and Hybrid Caching
by: Lee, Sanghyeon, et al.
Published: (2025)
by: Lee, Sanghyeon, et al.
Published: (2025)
RHAPSODY: Execution of Hybrid AI-HPC Workflows at Scale
by: Alsaadi, Aymen, et al.
Published: (2025)
by: Alsaadi, Aymen, et al.
Published: (2025)
Adaptra: Straggler-Resilient Hybrid-Parallel Training with Pipeline Adaptation
by: Wu, Tianyuan, et al.
Published: (2025)
by: Wu, Tianyuan, et al.
Published: (2025)
Review of Hybrid Load Balancing Algorithms in Cloud Computing Environment
by: Ijeoma, Chukwuneke Chiamaka, et al.
Published: (2022)
by: Ijeoma, Chukwuneke Chiamaka, et al.
Published: (2022)
HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference
by: Lin, Haoran, et al.
Published: (2025)
by: Lin, Haoran, et al.
Published: (2025)
Hybrid Cloud Architectures for Research Computing: Applications and Use Cases
by: Stiensmeier, Xaver, et al.
Published: (2026)
by: Stiensmeier, Xaver, et al.
Published: (2026)
Edge Intelligence in Satellite-Terrestrial Networks with Hybrid Quantum Computing
by: Huang, Siyue, et al.
Published: (2024)
by: Huang, Siyue, et al.
Published: (2024)
Zipage: Maintain High Request Concurrency for LLM Reasoning through Compressed PagedAttention
by: Liao, Mengqi, et al.
Published: (2026)
by: Liao, Mengqi, et al.
Published: (2026)
CoEdge-RAG: Optimizing Hierarchical Scheduling for Retrieval-Augmented LLMs in Collaborative Edge Computing
by: Hong, Guihang, et al.
Published: (2025)
by: Hong, Guihang, et al.
Published: (2025)
MPI-Q: A Message Communication Library for Large-Scale Classical-Quantum Heterogeneous Hybrid Distributed Computing
by: Wang, Feng, et al.
Published: (2026)
by: Wang, Feng, et al.
Published: (2026)
ResiHP: Taming LLM Training Failures with Dynamic Hybrid Parallelism
by: Ma, Tenghui, et al.
Published: (2026)
by: Ma, Tenghui, et al.
Published: (2026)
Wave-Based Dispatch for Circuit Cutting in Hybrid HPC--Quantum Systems
by: García-Raigada, Ricard S., et al.
Published: (2026)
by: García-Raigada, Ricard S., et al.
Published: (2026)
Justin: Hybrid CPU/Memory Elastic Scaling for Distributed Stream Processing
by: Schmitz, Donatien, et al.
Published: (2025)
by: Schmitz, Donatien, et al.
Published: (2025)
SageSched: Efficient LLM Scheduling Confronting Demand Uncertainty and Hybridity
by: Gan, Zhenghao, et al.
Published: (2026)
by: Gan, Zhenghao, et al.
Published: (2026)
Strategies for Molecular Dynamics using Hybrid Systems: LAMMPS Use Case
by: Ramalho, Paulo Henrique Leme, et al.
Published: (2026)
by: Ramalho, Paulo Henrique Leme, et al.
Published: (2026)
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection
by: Zhou, Yuhang, et al.
Published: (2025)
by: Zhou, Yuhang, et al.
Published: (2025)
FRESCO: Fast and Reliable Edge Offloading with Reputation-based Hybrid Smart Contracts
by: Zilic, Josip, et al.
Published: (2024)
by: Zilic, Josip, et al.
Published: (2024)
HydraInfer: Hybrid Disaggregated Scheduling for Multimodal Large Language Model Serving
by: Dong, Xianzhe, et al.
Published: (2025)
by: Dong, Xianzhe, et al.
Published: (2025)
A Hybrid Communication Approach for Metadata Exchange in Geo-Distributed Fog Environments
by: Kruber, Marvin, et al.
Published: (2023)
by: Kruber, Marvin, et al.
Published: (2023)
Octopus: Experiences with a Hybrid Event-Driven Architecture for Distributed Scientific Computing
by: Pan, Haochen, et al.
Published: (2024)
by: Pan, Haochen, et al.
Published: (2024)
Similar Items
-
SageServe: Optimizing LLM Serving on Cloud Data Centers with Forecast Aware Auto-Scaling
by: Jaiswal, Shashwat, et al.
Published: (2025) -
Distributed Retrieval-Augmented Generation
by: Xu, Chenhao, et al.
Published: (2025) -
Intelligent Router for LLM Workloads: Improving Performance Through Workload-Aware Load Balancing
by: Jain, Kunal, et al.
Published: (2024) -
HADIS: Hybrid Adaptive Diffusion Model Serving for Efficient Text-to-Image Generation
by: Yang, Qizheng, et al.
Published: (2025) -
Dependency Aware Incident Linking in Large Cloud Systems
by: Ghosh, Supriyo, et al.
Published: (2024)