HA-RAG: Hotness-Aware RAG Acceleration via Mixed Precision and Data Placement
Fuente:
arXiv
Guardado en:
| Autores principales: | Ge, Danying, Gao, Jianhua, Yang, Yixue, Ji, Weixing |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Toward Storage-Aware Learning with Compressed Data An Empirical Exploratory Study on JPEG
por: Lee, Kichang, et al.
Publicado: (2025)
por: Lee, Kichang, et al.
Publicado: (2025)
Understanding The Effectiveness of Lossy Compression in Machine Learning Training Sets
por: Underwood, Robert, et al.
Publicado: (2024)
por: Underwood, Robert, et al.
Publicado: (2024)
Rotation Invariant Quantization for Model Compression
por: Kampeas, Joseph, et al.
Publicado: (2023)
por: Kampeas, Joseph, et al.
Publicado: (2023)
FedHypeVAE: Federated Learning with Hypernetwork Generated Conditional VAEs for Differentially Private Embedding Sharing
por: Gupta, Sunny, et al.
Publicado: (2026)
por: Gupta, Sunny, et al.
Publicado: (2026)
Tensor Generalized Approximate Message Passing
por: Li, Yinchuan, et al.
Publicado: (2025)
por: Li, Yinchuan, et al.
Publicado: (2025)
QPMeL - Quantum-Aware Classically-Trained Embeddings via Projective Metric Learning
por: Sharma, Vinayak, et al.
Publicado: (2023)
por: Sharma, Vinayak, et al.
Publicado: (2023)
KGiRAG: An Iterative GraphRAG Approach for Responding Sensemaking Queries
por: Iacob, Isabela, et al.
Publicado: (2026)
por: Iacob, Isabela, et al.
Publicado: (2026)
Interpretable classifiers for tabular data via discretization and feature selection
por: Jaakkola, Reijo, et al.
Publicado: (2024)
por: Jaakkola, Reijo, et al.
Publicado: (2024)
GOT-Edit: Geometry-Aware Generic Object Tracking via Online Model Editing
por: Chen, Shih-Fang, et al.
Publicado: (2026)
por: Chen, Shih-Fang, et al.
Publicado: (2026)
Predictive Modeling of I/O Performance for Machine Learning Training Pipelines: A Data-Driven Approach to Storage Optimization
por: Prabhakar, Karthik, et al.
Publicado: (2025)
por: Prabhakar, Karthik, et al.
Publicado: (2025)
GeoGPT-RAG Technical Report
por: Huang, Fei, et al.
Publicado: (2025)
por: Huang, Fei, et al.
Publicado: (2025)
OG-RAG: Ontology-Grounded Retrieval-Augmented Generation For Large Language Models
por: Sharma, Kartik, et al.
Publicado: (2024)
por: Sharma, Kartik, et al.
Publicado: (2024)
Scalable Inference Architectures for Compound AI Systems: A Production Deployment Study
por: S V, Srikanta Prasad, et al.
Publicado: (2026)
por: S V, Srikanta Prasad, et al.
Publicado: (2026)
ROZA Graphs: Self-Improving Near-Deterministic RAG through Evidence-Centric Feedback
por: Penaroza, Matthew
Publicado: (2026)
por: Penaroza, Matthew
Publicado: (2026)
Agent Memory Below the Prompt: Persistent Q4 KV Cache for Multi-Agent LLM Inference on Edge Devices
por: Shkolnikov, Yakov Pyotr
Publicado: (2026)
por: Shkolnikov, Yakov Pyotr
Publicado: (2026)
Accelerating Cerebral Diagnostics with BrainFusion: A Comprehensive MRI Tumor Framework
por: Houmaidi, Walid, et al.
Publicado: (2025)
por: Houmaidi, Walid, et al.
Publicado: (2025)
WaveMixSR-V2: Enhancing Super-resolution with Higher Efficiency
por: Jeevan, Pranav, et al.
Publicado: (2024)
por: Jeevan, Pranav, et al.
Publicado: (2024)
Byzantine-Resilient Federated Learning via QUBO-Based Client Selection on Quantum Annealers
por: Ferenczi, Andras, et al.
Publicado: (2026)
por: Ferenczi, Andras, et al.
Publicado: (2026)
$\mathsf{OPA}$: One-shot Private Aggregation with Single Client Interaction and its Applications to Federated Learning
por: Karthikeyan, Harish, et al.
Publicado: (2024)
por: Karthikeyan, Harish, et al.
Publicado: (2024)
WaveMix: A Resource-efficient Neural Network for Image Analysis
por: Jeevan, Pranav, et al.
Publicado: (2022)
por: Jeevan, Pranav, et al.
Publicado: (2022)
GREEN-CODE: Learning to Optimize Energy Efficiency in LLM-based Code Generation
por: Ilager, Shashikant, et al.
Publicado: (2025)
por: Ilager, Shashikant, et al.
Publicado: (2025)
SemanticAgent: A Semantics-Aware Framework for Text-to-SQL Data Synthesis
por: Gao, Qiang, et al.
Publicado: (2026)
por: Gao, Qiang, et al.
Publicado: (2026)
Accelerating Large Language Models through Partially Linear Feed-Forward Network
por: Hu, Gansen, et al.
Publicado: (2025)
por: Hu, Gansen, et al.
Publicado: (2025)
UnWeaving the knots of GraphRAG -- turns out VectorRAG is almost enough
por: Tuora, Ryszard, et al.
Publicado: (2026)
por: Tuora, Ryszard, et al.
Publicado: (2026)
Inductive Link Prediction on N-ary Relational Facts via Semantic Hypergraph Reasoning
por: Yin, Gongzhu, et al.
Publicado: (2025)
por: Yin, Gongzhu, et al.
Publicado: (2025)
Hebbian Memory-Augmented Recurrent Networks: Engram Neurons in Deep Learning
por: Szelogowski, Daniel
Publicado: (2025)
por: Szelogowski, Daniel
Publicado: (2025)
Engram Memory Encoding and Retrieval: A Neurocomputational Perspective
por: Szelogowski, Daniel
Publicado: (2025)
por: Szelogowski, Daniel
Publicado: (2025)
GenCeption: Evaluate Vision LLMs with Unlabeled Unimodal Data
por: Cao, Lele, et al.
Publicado: (2024)
por: Cao, Lele, et al.
Publicado: (2024)
Generative Modeling in Protein Design: Neural Representations, Conditional Generation, and Evaluation Standards
por: Wanasekara, Senura Hansaja, et al.
Publicado: (2026)
por: Wanasekara, Senura Hansaja, et al.
Publicado: (2026)
FED-FSTQ: Fisher-Guided Token Quantization for Communication-Efficient Federated Fine-Tuning of LLMs on Edge Devices
por: Li, Changyu, et al.
Publicado: (2026)
por: Li, Changyu, et al.
Publicado: (2026)
Phase-Aware Wavelet-Based-Scattering Encoder-Decoder for Dense Predictions
por: Marrakchi, Ghassen, et al.
Publicado: (2026)
por: Marrakchi, Ghassen, et al.
Publicado: (2026)
MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices
por: Shakerdargah, Mohammadali, et al.
Publicado: (2024)
por: Shakerdargah, Mohammadali, et al.
Publicado: (2024)
PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models
por: Friedland, Gerald, et al.
Publicado: (2024)
por: Friedland, Gerald, et al.
Publicado: (2024)
SLA Management in Reconfigurable Multi-Agent RAG: A Systems Approach to Question Answering
por: Iannelli, Michael, et al.
Publicado: (2024)
por: Iannelli, Michael, et al.
Publicado: (2024)
Kunlun Anomaly Troubleshooter: Enabling Kernel-Level Anomaly Detection and Causal Reasoning for Large Model Distributed Inference
por: Liu, Yuyang, et al.
Publicado: (2025)
por: Liu, Yuyang, et al.
Publicado: (2025)
Evaluating Temporal Semantic Caching and Workflow Optimization in Agentic Plan-Execute Pipelines
por: Merchant, Alimurtaza Mustafa, et al.
Publicado: (2026)
por: Merchant, Alimurtaza Mustafa, et al.
Publicado: (2026)
Vertical Federated Image Segmentation
por: Mandal, Paul K., et al.
Publicado: (2024)
por: Mandal, Paul K., et al.
Publicado: (2024)
Horizontal Federated Computer Vision
por: Mandal, Paul K., et al.
Publicado: (2023)
por: Mandal, Paul K., et al.
Publicado: (2023)
PEM: Perception Error Model for Virtual Testing of Autonomous Vehicles
por: Piazzoni, Andrea, et al.
Publicado: (2023)
por: Piazzoni, Andrea, et al.
Publicado: (2023)
Cold-RL: Learning Cache Eviction with Offline Reinforcement Learning for NGINX
por: Gupta, Aayush, et al.
Publicado: (2025)
por: Gupta, Aayush, et al.
Publicado: (2025)
Ejemplares similares
-
Toward Storage-Aware Learning with Compressed Data An Empirical Exploratory Study on JPEG
por: Lee, Kichang, et al.
Publicado: (2025) -
Understanding The Effectiveness of Lossy Compression in Machine Learning Training Sets
por: Underwood, Robert, et al.
Publicado: (2024) -
Rotation Invariant Quantization for Model Compression
por: Kampeas, Joseph, et al.
Publicado: (2023) -
FedHypeVAE: Federated Learning with Hypernetwork Generated Conditional VAEs for Differentially Private Embedding Sharing
por: Gupta, Sunny, et al.
Publicado: (2026) -
Tensor Generalized Approximate Message Passing
por: Li, Yinchuan, et al.
Publicado: (2025)