Accelerating Retrieval-Augmented Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Quinn, Derrick, Nouri, Mohammad, Patel, Neel, Salihu, John, Salemi, Alireza, Lee, Sukhan, Zamani, Hamed, Alian, Mohammad |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Accelerating Triangle Counting with Real Processing-in-Memory Systems
von: Asquini, Lorenzo, et al.
Veröffentlicht: (2025)
von: Asquini, Lorenzo, et al.
Veröffentlicht: (2025)
Leveraging SIMD for Accelerating Large-number Arithmetic
von: Das, Subhrajit, et al.
Veröffentlicht: (2026)
von: Das, Subhrajit, et al.
Veröffentlicht: (2026)
Optimizing Task Scheduling in Fog Computing with Deadline Awareness
von: Sirjani, Mohammad Sadegh, et al.
Veröffentlicht: (2025)
von: Sirjani, Mohammad Sadegh, et al.
Veröffentlicht: (2025)
ALPHA-PIM: Analysis of Linear Algebraic Processing for High-Performance Graph Applications on a Real Processing-In-Memory System
von: Barkhordar, Marzieh, et al.
Veröffentlicht: (2026)
von: Barkhordar, Marzieh, et al.
Veröffentlicht: (2026)
HetGPU: The pursuit of making binary compatibility towards GPUs
von: Yang, Yiwei, et al.
Veröffentlicht: (2025)
von: Yang, Yiwei, et al.
Veröffentlicht: (2025)
Sparse MTTKRP Acceleration for Tensor Decomposition on GPU
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2024)
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2024)
A Modern Primer on Processing in Memory
von: Mutlu, Onur, et al.
Veröffentlicht: (2020)
von: Mutlu, Onur, et al.
Veröffentlicht: (2020)
Application Experiences on a GPU-Accelerated Arm-based HPC Testbed
von: Elwasif, Wael, et al.
Veröffentlicht: (2022)
von: Elwasif, Wael, et al.
Veröffentlicht: (2022)
Balanced Data Placement for GEMV Acceleration with Processing-In-Memory
von: Ibrahim, Mohamed Assem, et al.
Veröffentlicht: (2024)
von: Ibrahim, Mohamed Assem, et al.
Veröffentlicht: (2024)
Accelerating Data Chunking in Deduplication Systems using Vector Instructions
von: Udayashankar, Sreeharsha, et al.
Veröffentlicht: (2025)
von: Udayashankar, Sreeharsha, et al.
Veröffentlicht: (2025)
Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUs
von: Zhang, Qijun, et al.
Veröffentlicht: (2026)
von: Zhang, Qijun, et al.
Veröffentlicht: (2026)
A Heterogeneous Chiplet Architecture for Accelerating End-to-End Transformer Models
von: Sharma, Harsh, et al.
Veröffentlicht: (2023)
von: Sharma, Harsh, et al.
Veröffentlicht: (2023)
Workload-Aware Hardware Accelerator Mining for Distributed Deep Learning Training
von: Adnan, Muhammad, et al.
Veröffentlicht: (2024)
von: Adnan, Muhammad, et al.
Veröffentlicht: (2024)
TAPA-CS: Enabling Scalable Accelerator Design on Distributed HBM-FPGAs
von: Prakriya, Neha, et al.
Veröffentlicht: (2023)
von: Prakriya, Neha, et al.
Veröffentlicht: (2023)
Design in Tiles: Automating GEMM Deployment on Tile-Based Many-PE Accelerators
von: Shen, Aofeng, et al.
Veröffentlicht: (2025)
von: Shen, Aofeng, et al.
Veröffentlicht: (2025)
FLEX: Leveraging FPGA-CPU Synergy for Mixed-Cell-Height Legalization Acceleration
von: Liu, Xingyu, et al.
Veröffentlicht: (2025)
von: Liu, Xingyu, et al.
Veröffentlicht: (2025)
An Evaluation and Comparison of GPU Hardware and Solver Libraries for Accelerating the OPM Flow Reservoir Simulator
von: Qiu, Tong Dong, et al.
Veröffentlicht: (2023)
von: Qiu, Tong Dong, et al.
Veröffentlicht: (2023)
MVDRAM: Enabling GeMV Execution in Unmodified DRAM for Low-Bit LLM Acceleration
von: Kubo, Tatsuya, et al.
Veröffentlicht: (2025)
von: Kubo, Tatsuya, et al.
Veröffentlicht: (2025)
CMDS: Cross-layer Dataflow Optimization for DNN Accelerators Exploiting Multi-bank Memories
von: Shi, Man, et al.
Veröffentlicht: (2024)
von: Shi, Man, et al.
Veröffentlicht: (2024)
DiP: A Scalable, Energy-Efficient Systolic Array for Matrix Multiplication Acceleration
von: Abdelmaksoud, Ahmed J., et al.
Veröffentlicht: (2024)
von: Abdelmaksoud, Ahmed J., et al.
Veröffentlicht: (2024)
EDEA: Efficient Dual-Engine Accelerator for Depthwise Separable Convolution with Direct Data Transfer
von: Chen, Yi, et al.
Veröffentlicht: (2025)
von: Chen, Yi, et al.
Veröffentlicht: (2025)
DP-HLS: A High-Level Synthesis Framework for Accelerating Dynamic Programming Algorithms in Bioinformatics
von: Cao, Yingqi, et al.
Veröffentlicht: (2024)
von: Cao, Yingqi, et al.
Veröffentlicht: (2024)
CCSS: Hardware-Accelerated RTL Simulation with Fast Combinational Logic Computing and Sequential Logic Synchronization
von: Feng, Weigang, et al.
Veröffentlicht: (2025)
von: Feng, Weigang, et al.
Veröffentlicht: (2025)
A Survey of Real-time Scheduling on Accelerator-based Heterogeneous Architecture for Time Critical Applications
von: Zou, An, et al.
Veröffentlicht: (2025)
von: Zou, An, et al.
Veröffentlicht: (2025)
A Lightweight High-Throughput Collective-Capable NoC for Large-Scale ML Accelerators
von: Colagrande, Luca, et al.
Veröffentlicht: (2026)
von: Colagrande, Luca, et al.
Veröffentlicht: (2026)
SLIM: A Heterogeneous Accelerator for Edge Inference of Sparse Large Language Model via Adaptive Thresholding
von: Xu, Weihong, et al.
Veröffentlicht: (2025)
von: Xu, Weihong, et al.
Veröffentlicht: (2025)
DeepStack: Scalable and Accurate Design Space Exploration for Distributed 3D-Stacked AI Accelerators
von: Mo, Zhiwen, et al.
Veröffentlicht: (2026)
von: Mo, Zhiwen, et al.
Veröffentlicht: (2026)
MANOJAVAM: A Scalable, Unified FPGA Accelerator for Matrix Multiplication and Singular Value Decomposition in Principal Component Analysis
von: Ramasubramanian, Srivaths, et al.
Veröffentlicht: (2026)
von: Ramasubramanian, Srivaths, et al.
Veröffentlicht: (2026)
HgPCN: A Heterogeneous Architecture for E2E Embedded Point Cloud Inference
von: Gao, Yiming, et al.
Veröffentlicht: (2025)
von: Gao, Yiming, et al.
Veröffentlicht: (2025)
XDMA: A Distributed, Extensible DMA Architecture for Layout-Flexible Data Movements in Heterogeneous Multi-Accelerator SoCs
von: Kong, Fanchen, et al.
Veröffentlicht: (2025)
von: Kong, Fanchen, et al.
Veröffentlicht: (2025)
DeFiNES: Enabling Fast Exploration of the Depth-first Scheduling Space for DNN Accelerators through Analytical Modeling
von: Mei, Linyan, et al.
Veröffentlicht: (2022)
von: Mei, Linyan, et al.
Veröffentlicht: (2022)
MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU Systems
von: Zhou, Zhuoshan, et al.
Veröffentlicht: (2026)
von: Zhou, Zhuoshan, et al.
Veröffentlicht: (2026)
Conduit: Programmer-Transparent Near-Data Processing Using Multiple Compute-Capable Resources in Solid State Drives
von: Nadig, Rakesh, et al.
Veröffentlicht: (2026)
von: Nadig, Rakesh, et al.
Veröffentlicht: (2026)
Functionally-Complete Boolean Logic in Real DRAM Chips: Experimental Characterization and Analysis
von: Yuksel, Ismail Emir, et al.
Veröffentlicht: (2024)
von: Yuksel, Ismail Emir, et al.
Veröffentlicht: (2024)
Proteus: Enabling High-Performance Processing-Using-DRAM with Dynamic Bit-Precision, Adaptive Data Representation, and Flexible Arithmetic
von: Oliveira, Geraldo F., et al.
Veröffentlicht: (2025)
von: Oliveira, Geraldo F., et al.
Veröffentlicht: (2025)
RevaMp3D: Architecting the Processor Core and Cache Hierarchy for Systems with Monolithically-Integrated Logic and Memory
von: Ghiasi, Nika Mansouri, et al.
Veröffentlicht: (2022)
von: Ghiasi, Nika Mansouri, et al.
Veröffentlicht: (2022)
PULSAR: Simultaneous Many-Row Activation for Reliable and High-Performance Computing in Off-the-Shelf DRAM Chips
von: Yuksel, Ismail Emir, et al.
Veröffentlicht: (2023)
von: Yuksel, Ismail Emir, et al.
Veröffentlicht: (2023)
Simultaneous Many-Row Activation in Off-the-Shelf DRAM Chips: Experimental Characterization and Analysis
von: Yuksel, Ismail Emir, et al.
Veröffentlicht: (2024)
von: Yuksel, Ismail Emir, et al.
Veröffentlicht: (2024)
Multi-Partner Project: Multi-GPU Performance Portability Analysis for CFD Simulations at Scale
von: Eleftherakis, Panagiotis-Eleftherios, et al.
Veröffentlicht: (2026)
von: Eleftherakis, Panagiotis-Eleftherios, et al.
Veröffentlicht: (2026)
TT-Edge: A Hardware-Software Co-Design for Energy-Efficient Tensor-Train Decomposition on Edge AI
von: Kwak, Hyunseok, et al.
Veröffentlicht: (2025)
von: Kwak, Hyunseok, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Accelerating Triangle Counting with Real Processing-in-Memory Systems
von: Asquini, Lorenzo, et al.
Veröffentlicht: (2025) -
Leveraging SIMD for Accelerating Large-number Arithmetic
von: Das, Subhrajit, et al.
Veröffentlicht: (2026) -
Optimizing Task Scheduling in Fog Computing with Deadline Awareness
von: Sirjani, Mohammad Sadegh, et al.
Veröffentlicht: (2025) -
ALPHA-PIM: Analysis of Linear Algebraic Processing for High-Performance Graph Applications on a Real Processing-In-Memory System
von: Barkhordar, Marzieh, et al.
Veröffentlicht: (2026) -
HetGPU: The pursuit of making binary compatibility towards GPUs
von: Yang, Yiwei, et al.
Veröffentlicht: (2025)