Accelerating Data Chunking in Deduplication Systems using Vector Instructions
Fuente:
arXiv
Guardado en:
| Autores principales: | Udayashankar, Sreeharsha, Baba, Abdelrahman, Al-Kiswany, Samer |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Vectorized Sequence-Based Chunking for Data Deduplication
por: Udayashankar, Sreeharsha, et al.
Publicado: (2025)
por: Udayashankar, Sreeharsha, et al.
Publicado: (2025)
Balanced Data Placement for GEMV Acceleration with Processing-In-Memory
por: Ibrahim, Mohamed Assem, et al.
Publicado: (2024)
por: Ibrahim, Mohamed Assem, et al.
Publicado: (2024)
Accelerating Triangle Counting with Real Processing-in-Memory Systems
por: Asquini, Lorenzo, et al.
Publicado: (2025)
por: Asquini, Lorenzo, et al.
Publicado: (2025)
RISC-V Word-Size Modular Instructions for Residue Number Systems
por: Didier, Laurent-Stéphane, et al.
Publicado: (2024)
por: Didier, Laurent-Stéphane, et al.
Publicado: (2024)
MIMDRAM: An End-to-End Processing-Using-DRAM System for High-Throughput, Energy-Efficient and Programmer-Transparent Multiple-Instruction Multiple-Data Processing
por: Oliveira, Geraldo F., et al.
Publicado: (2024)
por: Oliveira, Geraldo F., et al.
Publicado: (2024)
FlexVector: A SpMM Vector Processor with Flexible VRF for GCNs on Varying-Sparsity Graphs
por: Li, Bohan, et al.
Publicado: (2026)
por: Li, Bohan, et al.
Publicado: (2026)
EDEA: Efficient Dual-Engine Accelerator for Depthwise Separable Convolution with Direct Data Transfer
por: Chen, Yi, et al.
Publicado: (2025)
por: Chen, Yi, et al.
Publicado: (2025)
Efficient Architecture for RISC-V Vector Memory Access
por: Guan, Hongyi, et al.
Publicado: (2025)
por: Guan, Hongyi, et al.
Publicado: (2025)
XDMA: A Distributed, Extensible DMA Architecture for Layout-Flexible Data Movements in Heterogeneous Multi-Accelerator SoCs
por: Kong, Fanchen, et al.
Publicado: (2025)
por: Kong, Fanchen, et al.
Publicado: (2025)
MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU Systems
por: Zhou, Zhuoshan, et al.
Publicado: (2026)
por: Zhou, Zhuoshan, et al.
Publicado: (2026)
Data-aware Dynamic Execution of Irregular Workloads on Heterogeneous Systems
por: Bai, Zhenyu, et al.
Publicado: (2025)
por: Bai, Zhenyu, et al.
Publicado: (2025)
Sparse MTTKRP Acceleration for Tensor Decomposition on GPU
por: Wijeratne, Sasindu, et al.
Publicado: (2024)
por: Wijeratne, Sasindu, et al.
Publicado: (2024)
Leveraging SIMD for Accelerating Large-number Arithmetic
por: Das, Subhrajit, et al.
Publicado: (2026)
por: Das, Subhrajit, et al.
Publicado: (2026)
Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUs
por: Zhang, Qijun, et al.
Publicado: (2026)
por: Zhang, Qijun, et al.
Publicado: (2026)
A Heterogeneous Chiplet Architecture for Accelerating End-to-End Transformer Models
por: Sharma, Harsh, et al.
Publicado: (2023)
por: Sharma, Harsh, et al.
Publicado: (2023)
Application Experiences on a GPU-Accelerated Arm-based HPC Testbed
por: Elwasif, Wael, et al.
Publicado: (2022)
por: Elwasif, Wael, et al.
Publicado: (2022)
Workload-Aware Hardware Accelerator Mining for Distributed Deep Learning Training
por: Adnan, Muhammad, et al.
Publicado: (2024)
por: Adnan, Muhammad, et al.
Publicado: (2024)
TAPA-CS: Enabling Scalable Accelerator Design on Distributed HBM-FPGAs
por: Prakriya, Neha, et al.
Publicado: (2023)
por: Prakriya, Neha, et al.
Publicado: (2023)
Design in Tiles: Automating GEMM Deployment on Tile-Based Many-PE Accelerators
por: Shen, Aofeng, et al.
Publicado: (2025)
por: Shen, Aofeng, et al.
Publicado: (2025)
FLEX: Leveraging FPGA-CPU Synergy for Mixed-Cell-Height Legalization Acceleration
por: Liu, Xingyu, et al.
Publicado: (2025)
por: Liu, Xingyu, et al.
Publicado: (2025)
MVDRAM: Enabling GeMV Execution in Unmodified DRAM for Low-Bit LLM Acceleration
por: Kubo, Tatsuya, et al.
Publicado: (2025)
por: Kubo, Tatsuya, et al.
Publicado: (2025)
An Evaluation and Comparison of GPU Hardware and Solver Libraries for Accelerating the OPM Flow Reservoir Simulator
por: Qiu, Tong Dong, et al.
Publicado: (2023)
por: Qiu, Tong Dong, et al.
Publicado: (2023)
CMDS: Cross-layer Dataflow Optimization for DNN Accelerators Exploiting Multi-bank Memories
por: Shi, Man, et al.
Publicado: (2024)
por: Shi, Man, et al.
Publicado: (2024)
DiP: A Scalable, Energy-Efficient Systolic Array for Matrix Multiplication Acceleration
por: Abdelmaksoud, Ahmed J., et al.
Publicado: (2024)
por: Abdelmaksoud, Ahmed J., et al.
Publicado: (2024)
CCSS: Hardware-Accelerated RTL Simulation with Fast Combinational Logic Computing and Sequential Logic Synchronization
por: Feng, Weigang, et al.
Publicado: (2025)
por: Feng, Weigang, et al.
Publicado: (2025)
A Survey of Real-time Scheduling on Accelerator-based Heterogeneous Architecture for Time Critical Applications
por: Zou, An, et al.
Publicado: (2025)
por: Zou, An, et al.
Publicado: (2025)
DP-HLS: A High-Level Synthesis Framework for Accelerating Dynamic Programming Algorithms in Bioinformatics
por: Cao, Yingqi, et al.
Publicado: (2024)
por: Cao, Yingqi, et al.
Publicado: (2024)
A Lightweight High-Throughput Collective-Capable NoC for Large-Scale ML Accelerators
por: Colagrande, Luca, et al.
Publicado: (2026)
por: Colagrande, Luca, et al.
Publicado: (2026)
SLIM: A Heterogeneous Accelerator for Edge Inference of Sparse Large Language Model via Adaptive Thresholding
por: Xu, Weihong, et al.
Publicado: (2025)
por: Xu, Weihong, et al.
Publicado: (2025)
DeepStack: Scalable and Accurate Design Space Exploration for Distributed 3D-Stacked AI Accelerators
por: Mo, Zhiwen, et al.
Publicado: (2026)
por: Mo, Zhiwen, et al.
Publicado: (2026)
MANOJAVAM: A Scalable, Unified FPGA Accelerator for Matrix Multiplication and Singular Value Decomposition in Principal Component Analysis
por: Ramasubramanian, Srivaths, et al.
Publicado: (2026)
por: Ramasubramanian, Srivaths, et al.
Publicado: (2026)
DeFiNES: Enabling Fast Exploration of the Depth-first Scheduling Space for DNN Accelerators through Analytical Modeling
por: Mei, Linyan, et al.
Publicado: (2022)
por: Mei, Linyan, et al.
Publicado: (2022)
PIMDAL: Mitigating the Memory Bottleneck in Data Analytics using a Real Processing-in-Memory System
por: Frouzakis, Manos, et al.
Publicado: (2025)
por: Frouzakis, Manos, et al.
Publicado: (2025)
Revisiting Computational Storage for Data Integrity and Security
por: Shi, Chao, et al.
Publicado: (2025)
por: Shi, Chao, et al.
Publicado: (2025)
Torrent: A Distributed DMA for Efficient and Flexible Point-to-Multipoint Data Movement
por: Deng, Yunhao, et al.
Publicado: (2025)
por: Deng, Yunhao, et al.
Publicado: (2025)
Evaluating Rapid Makespan Predictions for Heterogeneous Systems with Programmable Logic
por: Wilhelm, Martin, et al.
Publicado: (2025)
por: Wilhelm, Martin, et al.
Publicado: (2025)
Navigating the Landscape of Distributed File Systems: Architectures, Implementations, and Considerations
por: Pan, Xueting, et al.
Publicado: (2024)
por: Pan, Xueting, et al.
Publicado: (2024)
FpgaHub: Fpga-centric Hyper-heterogeneous Computing Platform for Big Data Analytics
por: Wang, Zeke, et al.
Publicado: (2025)
por: Wang, Zeke, et al.
Publicado: (2025)
EPOCH: Enabling Preemption Operation for Context Saving in Heterogeneous FPGA Systems
por: Malik, Arsalan Ali, et al.
Publicado: (2025)
por: Malik, Arsalan Ali, et al.
Publicado: (2025)
New Tools, Programming Models, and System Support for Processing-in-Memory Architectures
por: Oliveira, Geraldo F.
Publicado: (2025)
por: Oliveira, Geraldo F.
Publicado: (2025)
Ejemplares similares
-
Vectorized Sequence-Based Chunking for Data Deduplication
por: Udayashankar, Sreeharsha, et al.
Publicado: (2025) -
Balanced Data Placement for GEMV Acceleration with Processing-In-Memory
por: Ibrahim, Mohamed Assem, et al.
Publicado: (2024) -
Accelerating Triangle Counting with Real Processing-in-Memory Systems
por: Asquini, Lorenzo, et al.
Publicado: (2025) -
RISC-V Word-Size Modular Instructions for Residue Number Systems
por: Didier, Laurent-Stéphane, et al.
Publicado: (2024) -
MIMDRAM: An End-to-End Processing-Using-DRAM System for High-Throughput, Energy-Efficient and Programmer-Transparent Multiple-Instruction Multiple-Data Processing
por: Oliveira, Geraldo F., et al.
Publicado: (2024)