Performance Optimization of 3D Stencil Computation on ARM Scalable Vector Extension
Fuente:
arXiv
Saved in:
| Main Author: | Chen, Hongguang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
oneDAL Optimization for ARM Scalable Vector Extension: Maximizing Efficiency for High-Performance Data Science
by: Sharma, Chandan, et al.
Published: (2025)
by: Sharma, Chandan, et al.
Published: (2025)
Stencil-Lifting: Hierarchical Recursive Lifting System for Extracting Summary of Stencil Kernel in Legacy Codes
by: Li, Mingyi, et al.
Published: (2025)
by: Li, Mingyi, et al.
Published: (2025)
Comparison of Vectorization Capabilities of Different Compilers for X86 and ARM CPUs
by: Sakib, Nazmus, et al.
Published: (2025)
by: Sakib, Nazmus, et al.
Published: (2025)
Scalable Packed Layouts for Vector-Length-Agnostic ML Code Generation
by: Beysel, Ege, et al.
Published: (2026)
by: Beysel, Ege, et al.
Published: (2026)
Cyclic Data Streaming on GPUs for Short Range Stencils Applied to Molecular Dynamics
by: Rose, Martin, et al.
Published: (2025)
by: Rose, Martin, et al.
Published: (2025)
Performance Evaluation of CMOS Annealing with Support Vector Machine
by: Fukuhara, Ryoga, et al.
Published: (2024)
by: Fukuhara, Ryoga, et al.
Published: (2024)
On the Performance of Cloud-based ARM SVE for Zero-Knowledge Proving Systems
by: Loghin, Dumitrel, et al.
Published: (2025)
by: Loghin, Dumitrel, et al.
Published: (2025)
Acceleration and energy consumption optimization in cascading classifiers for face detection on low-cost ARM big.LITTLE asymmetric architectures
by: Corpas, Alberto, et al.
Published: (2024)
by: Corpas, Alberto, et al.
Published: (2024)
Performance of Confidential Computing GPUs
by: Ibarra, Antonio Martínez, et al.
Published: (2025)
by: Ibarra, Antonio Martínez, et al.
Published: (2025)
Scalable Systems and Software Architectures for High-Performance Computing on cloud platforms
by: Ramesh, Risshab Srinivas
Published: (2024)
by: Ramesh, Risshab Srinivas
Published: (2024)
Performance Characterization of Containers in Edge Computing
by: Gupta, Ragini, et al.
Published: (2025)
by: Gupta, Ragini, et al.
Published: (2025)
Two Criteria for Performance Analysis of Optimization Algorithms
by: Jing, Yunpeng, et al.
Published: (2024)
by: Jing, Yunpeng, et al.
Published: (2024)
Performance Characterization and Optimizations of Traditional ML Applications
by: Kumar, Harsh, et al.
Published: (2024)
by: Kumar, Harsh, et al.
Published: (2024)
Evaluating the Performance of the DeepSeek Model in Confidential Computing Environment
by: Dong, Ben, et al.
Published: (2025)
by: Dong, Ben, et al.
Published: (2025)
A Continuous Benchmarking Infrastructure for High-Performance Computing Applications
by: Alt, Christoph, et al.
Published: (2024)
by: Alt, Christoph, et al.
Published: (2024)
TideGS: Scalable Training of Over One Billion 3D Gaussian Splatting Primitives via Out-of-Core Optimization
by: Zhong, Chonghao, et al.
Published: (2026)
by: Zhong, Chonghao, et al.
Published: (2026)
Enhanced Scalability in Assessing Quantum Integer Factorization Performance
by: Lee, Junseo
Published: (2023)
by: Lee, Junseo
Published: (2023)
The Multiserver-Job Stochastic Recurrence Equation for Cloud Computing Performance Evaluation
by: Baccelli, Francois, et al.
Published: (2026)
by: Baccelli, Francois, et al.
Published: (2026)
Hierarchical Analyses Applied to Computer System Performance: Review and Call for Further Studies
by: Thomasian, Alexander
Published: (2024)
by: Thomasian, Alexander
Published: (2024)
Pinching-Antenna Systems For Indoor Immersive Communications: A 3D-Modeling Based Performance Analysis
by: Wang, Yulei, et al.
Published: (2025)
by: Wang, Yulei, et al.
Published: (2025)
VDTuner: Automated Performance Tuning for Vector Data Management Systems
by: Yang, Tiannuo, et al.
Published: (2024)
by: Yang, Tiannuo, et al.
Published: (2024)
Opal: A Modular Framework for Optimizing Performance using Analytics and LLMs
by: Zaeed, Mohammad, et al.
Published: (2025)
by: Zaeed, Mohammad, et al.
Published: (2025)
ACALSim: A Scalable Parallel Simulation Framework for High-Performance System Design Space Exploration
by: Lin, Wei-Fen, et al.
Published: (2026)
by: Lin, Wei-Fen, et al.
Published: (2026)
A Scalable k-Medoids Clustering via Whale Optimization Algorithm
by: Chenan, Huang, et al.
Published: (2024)
by: Chenan, Huang, et al.
Published: (2024)
OSCAR-P and aMLLibrary: Profiling and Predicting the Performance of FaaS-based Applications in Computing Continua
by: Sala, Roberto, et al.
Published: (2024)
by: Sala, Roberto, et al.
Published: (2024)
Accurate and Scalable Many-Node Simulation
by: Eyerman, Stijn, et al.
Published: (2024)
by: Eyerman, Stijn, et al.
Published: (2024)
From HNSW to Information-Theoretic Binarization: Rethinking the Architecture of Scalable Vector Search
by: Abtahi, Seyed Moein, et al.
Published: (2025)
by: Abtahi, Seyed Moein, et al.
Published: (2025)
Scalable GPU Performance Variability Analysis framework
by: Lahiry, Ankur, et al.
Published: (2025)
by: Lahiry, Ankur, et al.
Published: (2025)
Characterizing and Optimizing Realistic Workloads on a Commercial Compute-in-SRAM Device
by: Zhang, Niansong, et al.
Published: (2025)
by: Zhang, Niansong, et al.
Published: (2025)
Toward Efficient and Scalable Design of In-Memory Graph-Based Vector Search
by: Azizi, Ilias, et al.
Published: (2025)
by: Azizi, Ilias, et al.
Published: (2025)
Impact of Extensions on Browser Performance: An Empirical Study on Google Chrome
by: Jin, Bihui, et al.
Published: (2024)
by: Jin, Bihui, et al.
Published: (2024)
RAVE: RISC-V Analyzer of Vector Executions, a QEMU tracing plugin
by: Vizcaino, Pablo, et al.
Published: (2024)
by: Vizcaino, Pablo, et al.
Published: (2024)
Vectorized FlashAttention with Low-cost Exponential Computation in RISC-V Vector Processors
by: Titopoulos, Vasileios, et al.
Published: (2025)
by: Titopoulos, Vasileios, et al.
Published: (2025)
Performance Characterization of Expert Router for Scalable LLM Inference
by: Pichlmeier, Josef, et al.
Published: (2024)
by: Pichlmeier, Josef, et al.
Published: (2024)
gpu tracker: Python Package for Tracking and Profiling GPU and Other Hardware Utilization in Both Desktop and High-Performance Computing Environments
by: Huckvale, Erik D., et al.
Published: (2024)
by: Huckvale, Erik D., et al.
Published: (2024)
Cloud Computing Energy Consumption Prediction Based on Kernel Extreme Learning Machine Algorithm Improved by Vector Weighted Average Algorithm
by: Wang, Yuqing, et al.
Published: (2025)
by: Wang, Yuqing, et al.
Published: (2025)
Towards High-Performance and Portable Molecular Docking on CPUs through Vectorization
by: Accordi, Gianmarco, et al.
Published: (2025)
by: Accordi, Gianmarco, et al.
Published: (2025)
Towards Computational Performance Engineering for Unsupervised Concept Drift Detection -- Complexities, Benchmarking, Performance Analysis
by: Werner, Elias, et al.
Published: (2023)
by: Werner, Elias, et al.
Published: (2023)
Tracing Optimization for Performance Modeling and Regression Detection
by: Shahedi, Kaveh, et al.
Published: (2024)
by: Shahedi, Kaveh, et al.
Published: (2024)
Delegation with Trust<T>: A Scalable, Type- and Memory-Safe Alternative to Locks
by: Ahmad, Noaman, et al.
Published: (2024)
by: Ahmad, Noaman, et al.
Published: (2024)
Similar Items
-
oneDAL Optimization for ARM Scalable Vector Extension: Maximizing Efficiency for High-Performance Data Science
by: Sharma, Chandan, et al.
Published: (2025) -
Stencil-Lifting: Hierarchical Recursive Lifting System for Extracting Summary of Stencil Kernel in Legacy Codes
by: Li, Mingyi, et al.
Published: (2025) -
Comparison of Vectorization Capabilities of Different Compilers for X86 and ARM CPUs
by: Sakib, Nazmus, et al.
Published: (2025) -
Scalable Packed Layouts for Vector-Length-Agnostic ML Code Generation
by: Beysel, Ege, et al.
Published: (2026) -
Cyclic Data Streaming on GPUs for Short Range Stencils Applied to Molecular Dynamics
by: Rose, Martin, et al.
Published: (2025)