Hardware Acceleration for Knowledge Graph Processing: Challenges & Recent Developments
Fuente:
arXiv
Guardado en:
| Autores principales: | Besta, Maciej, Gerstenberger, Robert, Iff, Patrick, Sonawane, Pournima, Luna, Juan Gómez, Kanakagiri, Raghavendra, Min, Rui, Kwaśniewski, Grzegorz, Mutlu, Onur, Hoefler, Torsten, Appuswamy, Raja, Mahony, Aidan O |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
FoldedHexaTorus: An Inter-Chiplet Interconnect Topology for Chiplet-based Systems using Organic and Glass Substrates
por: Iff, Patrick, et al.
Publicado: (2025)
por: Iff, Patrick, et al.
Publicado: (2025)
EvalNet: A Practical Toolchain for Generation and Analysis of Extreme-Scale Interconnects
por: Besta, Maciej, et al.
Publicado: (2021)
por: Besta, Maciej, et al.
Publicado: (2021)
Inductive Loop Analysis for Practical HPC Application Optimization
por: Schaad, Philipp, et al.
Publicado: (2025)
por: Schaad, Philipp, et al.
Publicado: (2025)
Minimum Cost Loop Nests for Contraction of a Sparse Tensor with a Tensor Network
por: Kanakagiri, Raghavendra, et al.
Publicado: (2023)
por: Kanakagiri, Raghavendra, et al.
Publicado: (2023)
FPsPIN: An FPGA-based Open-Hardware Research Platform for Processing in the Network
por: Schneider, Timo, et al.
Publicado: (2024)
por: Schneider, Timo, et al.
Publicado: (2024)
Near-Optimal Wafer-Scale Reduce
por: Luczynski, Piotr, et al.
Publicado: (2024)
por: Luczynski, Piotr, et al.
Publicado: (2024)
PlaceIT: Placement-based Inter-Chiplet Interconnect Topologies
por: Iff, Patrick, et al.
Publicado: (2025)
por: Iff, Patrick, et al.
Publicado: (2025)
Network Design for Wafer-Scale Systems with Wafer-on-Wafer Hybrid Bonding
por: Iff, Patrick, et al.
Publicado: (2026)
por: Iff, Patrick, et al.
Publicado: (2026)
Demystifying Chains, Trees, and Graphs of Thoughts
por: Besta, Maciej, et al.
Publicado: (2024)
por: Besta, Maciej, et al.
Publicado: (2024)
High Performance Unstructured SpMM Computation Using Tensor Cores
por: Okanovic, Patrik, et al.
Publicado: (2024)
por: Okanovic, Patrik, et al.
Publicado: (2024)
Demystifying Higher-Order Graph Neural Networks
por: Besta, Maciej, et al.
Publicado: (2024)
por: Besta, Maciej, et al.
Publicado: (2024)
Higher-Order Graph Databases
por: Besta, Maciej, et al.
Publicado: (2025)
por: Besta, Maciej, et al.
Publicado: (2025)
Benchmarking Filtered Approximate Nearest Neighbor Search Algorithms on Transformer-based Embedding Vectors
por: Iff, Patrick, et al.
Publicado: (2025)
por: Iff, Patrick, et al.
Publicado: (2025)
RapidChiplet: A Toolchain for Rapid Design Space Exploration of Chiplet Architectures
por: Iff, Patrick, et al.
Publicado: (2023)
por: Iff, Patrick, et al.
Publicado: (2023)
A Priori Loop Nest Normalization: Automatic Loop Scheduling in Complex Applications
por: Trümper, Lukas, et al.
Publicado: (2024)
por: Trümper, Lukas, et al.
Publicado: (2024)
Iterating Pointers: Enabling Static Analysis for Loop-based Pointers
por: Lepori, Andrea, et al.
Publicado: (2025)
por: Lepori, Andrea, et al.
Publicado: (2025)
DaCe AD: Unifying High-Performance Automatic Differentiation for Machine Learning and Scientific Computing
por: Boudaoud, Afif, et al.
Publicado: (2025)
por: Boudaoud, Afif, et al.
Publicado: (2025)
Confidential LLM Inference: Performance and Cost Across CPU and GPU TEEs
por: Chrapek, Marcin, et al.
Publicado: (2025)
por: Chrapek, Marcin, et al.
Publicado: (2025)
EDAN: Towards Understanding Memory Parallelism and Latency Sensitivity in HPC
por: Shen, Siyuan, et al.
Publicado: (2025)
por: Shen, Siyuan, et al.
Publicado: (2025)
PerfDojo: Automated ML Library Generation for Heterogeneous Architectures
por: Ivanov, Andrei, et al.
Publicado: (2025)
por: Ivanov, Andrei, et al.
Publicado: (2025)
Affordable AI Assistants with Knowledge Graph of Thoughts
por: Besta, Maciej, et al.
Publicado: (2025)
por: Besta, Maciej, et al.
Publicado: (2025)
PICO: Performance Insights for Collective Operations
por: Pasqualoni, Saverio, et al.
Publicado: (2025)
por: Pasqualoni, Saverio, et al.
Publicado: (2025)
Assessing the Performance of OpenTitan as Cryptographic Accelerator in Secure Open-Hardware System-on-Chips
por: Parisi, Emanuele, et al.
Publicado: (2024)
por: Parisi, Emanuele, et al.
Publicado: (2024)
Cleaning up the Mess: Re-Evaluating the Real-System Modeling Accuracy of Ramulator 2.0
por: Bostanci, F. Nisa, et al.
Publicado: (2025)
por: Bostanci, F. Nisa, et al.
Publicado: (2025)
CheckEmbed: Effective Verification of LLM Solutions to Open-Ended Tasks
por: Besta, Maciej, et al.
Publicado: (2024)
por: Besta, Maciej, et al.
Publicado: (2024)
Multi-Strided Access Patterns to Boost Hardware Prefetching
por: Blom, Miguel O., et al.
Publicado: (2024)
por: Blom, Miguel O., et al.
Publicado: (2024)
Denoising Application Performance Models with Noise-Resilient Priors
por: de Morais, Gustavo, et al.
Publicado: (2025)
por: de Morais, Gustavo, et al.
Publicado: (2025)
Hardware-Agnostic and Insightful Efficiency Metrics for Accelerated Systems: Definition and Implementation within TALP
por: Rahimi, Ghazal, et al.
Publicado: (2026)
por: Rahimi, Ghazal, et al.
Publicado: (2026)
Long-term Monitoring of Kernel and Hardware Events to Understand Latency Variance
por: Zhou, Fang, et al.
Publicado: (2026)
por: Zhou, Fang, et al.
Publicado: (2026)
Psychologically Enhanced AI Agents
por: Besta, Maciej, et al.
Publicado: (2025)
por: Besta, Maciej, et al.
Publicado: (2025)
TINA: Acceleration of Non-NN Signal Processing Algorithms Using NN Accelerators
por: Boerkamp, Christiaan, et al.
Publicado: (2024)
por: Boerkamp, Christiaan, et al.
Publicado: (2024)
KForge: Program Synthesis for Diverse AI Hardware Accelerators
por: Sereda, Taras, et al.
Publicado: (2025)
por: Sereda, Taras, et al.
Publicado: (2025)
Solving Combinatorial Optimization Problems on a Photonic Quantum Computer
por: Slysz, Mateusz, et al.
Publicado: (2024)
por: Slysz, Mateusz, et al.
Publicado: (2024)
ADELIA: Automatic Differentiation for Efficient Laplace Inference Approximations
por: Boudaoud, Afif, et al.
Publicado: (2026)
por: Boudaoud, Afif, et al.
Publicado: (2026)
Dissecting RISC-V Performance: Practical PMU Profiling and Hardware-Agnostic Roofline Analysis on Emerging Platforms
por: Batashev, Alexander
Publicado: (2025)
por: Batashev, Alexander
Publicado: (2025)
In-Network Collective Operations: Game Changer or Challenge for AI Workloads?
por: Hoefler, Torsten, et al.
Publicado: (2026)
por: Hoefler, Torsten, et al.
Publicado: (2026)
Hardware optimization on Android for inference of AI models
por: Gherasim, Iulius, et al.
Publicado: (2025)
por: Gherasim, Iulius, et al.
Publicado: (2025)
gpu tracker: Python Package for Tracking and Profiling GPU and Other Hardware Utilization in Both Desktop and High-Performance Computing Environments
por: Huckvale, Erik D., et al.
Publicado: (2024)
por: Huckvale, Erik D., et al.
Publicado: (2024)
Performance-Driven Optimization of Parallel Breadth-First Search
por: Bhaskar, Marati, et al.
Publicado: (2025)
por: Bhaskar, Marati, et al.
Publicado: (2025)
DCC: Data-Centric Compilation of Machine Learning Kernels for Processing-In-Memory Architectures
por: Yang, Peiming, et al.
Publicado: (2025)
por: Yang, Peiming, et al.
Publicado: (2025)
Ejemplares similares
-
FoldedHexaTorus: An Inter-Chiplet Interconnect Topology for Chiplet-based Systems using Organic and Glass Substrates
por: Iff, Patrick, et al.
Publicado: (2025) -
EvalNet: A Practical Toolchain for Generation and Analysis of Extreme-Scale Interconnects
por: Besta, Maciej, et al.
Publicado: (2021) -
Inductive Loop Analysis for Practical HPC Application Optimization
por: Schaad, Philipp, et al.
Publicado: (2025) -
Minimum Cost Loop Nests for Contraction of a Sparse Tensor with a Tensor Network
por: Kanakagiri, Raghavendra, et al.
Publicado: (2023) -
FPsPIN: An FPGA-based Open-Hardware Research Platform for Processing in the Network
por: Schneider, Timo, et al.
Publicado: (2024)