CPU-less parallel execution of lambda calculus in digital logic
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fitchett, Harry, Fox, Charles |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Extending CPU-less parallel execution of lambda calculus in digital logic with lists and arithmetic
von: Fitchett, Harry, et al.
Veröffentlicht: (2026)
von: Fitchett, Harry, et al.
Veröffentlicht: (2026)
Massimult: A Novel Parallel CPU Architecture Based on Combinator Reduction
von: Nicklisch-Franken, Jurgen, et al.
Veröffentlicht: (2024)
von: Nicklisch-Franken, Jurgen, et al.
Veröffentlicht: (2024)
Sandwich: Joint Configuration Search and Hot-Switching for Efficient CPU LLM Serving
von: Zhao, Juntao, et al.
Veröffentlicht: (2025)
von: Zhao, Juntao, et al.
Veröffentlicht: (2025)
Agentic Operator Generation for ML ASICs
von: Hammond, Alec M., et al.
Veröffentlicht: (2025)
von: Hammond, Alec M., et al.
Veröffentlicht: (2025)
Characterizing CPU-Induced Slowdowns in Multi-GPU LLM Inference
von: Chung, Euijun, et al.
Veröffentlicht: (2026)
von: Chung, Euijun, et al.
Veröffentlicht: (2026)
FLEX: Leveraging FPGA-CPU Synergy for Mixed-Cell-Height Legalization Acceleration
von: Liu, Xingyu, et al.
Veröffentlicht: (2025)
von: Liu, Xingyu, et al.
Veröffentlicht: (2025)
Evaluation of computational and energy performance in matrix multiplication algorithms on CPU and GPU using MKL, cuBLAS and SYCL
von: Torres, L. A., et al.
Veröffentlicht: (2024)
von: Torres, L. A., et al.
Veröffentlicht: (2024)
FPGA or GPU? Analyzing comparative research for application-specific guidance
von: Purkayastha, Arnab A, et al.
Veröffentlicht: (2025)
von: Purkayastha, Arnab A, et al.
Veröffentlicht: (2025)
TAPA: A Scalable Task-Parallel Dataflow Programming Framework for Modern FPGAs with Co-Optimization of HLS and Physical Design
von: Guo, Licheng, et al.
Veröffentlicht: (2022)
von: Guo, Licheng, et al.
Veröffentlicht: (2022)
Linear Layouts: Robust Code Generation of Efficient Tensor Computation Using $\mathbb{F}_2$
von: Zhou, Keren, et al.
Veröffentlicht: (2025)
von: Zhou, Keren, et al.
Veröffentlicht: (2025)
Investigating Memory Failure Prediction Across CPU Architectures
von: Yu, Qiao, et al.
Veröffentlicht: (2024)
von: Yu, Qiao, et al.
Veröffentlicht: (2024)
ODIN-Based CPU-GPU Architecture with Replay-Driven Simulation and Emulation
von: Dorairaj, Nij, et al.
Veröffentlicht: (2026)
von: Dorairaj, Nij, et al.
Veröffentlicht: (2026)
Taming Asynchronous CPU-GPU Coupling for Frequency-aware Latency Estimation on Mobile Edge
von: Chen, Jiesong, et al.
Veröffentlicht: (2026)
von: Chen, Jiesong, et al.
Veröffentlicht: (2026)
Exploring energy consumption of AI frameworks on a 64-core RV64 Server CPU
von: Malenza, Giulio, et al.
Veröffentlicht: (2025)
von: Malenza, Giulio, et al.
Veröffentlicht: (2025)
How Fast Can Graph Computations Go on Fine-grained Parallel Architectures
von: Wang, Yuqing, et al.
Veröffentlicht: (2025)
von: Wang, Yuqing, et al.
Veröffentlicht: (2025)
TriMoE: Augmenting GPU with AMX-Enabled CPU and DIMM-NDP for High-Throughput MoE Inference via Offloading
von: Pan, Yudong, et al.
Veröffentlicht: (2026)
von: Pan, Yudong, et al.
Veröffentlicht: (2026)
Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2025)
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2025)
PAM: Processing Across Memory Hierarchy for Efficient KV-centric LLM Serving System
von: Liu, Lian, et al.
Veröffentlicht: (2026)
von: Liu, Lian, et al.
Veröffentlicht: (2026)
Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems
von: Zhang, Chen, et al.
Veröffentlicht: (2026)
von: Zhang, Chen, et al.
Veröffentlicht: (2026)
FlexVector: A SpMM Vector Processor with Flexible VRF for GCNs on Varying-Sparsity Graphs
von: Li, Bohan, et al.
Veröffentlicht: (2026)
von: Li, Bohan, et al.
Veröffentlicht: (2026)
Sequence-Aware Split Heuristic to Mitigate SM Underutilization in FlashAttention-3 Low-Head-Count Decoding
von: Font, Martí Llopart, et al.
Veröffentlicht: (2026)
von: Font, Martí Llopart, et al.
Veröffentlicht: (2026)
iHAC: A Hybrid Cluster Architecture for Enhanced Performance and Resilience
von: Muntaka, Siddique Abubakr, et al.
Veröffentlicht: (2026)
von: Muntaka, Siddique Abubakr, et al.
Veröffentlicht: (2026)
Context-aware Simopt-Power: Using structural data with simulation metadata to optimise FPGA designs
von: Wadhwa, Eashan, et al.
Veröffentlicht: (2026)
von: Wadhwa, Eashan, et al.
Veröffentlicht: (2026)
Conduit: Programmer-Transparent Near-Data Processing Using Multiple Compute-Capable Resources in Solid State Drives
von: Nadig, Rakesh, et al.
Veröffentlicht: (2026)
von: Nadig, Rakesh, et al.
Veröffentlicht: (2026)
Exploring the Efficiency of 3D-Stacked AI Chip Architecture for LLM Inference with Voxel
von: Liu, Yiqi, et al.
Veröffentlicht: (2026)
von: Liu, Yiqi, et al.
Veröffentlicht: (2026)
Microbenchmark-Driven Analytical Performance Modeling Across Modern GPU Architectures
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2026)
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2026)
EPAC: The Last Dance
von: Mantovani, Filippo, et al.
Veröffentlicht: (2026)
von: Mantovani, Filippo, et al.
Veröffentlicht: (2026)
HieraSparse: Hierarchical Semi-Structured Sparse KV Attention
von: Wang, Haoxuan, et al.
Veröffentlicht: (2026)
von: Wang, Haoxuan, et al.
Veröffentlicht: (2026)
Enabling Mixed criticality applications for the Versal AI-Engines
von: Sprave, Vincent, et al.
Veröffentlicht: (2026)
von: Sprave, Vincent, et al.
Veröffentlicht: (2026)
MANOJAVAM: A Scalable, Unified FPGA Accelerator for Matrix Multiplication and Singular Value Decomposition in Principal Component Analysis
von: Ramasubramanian, Srivaths, et al.
Veröffentlicht: (2026)
von: Ramasubramanian, Srivaths, et al.
Veröffentlicht: (2026)
Adaptive KV Cache Reuse for Fast Long-Context LLM Serving
von: li, Fei, et al.
Veröffentlicht: (2026)
von: li, Fei, et al.
Veröffentlicht: (2026)
MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU Systems
von: Zhou, Zhuoshan, et al.
Veröffentlicht: (2026)
von: Zhou, Zhuoshan, et al.
Veröffentlicht: (2026)
Analyzing Reverse Address Translation Overheads in Multi-GPU Scale-Up Pods
von: Fatima, Amel, et al.
Veröffentlicht: (2026)
von: Fatima, Amel, et al.
Veröffentlicht: (2026)
Fine-Grained Power and Energy Attribution on AMD GPU/APU-Based Exascale Nodes
von: McDaniel, Adam, et al.
Veröffentlicht: (2026)
von: McDaniel, Adam, et al.
Veröffentlicht: (2026)
Pooling Engram Conditional Memory in Large Language Models using CXL
von: Ma, Ruiyang, et al.
Veröffentlicht: (2026)
von: Ma, Ruiyang, et al.
Veröffentlicht: (2026)
Execution-Centric Characterization of FP8 Matrix Cores, Asynchronous Execution, and Structured Sparsity on AMD MI300A
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2026)
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2026)
DeepStack: Scalable and Accurate Design Space Exploration for Distributed 3D-Stacked AI Accelerators
von: Mo, Zhiwen, et al.
Veröffentlicht: (2026)
von: Mo, Zhiwen, et al.
Veröffentlicht: (2026)
SwarmIO: Towards 100 Million IOPS SSD Emulation for Next-generation GPU-centric Storage Systems
von: Kim, Hyeseong, et al.
Veröffentlicht: (2026)
von: Kim, Hyeseong, et al.
Veröffentlicht: (2026)
Multi-Partner Project: Multi-GPU Performance Portability Analysis for CFD Simulations at Scale
von: Eleftherakis, Panagiotis-Eleftherios, et al.
Veröffentlicht: (2026)
von: Eleftherakis, Panagiotis-Eleftherios, et al.
Veröffentlicht: (2026)
ALPHA-PIM: Analysis of Linear Algebraic Processing for High-Performance Graph Applications on a Real Processing-In-Memory System
von: Barkhordar, Marzieh, et al.
Veröffentlicht: (2026)
von: Barkhordar, Marzieh, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Extending CPU-less parallel execution of lambda calculus in digital logic with lists and arithmetic
von: Fitchett, Harry, et al.
Veröffentlicht: (2026) -
Massimult: A Novel Parallel CPU Architecture Based on Combinator Reduction
von: Nicklisch-Franken, Jurgen, et al.
Veröffentlicht: (2024) -
Sandwich: Joint Configuration Search and Hot-Switching for Efficient CPU LLM Serving
von: Zhao, Juntao, et al.
Veröffentlicht: (2025) -
Agentic Operator Generation for ML ASICs
von: Hammond, Alec M., et al.
Veröffentlicht: (2025) -
Characterizing CPU-Induced Slowdowns in Multi-GPU LLM Inference
von: Chung, Euijun, et al.
Veröffentlicht: (2026)