CAT: Cellular Automata on Tensor cores
Fuente:
arXiv
Saved in:
| Main Authors: | Navarro, Cristóbal A., Quezada, Felipe A., Meneses, Enzo, Ferrada, Héctor, Hitschfeld, Nancy |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Convex Hull 3D Filtering with GPU Ray Tracing and Tensor Cores
by: Carrasco, Roberto, et al.
Published: (2026)
by: Carrasco, Roberto, et al.
Published: (2026)
Advancing RT Core-Accelerated Fixed-Radius Nearest Neighbor Search
by: Meneses, Enzo, et al.
Published: (2026)
by: Meneses, Enzo, et al.
Published: (2026)
Ray Tracing Cores for General-Purpose Computing: A Literature Review
by: Meneses, Enzo, et al.
Published: (2026)
by: Meneses, Enzo, et al.
Published: (2026)
Leveraging Mathematical Reasoning of LLMs for Efficient GPU Thread Mapping
by: Maureira, Jose, et al.
Published: (2026)
by: Maureira, Jose, et al.
Published: (2026)
Coefficient Synthesis for Threshold Automata
by: Balasubramanian, A. R.
Published: (2023)
by: Balasubramanian, A. R.
Published: (2023)
TACO: A Toolsuite for the Verification of Threshold Automata
by: Eichler, Paul, et al.
Published: (2026)
by: Eichler, Paul, et al.
Published: (2026)
Enhanced OpenMP Algorithm to Compute All-Pairs Shortest Path on x86 Architectures
by: Calderón, Sergio, et al.
Published: (2024)
by: Calderón, Sergio, et al.
Published: (2024)
Parameterized Verification of Round-based Distributed Algorithms via Extended Threshold Automata
by: Baumeister, Tom, et al.
Published: (2024)
by: Baumeister, Tom, et al.
Published: (2024)
Synergistic Tensor and Pipeline Parallelism
by: Qi, Mengshi, et al.
Published: (2025)
by: Qi, Mengshi, et al.
Published: (2025)
Minimizing CGYRO HPC Communication Costs in Ensembles with XGYRO by Sharing the Collisional Constant Tensor Structure
by: Sfiligoi, Igor, et al.
Published: (2025)
by: Sfiligoi, Igor, et al.
Published: (2025)
Building a Theory of Distributed Systems: Work by Nancy Lynch and Collaborators
by: Lynch, Nancy
Published: (2025)
by: Lynch, Nancy
Published: (2025)
Complexity of Verification and Synthesis of Threshold Automata
by: Balasubramanian, A. R., et al.
Published: (2020)
by: Balasubramanian, A. R., et al.
Published: (2020)
Federated Learning Using Coupled Tensor Train Decomposition
by: Zhang, Xiangtao, et al.
Published: (2024)
by: Zhang, Xiangtao, et al.
Published: (2024)
Accelerating Sparse MTTKRP for Small Tensor Decomposition on GPU
by: Wijeratne, Sasindu, et al.
Published: (2025)
by: Wijeratne, Sasindu, et al.
Published: (2025)
Collaborative Inference Acceleration with Non-Penetrative Tensor Partitioning
by: Liu, Zhibang, et al.
Published: (2025)
by: Liu, Zhibang, et al.
Published: (2025)
Do We Need Tensor Cores for Stencil Computations?
by: Gu, Qiqi, et al.
Published: (2026)
by: Gu, Qiqi, et al.
Published: (2026)
Accelerating Heterogeneous Tensor Parallelism via Flexible Workload Control
by: Wang, Zhigang, et al.
Published: (2024)
by: Wang, Zhigang, et al.
Published: (2024)
Accelerating Drug Discovery in AutoDock-GPU with Tensor Cores
by: Schieffer, Gabin, et al.
Published: (2024)
by: Schieffer, Gabin, et al.
Published: (2024)
Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services
by: Chen, Haoyu, et al.
Published: (2025)
by: Chen, Haoyu, et al.
Published: (2025)
PRISM: Processing-In-Memory Sparse MTTKRP for Tensor Decomposition Acceleration
by: Pacheco, Daniel, et al.
Published: (2026)
by: Pacheco, Daniel, et al.
Published: (2026)
Nitsum: Serving Tiered LLM Requests with Adaptive Tensor Parallelism
by: Srivatsa, Vikranth, et al.
Published: (2026)
by: Srivatsa, Vikranth, et al.
Published: (2026)
Achieving High-Performance Fault-Tolerant Routing in HyperX Interconnection Networks
by: Camarero, Cristóbal, et al.
Published: (2024)
by: Camarero, Cristóbal, et al.
Published: (2024)
Resource Allocation in HyperX Networks
by: Cano, Alejandro, et al.
Published: (2026)
by: Cano, Alejandro, et al.
Published: (2026)
High Performance Unstructured SpMM Computation Using Tensor Cores
by: Okanovic, Patrik, et al.
Published: (2024)
by: Okanovic, Patrik, et al.
Published: (2024)
Predictive Performance of Photonic SRAM-based In-Memory Computing for Tensor Decomposition
by: Wijeratne, Sasindu, et al.
Published: (2025)
by: Wijeratne, Sasindu, et al.
Published: (2025)
Minimizing Communication for Parallel Symmetric Tensor Times Same Vector Computation
by: Daas, Hussam Al, et al.
Published: (2025)
by: Daas, Hussam Al, et al.
Published: (2025)
Analyzing the Performance Portability of SYCL across CPUs, GPUs, and Hybrid Systems with SW Sequence Alignment
by: Costanzo, Manuel, et al.
Published: (2024)
by: Costanzo, Manuel, et al.
Published: (2024)
AMPED: Accelerating MTTKRP for Billion-Scale Sparse Tensor Decomposition on Multiple GPUs
by: Wijeratne, Sasindu, et al.
Published: (2025)
by: Wijeratne, Sasindu, et al.
Published: (2025)
AnchorTP: Resilient LLM Inference with State-Preserving Elastic Tensor Parallelism
by: Xu, Wendong, et al.
Published: (2025)
by: Xu, Wendong, et al.
Published: (2025)
Gensor: A Graph-based Construction Tensor Compilation Method for Deep Learning
by: Liu, Hangda, et al.
Published: (2025)
by: Liu, Hangda, et al.
Published: (2025)
SPIDER: Unleashing Sparse Tensor Cores for Stencil Computation via Strided Swapping
by: GU, Qiqi, et al.
Published: (2025)
by: GU, Qiqi, et al.
Published: (2025)
VDKMS: Vehicular Decentralized Key Management System for Cellular Vehicular-to-Everything Networks, A Blockchain-Based Approach
by: Yao, Wei, et al.
Published: (2023)
by: Yao, Wei, et al.
Published: (2023)
ZeroPP: Unleashing Exceptional Parallelism Efficiency through Tensor-Parallelism-Free Methodology
by: Tang, Ding, et al.
Published: (2024)
by: Tang, Ding, et al.
Published: (2024)
NeutronTP: Load-Balanced Distributed Full-Graph GNN Training with Tensor Parallelism
by: Ai, Xin, et al.
Published: (2024)
by: Ai, Xin, et al.
Published: (2024)
SW-TNC : Reaching the Most Complex Random Quantum Circuit via Tensor Network Contraction
by: Chen, Yaojian, et al.
Published: (2025)
by: Chen, Yaojian, et al.
Published: (2025)
Guaranteed DGEMM Accuracy While Using Reduced Precision Tensor Cores Through Extensions of the Ozaki Scheme
by: Schwarz, Angelika, et al.
Published: (2025)
by: Schwarz, Angelika, et al.
Published: (2025)
Vortex: Efficient Sample-Free Dynamic Tensor Program Optimization via Hardware-aware Strategy Space Hierarchization
by: Zhou, Yangjie, et al.
Published: (2024)
by: Zhou, Yangjie, et al.
Published: (2024)
Oases: Efficient Large-Scale Model Training on Commodity Servers via Overlapped and Automated Tensor Model Parallelism
by: Li, Shengwei, et al.
Published: (2023)
by: Li, Shengwei, et al.
Published: (2023)
Neural Cellular Automata Can Respond to Signals
by: Stovold, James
Published: (2023)
by: Stovold, James
Published: (2023)
cuFastTuckerPlus: A Stochastic Parallel Sparse FastTucker Decomposition Using GPU Tensor Cores
by: Li, Zixuan, et al.
Published: (2024)
by: Li, Zixuan, et al.
Published: (2024)
Similar Items
-
Convex Hull 3D Filtering with GPU Ray Tracing and Tensor Cores
by: Carrasco, Roberto, et al.
Published: (2026) -
Advancing RT Core-Accelerated Fixed-Radius Nearest Neighbor Search
by: Meneses, Enzo, et al.
Published: (2026) -
Ray Tracing Cores for General-Purpose Computing: A Literature Review
by: Meneses, Enzo, et al.
Published: (2026) -
Leveraging Mathematical Reasoning of LLMs for Efficient GPU Thread Mapping
by: Maureira, Jose, et al.
Published: (2026) -
Coefficient Synthesis for Threshold Automata
by: Balasubramanian, A. R.
Published: (2023)