Nexus Machine: An Active Message Inspired Reconfigurable Architecture for Irregular Workloads
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Juneja, Rohan, Dangi, Pranav, Bandara, Thilini Kaushalya, Mitra, Tulika, Peh, Li-shiuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Building an Open CGRA Ecosystem for Agile Innovation
von: Juneja, Rohan, et al.
Veröffentlicht: (2025)
von: Juneja, Rohan, et al.
Veröffentlicht: (2025)
Sustainable Hardware Specialization
von: Dangi, Pranav, et al.
Veröffentlicht: (2024)
von: Dangi, Pranav, et al.
Veröffentlicht: (2024)
Enhancing CGRA Efficiency Through Aligned Compute and Communication Provisioning
von: Li, Zhaoying, et al.
Veröffentlicht: (2024)
von: Li, Zhaoying, et al.
Veröffentlicht: (2024)
A Data-Driven Dynamic Execution Orchestration Architecture
von: Bai, Zhenyu, et al.
Veröffentlicht: (2026)
von: Bai, Zhenyu, et al.
Veröffentlicht: (2026)
Data-aware Dynamic Execution of Irregular Workloads on Heterogeneous Systems
von: Bai, Zhenyu, et al.
Veröffentlicht: (2025)
von: Bai, Zhenyu, et al.
Veröffentlicht: (2025)
HALO: Hardware-aware quantization with low critical-path-delay weights for LLM acceleration
von: Juneja, Rohan, et al.
Veröffentlicht: (2025)
von: Juneja, Rohan, et al.
Veröffentlicht: (2025)
SWAT: Scalable and Efficient Window Attention-based Transformers Acceleration on FPGAs
von: Bai, Zhenyu, et al.
Veröffentlicht: (2024)
von: Bai, Zhenyu, et al.
Veröffentlicht: (2024)
TerEffic: Highly Efficient Ternary LLM Inference on FPGA
von: Yin, Chenyang, et al.
Veröffentlicht: (2025)
von: Yin, Chenyang, et al.
Veröffentlicht: (2025)
Hybrid Photonic-digital Accelerator for Attention Mechanism
von: Li, Huize, et al.
Veröffentlicht: (2025)
von: Li, Huize, et al.
Veröffentlicht: (2025)
NOVA: NoC-based Vector Unit for Mapping Attention Layers on a CNN Accelerator
von: Upadhyay, Mohit, et al.
Veröffentlicht: (2024)
von: Upadhyay, Mohit, et al.
Veröffentlicht: (2024)
SLTarch: Towards Scalable Point-Based Neural Rendering by Taming Workload Imbalance and Memory Irregularity
von: Li, Xingyang, et al.
Veröffentlicht: (2025)
von: Li, Xingyang, et al.
Veröffentlicht: (2025)
Messaging-based Adaptive Vector Computing (MAVeC) Accelerator for AI Workloads
von: Chowdhury, Md. Rownak Hossain, et al.
Veröffentlicht: (2024)
von: Chowdhury, Md. Rownak Hossain, et al.
Veröffentlicht: (2024)
Reconfigurable Stream Network Architecture
von: Wang, Chengyue, et al.
Veröffentlicht: (2024)
von: Wang, Chengyue, et al.
Veröffentlicht: (2024)
Architectural Classification of XR Workloads: Cross-Layer Archetypes and Implications
von: Shi, Xinyu, et al.
Veröffentlicht: (2026)
von: Shi, Xinyu, et al.
Veröffentlicht: (2026)
Spatzformer: An Efficient Reconfigurable Dual-Core RISC-V V Cluster for Mixed Scalar-Vector Workloads
von: Perotti, Matteo, et al.
Veröffentlicht: (2024)
von: Perotti, Matteo, et al.
Veröffentlicht: (2024)
DCRA: A Distributed Chiplet-based Reconfigurable Architecture for Irregular Applications
von: Orenes-Vera, Marcelo, et al.
Veröffentlicht: (2023)
von: Orenes-Vera, Marcelo, et al.
Veröffentlicht: (2023)
CIMinus: Empowering Sparse DNN Workloads Modeling and Exploration on SRAM-based CIM Architectures
von: Qi, Yingjie, et al.
Veröffentlicht: (2025)
von: Qi, Yingjie, et al.
Veröffentlicht: (2025)
3D-TrIM: A Memory-Efficient Spatial Computing Architecture for Convolution Workloads
von: Sestito, Cristian, et al.
Veröffentlicht: (2025)
von: Sestito, Cristian, et al.
Veröffentlicht: (2025)
Workload-Aware Early-Stage Power Delivery Network Optimization via Architectural Power Traces
von: Hayes, Oran, et al.
Veröffentlicht: (2026)
von: Hayes, Oran, et al.
Veröffentlicht: (2026)
Evolution, Challenges, and Optimization in Computer Architecture: The Role of Reconfigurable Systems
von: Ederhion, Jefferson, et al.
Veröffentlicht: (2024)
von: Ederhion, Jefferson, et al.
Veröffentlicht: (2024)
THERMOS: Thermally-Aware Multi-Objective Scheduling of AI Workloads on Heterogeneous Multi-Chiplet PIM Architectures
von: Kanani, Alish, et al.
Veröffentlicht: (2025)
von: Kanani, Alish, et al.
Veröffentlicht: (2025)
FILCO: Flexible Composing Architecture with Real-Time Reconfigurability for DNN Acceleration
von: Chen, Xingzhen, et al.
Veröffentlicht: (2026)
von: Chen, Xingzhen, et al.
Veröffentlicht: (2026)
FireFly-S: Exploiting Dual-Side Sparsity for Spiking Neural Networks Acceleration with Reconfigurable Spatial Architecture
von: Li, Tenglong, et al.
Veröffentlicht: (2024)
von: Li, Tenglong, et al.
Veröffentlicht: (2024)
Workload Characterization for Branch Predictability
von: Vikas, FNU, et al.
Veröffentlicht: (2025)
von: Vikas, FNU, et al.
Veröffentlicht: (2025)
A Fully Pipelined FIFO Based Polynomial Multiplication Hardware Architecture Based On Number Theoretic Transform
von: Heidarpur, Moslem, et al.
Veröffentlicht: (2025)
von: Heidarpur, Moslem, et al.
Veröffentlicht: (2025)
ICP: Exploiting Instruction Correlation for Prefetching Irregular Memory Accesses
von: Li, Mengming, et al.
Veröffentlicht: (2026)
von: Li, Mengming, et al.
Veröffentlicht: (2026)
SAT-MapIt: A SAT-based Modulo Scheduling Mapper for Coarse Grain Reconfigurable Architectures
von: Tirelli, Cristian, et al.
Veröffentlicht: (2025)
von: Tirelli, Cristian, et al.
Veröffentlicht: (2025)
Towards Efficient Neuro-Symbolic AI: From Workload Characterization to Hardware Architecture
von: Wan, Zishen, et al.
Veröffentlicht: (2024)
von: Wan, Zishen, et al.
Veröffentlicht: (2024)
CRISP: Hybrid Structured Sparsity for Class-aware Model Pruning
von: Aggarwal, Shivam, et al.
Veröffentlicht: (2023)
von: Aggarwal, Shivam, et al.
Veröffentlicht: (2023)
Mapping Space Exploration for Multi-Chiplet Accelerators Targeting LLM Inference Serving Workloads
von: Li, Boyu, et al.
Veröffentlicht: (2025)
von: Li, Boyu, et al.
Veröffentlicht: (2025)
Towards An Approach to Identify Divergences in Hardware Designs for HPC Workloads
von: Popovici, Doru Thom, et al.
Veröffentlicht: (2025)
von: Popovici, Doru Thom, et al.
Veröffentlicht: (2025)
Allspark: Workload Orchestration for Visual Transformers on Processing In-Memory Systems
von: Ge, Mengke, et al.
Veröffentlicht: (2024)
von: Ge, Mengke, et al.
Veröffentlicht: (2024)
UpDown: Programmable fine-grained Events for Scalable Performance on Irregular Applications
von: Rajasukumar, Andronicus, et al.
Veröffentlicht: (2024)
von: Rajasukumar, Andronicus, et al.
Veröffentlicht: (2024)
Energy-Efficient p-Bit-Based Fully-Connected Quantum-Inspired Simulated Annealer with Dual BRAM Architecture
von: Onizawa, Naoya, et al.
Veröffentlicht: (2026)
von: Onizawa, Naoya, et al.
Veröffentlicht: (2026)
ACS: Concurrent Kernel Execution on Irregular, Input-Dependent Computational Graphs
von: Durvasula, Sankeerth, et al.
Veröffentlicht: (2024)
von: Durvasula, Sankeerth, et al.
Veröffentlicht: (2024)
TROOP: At-the-Roofline Performance for Vector Processors on Low Operational Intensity Workloads
von: Purayil, Navaneeth Kunhi, et al.
Veröffentlicht: (2025)
von: Purayil, Navaneeth Kunhi, et al.
Veröffentlicht: (2025)
Communication Characterization of AI Workloads for Large-scale Multi-chiplet Accelerators
von: Musavi, Mariam, et al.
Veröffentlicht: (2024)
von: Musavi, Mariam, et al.
Veröffentlicht: (2024)
EnergAIzer: Fast and Accurate GPU Power Estimation Framework for AI Workloads
von: Lee, Kyungmi, et al.
Veröffentlicht: (2026)
von: Lee, Kyungmi, et al.
Veröffentlicht: (2026)
NEURA: A Unified and Retargetable Compilation Framework for Coarse-Grained Reconfigurable Architectures
von: Li, Shangkun, et al.
Veröffentlicht: (2026)
von: Li, Shangkun, et al.
Veröffentlicht: (2026)
MINISA: Minimal Instruction Set Architecture for Next-gen Reconfigurable Inference Accelerator
von: Tong, Jianming, et al.
Veröffentlicht: (2026)
von: Tong, Jianming, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Building an Open CGRA Ecosystem for Agile Innovation
von: Juneja, Rohan, et al.
Veröffentlicht: (2025) -
Sustainable Hardware Specialization
von: Dangi, Pranav, et al.
Veröffentlicht: (2024) -
Enhancing CGRA Efficiency Through Aligned Compute and Communication Provisioning
von: Li, Zhaoying, et al.
Veröffentlicht: (2024) -
A Data-Driven Dynamic Execution Orchestration Architecture
von: Bai, Zhenyu, et al.
Veröffentlicht: (2026) -
Data-aware Dynamic Execution of Irregular Workloads on Heterogeneous Systems
von: Bai, Zhenyu, et al.
Veröffentlicht: (2025)