Assessing the Performance of Analog Training for Transfer Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fagbohungbe, Omobayode, Lammie, Corey, Rasch, Malte J., Ando, Takashi, Gokmen, Tayfun, Narayanan, Vijay |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Neuromorphic Computing: A Theoretical Framework for Time, Space, and Energy Scaling
von: Aimone, James B
Veröffentlicht: (2025)
von: Aimone, James B
Veröffentlicht: (2025)
Wireless Sensor Networks as Parallel and Distributed Hardware Platform for Artificial Neural Networks
von: Serpen, Gursel
Veröffentlicht: (2025)
von: Serpen, Gursel
Veröffentlicht: (2025)
NeuroRing: Scaling Spiking Neural Networks via Multi-FPGA Bidirectional Ring Topologies and Stream-Dataflow Architectures
von: Hafiz, Muhammad Ihsan Al, et al.
Veröffentlicht: (2026)
von: Hafiz, Muhammad Ihsan Al, et al.
Veröffentlicht: (2026)
MAC-DO: An Efficient Output-Stationary GEMM Accelerator for CNNs Using DRAM Technology
von: Jeong, Minki, et al.
Veröffentlicht: (2022)
von: Jeong, Minki, et al.
Veröffentlicht: (2022)
NeuraChip: Accelerating GNN Computations with a Hash-based Decoupled Spatial Accelerator
von: Shivdikar, Kaustubh, et al.
Veröffentlicht: (2024)
von: Shivdikar, Kaustubh, et al.
Veröffentlicht: (2024)
STEMS: Spatial-Temporal Mapping For Spiking Neural Networks
von: Eissa, Sherif, et al.
Veröffentlicht: (2025)
von: Eissa, Sherif, et al.
Veröffentlicht: (2025)
EDEA: Efficient Dual-Engine Accelerator for Depthwise Separable Convolution with Direct Data Transfer
von: Chen, Yi, et al.
Veröffentlicht: (2025)
von: Chen, Yi, et al.
Veröffentlicht: (2025)
BlockAMC: Scalable In-Memory Analog Matrix Computing for Solving Linear Systems
von: Pan, Lunshuai, et al.
Veröffentlicht: (2024)
von: Pan, Lunshuai, et al.
Veröffentlicht: (2024)
OpenRASE: Service Function Chain Emulation
von: Krishnamohan, Theviyanthan, et al.
Veröffentlicht: (2025)
von: Krishnamohan, Theviyanthan, et al.
Veröffentlicht: (2025)
Towards a Decentralised Application-Centric Orchestration Framework in the Cloud-Edge Continuum
von: Ullah, Amjad, et al.
Veröffentlicht: (2025)
von: Ullah, Amjad, et al.
Veröffentlicht: (2025)
Workload-Aware Hardware Accelerator Mining for Distributed Deep Learning Training
von: Adnan, Muhammad, et al.
Veröffentlicht: (2024)
von: Adnan, Muhammad, et al.
Veröffentlicht: (2024)
SSDTrain: An Activation Offloading Framework to SSDs for Faster Large Language Model Training
von: Wu, Kun, et al.
Veröffentlicht: (2024)
von: Wu, Kun, et al.
Veröffentlicht: (2024)
Optimizing Offload Performance in Heterogeneous MPSoCs
von: Colagrande, Luca, et al.
Veröffentlicht: (2024)
von: Colagrande, Luca, et al.
Veröffentlicht: (2024)
Chopper: A Multi-Level GPU Characterization Tool & Derived Insights Into LLM Training Inefficiency
von: Kurzynski, Marco, et al.
Veröffentlicht: (2025)
von: Kurzynski, Marco, et al.
Veröffentlicht: (2025)
A Fresh Approach to Evaluate Performance in Distributed Parallel Genetic Algorithms
von: Harada, Tomohiro, et al.
Veröffentlicht: (2021)
von: Harada, Tomohiro, et al.
Veröffentlicht: (2021)
iHAC: A Hybrid Cluster Architecture for Enhanced Performance and Resilience
von: Muntaka, Siddique Abubakr, et al.
Veröffentlicht: (2026)
von: Muntaka, Siddique Abubakr, et al.
Veröffentlicht: (2026)
Microbenchmark-Driven Analytical Performance Modeling Across Modern GPU Architectures
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2026)
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2026)
TT-Edge: A Hardware-Software Co-Design for Energy-Efficient Tensor-Train Decomposition on Edge AI
von: Kwak, Hyunseok, et al.
Veröffentlicht: (2025)
von: Kwak, Hyunseok, et al.
Veröffentlicht: (2025)
Multi-Partner Project: Multi-GPU Performance Portability Analysis for CFD Simulations at Scale
von: Eleftherakis, Panagiotis-Eleftherios, et al.
Veröffentlicht: (2026)
von: Eleftherakis, Panagiotis-Eleftherios, et al.
Veröffentlicht: (2026)
PULSAR: Simultaneous Many-Row Activation for Reliable and High-Performance Computing in Off-the-Shelf DRAM Chips
von: Yuksel, Ismail Emir, et al.
Veröffentlicht: (2023)
von: Yuksel, Ismail Emir, et al.
Veröffentlicht: (2023)
Proteus: Enabling High-Performance Processing-Using-DRAM with Dynamic Bit-Precision, Adaptive Data Representation, and Flexible Arithmetic
von: Oliveira, Geraldo F., et al.
Veröffentlicht: (2025)
von: Oliveira, Geraldo F., et al.
Veröffentlicht: (2025)
ALPHA-PIM: Analysis of Linear Algebraic Processing for High-Performance Graph Applications on a Real Processing-In-Memory System
von: Barkhordar, Marzieh, et al.
Veröffentlicht: (2026)
von: Barkhordar, Marzieh, et al.
Veröffentlicht: (2026)
D-CODE: Data Colony Optimization for Dynamic Network Efficiency
von: Pandey, Tannu, et al.
Veröffentlicht: (2024)
von: Pandey, Tannu, et al.
Veröffentlicht: (2024)
A Frequency-based Parent Selection for Reducing the Effect of Evaluation Time Bias in Asynchronous Parallel Multi-objective Evolutionary Algorithms
von: Harada, Tomohiro
Veröffentlicht: (2021)
von: Harada, Tomohiro
Veröffentlicht: (2021)
Code generation and runtime techniques for enabling data-efficient deep learning training on GPUs
von: Wu, Kun
Veröffentlicht: (2024)
von: Wu, Kun
Veröffentlicht: (2024)
Neuromorphic Simulation of Drosophila Melanogaster Brain Connectome on Loihi 2
von: Wang, Felix, et al.
Veröffentlicht: (2025)
von: Wang, Felix, et al.
Veröffentlicht: (2025)
Evaluation and Efficiency Comparison of Evolutionary Algorithms for Service Placement Optimization in Fog Architectures
von: Guerrero, Carlos, et al.
Veröffentlicht: (2025)
von: Guerrero, Carlos, et al.
Veröffentlicht: (2025)
Reducing Data Bottlenecks in Distributed, Heterogeneous Neural Networks
von: Lin, Ruhai, et al.
Veröffentlicht: (2024)
von: Lin, Ruhai, et al.
Veröffentlicht: (2024)
Trackable Agent-based Evolution Models at Wafer Scale
von: Moreno, Matthew Andres, et al.
Veröffentlicht: (2024)
von: Moreno, Matthew Andres, et al.
Veröffentlicht: (2024)
Trackable Island-model Genetic Algorithms at Wafer Scale
von: Moreno, Matthew Andres, et al.
Veröffentlicht: (2024)
von: Moreno, Matthew Andres, et al.
Veröffentlicht: (2024)
Sparse Spiking Neural-like Membrane Systems on Graphics Processing Units
von: Hernández-Tello, Javier, et al.
Veröffentlicht: (2024)
von: Hernández-Tello, Javier, et al.
Veröffentlicht: (2024)
The DEEP-ER project: I/O and resiliency extensions for the Cluster-Booster architecture
von: Kreuzer, Anke, et al.
Veröffentlicht: (2019)
von: Kreuzer, Anke, et al.
Veröffentlicht: (2019)
An Evaluation and Comparison of GPU Hardware and Solver Libraries for Accelerating the OPM Flow Reservoir Simulator
von: Qiu, Tong Dong, et al.
Veröffentlicht: (2023)
von: Qiu, Tong Dong, et al.
Veröffentlicht: (2023)
SLIM: A Heterogeneous Accelerator for Edge Inference of Sparse Large Language Model via Adaptive Thresholding
von: Xu, Weihong, et al.
Veröffentlicht: (2025)
von: Xu, Weihong, et al.
Veröffentlicht: (2025)
COMET: A Framework for Modeling Compound Operation Dataflows with Explicit Collectives
von: Negi, Shubham, et al.
Veröffentlicht: (2025)
von: Negi, Shubham, et al.
Veröffentlicht: (2025)
MVDRAM: Enabling GeMV Execution in Unmodified DRAM for Low-Bit LLM Acceleration
von: Kubo, Tatsuya, et al.
Veröffentlicht: (2025)
von: Kubo, Tatsuya, et al.
Veröffentlicht: (2025)
PAM: Processing Across Memory Hierarchy for Efficient KV-centric LLM Serving System
von: Liu, Lian, et al.
Veröffentlicht: (2026)
von: Liu, Lian, et al.
Veröffentlicht: (2026)
RAPID-Graph: Recursive All-Pairs Shortest Paths Using Processing-in-Memory for Dynamic Programming on Graphs
von: Chen, Yanru, et al.
Veröffentlicht: (2025)
von: Chen, Yanru, et al.
Veröffentlicht: (2025)
Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems
von: Zhang, Chen, et al.
Veröffentlicht: (2026)
von: Zhang, Chen, et al.
Veröffentlicht: (2026)
Efficient deadlock avoidance for 2D mesh NoCs that use OQ or VOQ routers
von: Papaphilippou, Philippos, et al.
Veröffentlicht: (2023)
von: Papaphilippou, Philippos, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Neuromorphic Computing: A Theoretical Framework for Time, Space, and Energy Scaling
von: Aimone, James B
Veröffentlicht: (2025) -
Wireless Sensor Networks as Parallel and Distributed Hardware Platform for Artificial Neural Networks
von: Serpen, Gursel
Veröffentlicht: (2025) -
NeuroRing: Scaling Spiking Neural Networks via Multi-FPGA Bidirectional Ring Topologies and Stream-Dataflow Architectures
von: Hafiz, Muhammad Ihsan Al, et al.
Veröffentlicht: (2026) -
MAC-DO: An Efficient Output-Stationary GEMM Accelerator for CNNs Using DRAM Technology
von: Jeong, Minki, et al.
Veröffentlicht: (2022) -
NeuraChip: Accelerating GNN Computations with a Hash-based Decoupled Spatial Accelerator
von: Shivdikar, Kaustubh, et al.
Veröffentlicht: (2024)