Assessing the Performance of Analog Training for Transfer Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Fagbohungbe, Omobayode, Lammie, Corey, Rasch, Malte J., Ando, Takashi, Gokmen, Tayfun, Narayanan, Vijay |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Neuromorphic Computing: A Theoretical Framework for Time, Space, and Energy Scaling
por: Aimone, James B
Publicado: (2025)
por: Aimone, James B
Publicado: (2025)
Wireless Sensor Networks as Parallel and Distributed Hardware Platform for Artificial Neural Networks
por: Serpen, Gursel
Publicado: (2025)
por: Serpen, Gursel
Publicado: (2025)
NeuroRing: Scaling Spiking Neural Networks via Multi-FPGA Bidirectional Ring Topologies and Stream-Dataflow Architectures
por: Hafiz, Muhammad Ihsan Al, et al.
Publicado: (2026)
por: Hafiz, Muhammad Ihsan Al, et al.
Publicado: (2026)
MAC-DO: An Efficient Output-Stationary GEMM Accelerator for CNNs Using DRAM Technology
por: Jeong, Minki, et al.
Publicado: (2022)
por: Jeong, Minki, et al.
Publicado: (2022)
NeuraChip: Accelerating GNN Computations with a Hash-based Decoupled Spatial Accelerator
por: Shivdikar, Kaustubh, et al.
Publicado: (2024)
por: Shivdikar, Kaustubh, et al.
Publicado: (2024)
STEMS: Spatial-Temporal Mapping For Spiking Neural Networks
por: Eissa, Sherif, et al.
Publicado: (2025)
por: Eissa, Sherif, et al.
Publicado: (2025)
EDEA: Efficient Dual-Engine Accelerator for Depthwise Separable Convolution with Direct Data Transfer
por: Chen, Yi, et al.
Publicado: (2025)
por: Chen, Yi, et al.
Publicado: (2025)
BlockAMC: Scalable In-Memory Analog Matrix Computing for Solving Linear Systems
por: Pan, Lunshuai, et al.
Publicado: (2024)
por: Pan, Lunshuai, et al.
Publicado: (2024)
OpenRASE: Service Function Chain Emulation
por: Krishnamohan, Theviyanthan, et al.
Publicado: (2025)
por: Krishnamohan, Theviyanthan, et al.
Publicado: (2025)
Towards a Decentralised Application-Centric Orchestration Framework in the Cloud-Edge Continuum
por: Ullah, Amjad, et al.
Publicado: (2025)
por: Ullah, Amjad, et al.
Publicado: (2025)
Workload-Aware Hardware Accelerator Mining for Distributed Deep Learning Training
por: Adnan, Muhammad, et al.
Publicado: (2024)
por: Adnan, Muhammad, et al.
Publicado: (2024)
SSDTrain: An Activation Offloading Framework to SSDs for Faster Large Language Model Training
por: Wu, Kun, et al.
Publicado: (2024)
por: Wu, Kun, et al.
Publicado: (2024)
Optimizing Offload Performance in Heterogeneous MPSoCs
por: Colagrande, Luca, et al.
Publicado: (2024)
por: Colagrande, Luca, et al.
Publicado: (2024)
Chopper: A Multi-Level GPU Characterization Tool & Derived Insights Into LLM Training Inefficiency
por: Kurzynski, Marco, et al.
Publicado: (2025)
por: Kurzynski, Marco, et al.
Publicado: (2025)
A Fresh Approach to Evaluate Performance in Distributed Parallel Genetic Algorithms
por: Harada, Tomohiro, et al.
Publicado: (2021)
por: Harada, Tomohiro, et al.
Publicado: (2021)
iHAC: A Hybrid Cluster Architecture for Enhanced Performance and Resilience
por: Muntaka, Siddique Abubakr, et al.
Publicado: (2026)
por: Muntaka, Siddique Abubakr, et al.
Publicado: (2026)
Microbenchmark-Driven Analytical Performance Modeling Across Modern GPU Architectures
por: Jarmusch, Aaron, et al.
Publicado: (2026)
por: Jarmusch, Aaron, et al.
Publicado: (2026)
TT-Edge: A Hardware-Software Co-Design for Energy-Efficient Tensor-Train Decomposition on Edge AI
por: Kwak, Hyunseok, et al.
Publicado: (2025)
por: Kwak, Hyunseok, et al.
Publicado: (2025)
Multi-Partner Project: Multi-GPU Performance Portability Analysis for CFD Simulations at Scale
por: Eleftherakis, Panagiotis-Eleftherios, et al.
Publicado: (2026)
por: Eleftherakis, Panagiotis-Eleftherios, et al.
Publicado: (2026)
PULSAR: Simultaneous Many-Row Activation for Reliable and High-Performance Computing in Off-the-Shelf DRAM Chips
por: Yuksel, Ismail Emir, et al.
Publicado: (2023)
por: Yuksel, Ismail Emir, et al.
Publicado: (2023)
Proteus: Enabling High-Performance Processing-Using-DRAM with Dynamic Bit-Precision, Adaptive Data Representation, and Flexible Arithmetic
por: Oliveira, Geraldo F., et al.
Publicado: (2025)
por: Oliveira, Geraldo F., et al.
Publicado: (2025)
ALPHA-PIM: Analysis of Linear Algebraic Processing for High-Performance Graph Applications on a Real Processing-In-Memory System
por: Barkhordar, Marzieh, et al.
Publicado: (2026)
por: Barkhordar, Marzieh, et al.
Publicado: (2026)
D-CODE: Data Colony Optimization for Dynamic Network Efficiency
por: Pandey, Tannu, et al.
Publicado: (2024)
por: Pandey, Tannu, et al.
Publicado: (2024)
A Frequency-based Parent Selection for Reducing the Effect of Evaluation Time Bias in Asynchronous Parallel Multi-objective Evolutionary Algorithms
por: Harada, Tomohiro
Publicado: (2021)
por: Harada, Tomohiro
Publicado: (2021)
Code generation and runtime techniques for enabling data-efficient deep learning training on GPUs
por: Wu, Kun
Publicado: (2024)
por: Wu, Kun
Publicado: (2024)
Neuromorphic Simulation of Drosophila Melanogaster Brain Connectome on Loihi 2
por: Wang, Felix, et al.
Publicado: (2025)
por: Wang, Felix, et al.
Publicado: (2025)
Evaluation and Efficiency Comparison of Evolutionary Algorithms for Service Placement Optimization in Fog Architectures
por: Guerrero, Carlos, et al.
Publicado: (2025)
por: Guerrero, Carlos, et al.
Publicado: (2025)
Reducing Data Bottlenecks in Distributed, Heterogeneous Neural Networks
por: Lin, Ruhai, et al.
Publicado: (2024)
por: Lin, Ruhai, et al.
Publicado: (2024)
Trackable Agent-based Evolution Models at Wafer Scale
por: Moreno, Matthew Andres, et al.
Publicado: (2024)
por: Moreno, Matthew Andres, et al.
Publicado: (2024)
Trackable Island-model Genetic Algorithms at Wafer Scale
por: Moreno, Matthew Andres, et al.
Publicado: (2024)
por: Moreno, Matthew Andres, et al.
Publicado: (2024)
Sparse Spiking Neural-like Membrane Systems on Graphics Processing Units
por: Hernández-Tello, Javier, et al.
Publicado: (2024)
por: Hernández-Tello, Javier, et al.
Publicado: (2024)
The DEEP-ER project: I/O and resiliency extensions for the Cluster-Booster architecture
por: Kreuzer, Anke, et al.
Publicado: (2019)
por: Kreuzer, Anke, et al.
Publicado: (2019)
An Evaluation and Comparison of GPU Hardware and Solver Libraries for Accelerating the OPM Flow Reservoir Simulator
por: Qiu, Tong Dong, et al.
Publicado: (2023)
por: Qiu, Tong Dong, et al.
Publicado: (2023)
SLIM: A Heterogeneous Accelerator for Edge Inference of Sparse Large Language Model via Adaptive Thresholding
por: Xu, Weihong, et al.
Publicado: (2025)
por: Xu, Weihong, et al.
Publicado: (2025)
COMET: A Framework for Modeling Compound Operation Dataflows with Explicit Collectives
por: Negi, Shubham, et al.
Publicado: (2025)
por: Negi, Shubham, et al.
Publicado: (2025)
MVDRAM: Enabling GeMV Execution in Unmodified DRAM for Low-Bit LLM Acceleration
por: Kubo, Tatsuya, et al.
Publicado: (2025)
por: Kubo, Tatsuya, et al.
Publicado: (2025)
PAM: Processing Across Memory Hierarchy for Efficient KV-centric LLM Serving System
por: Liu, Lian, et al.
Publicado: (2026)
por: Liu, Lian, et al.
Publicado: (2026)
RAPID-Graph: Recursive All-Pairs Shortest Paths Using Processing-in-Memory for Dynamic Programming on Graphs
por: Chen, Yanru, et al.
Publicado: (2025)
por: Chen, Yanru, et al.
Publicado: (2025)
Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems
por: Zhang, Chen, et al.
Publicado: (2026)
por: Zhang, Chen, et al.
Publicado: (2026)
Efficient deadlock avoidance for 2D mesh NoCs that use OQ or VOQ routers
por: Papaphilippou, Philippos, et al.
Publicado: (2023)
por: Papaphilippou, Philippos, et al.
Publicado: (2023)
Ejemplares similares
-
Neuromorphic Computing: A Theoretical Framework for Time, Space, and Energy Scaling
por: Aimone, James B
Publicado: (2025) -
Wireless Sensor Networks as Parallel and Distributed Hardware Platform for Artificial Neural Networks
por: Serpen, Gursel
Publicado: (2025) -
NeuroRing: Scaling Spiking Neural Networks via Multi-FPGA Bidirectional Ring Topologies and Stream-Dataflow Architectures
por: Hafiz, Muhammad Ihsan Al, et al.
Publicado: (2026) -
MAC-DO: An Efficient Output-Stationary GEMM Accelerator for CNNs Using DRAM Technology
por: Jeong, Minki, et al.
Publicado: (2022) -
NeuraChip: Accelerating GNN Computations with a Hash-based Decoupled Spatial Accelerator
por: Shivdikar, Kaustubh, et al.
Publicado: (2024)