CiMBA: Accelerating Genome Sequencing through On-Device Basecalling via Compute-in-Memory

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Simon, William Andrew, Boybat, Irem, Kodra, Riselda, Ferro, Elena, Singh, Gagandeep, Alser, Mohammed, Jain, Shubham, Tsai, Hsinyu, Burr, Geoffrey W., Mutlu, Onur, Sebastian, Abu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912319414468608
author Simon, William Andrew
Boybat, Irem
Kodra, Riselda
Ferro, Elena
Singh, Gagandeep
Alser, Mohammed
Jain, Shubham
Tsai, Hsinyu
Burr, Geoffrey W.
Mutlu, Onur
Sebastian, Abu
author_facet Simon, William Andrew
Boybat, Irem
Kodra, Riselda
Ferro, Elena
Singh, Gagandeep
Alser, Mohammed
Jain, Shubham
Tsai, Hsinyu
Burr, Geoffrey W.
Mutlu, Onur
Sebastian, Abu
contents As genome sequencing is finding utility in a wide variety of domains beyond the confines of traditional medical settings, its computational pipeline faces two significant challenges. First, the creation of up to 0.5 GB of data per minute imposes substantial communication and storage overheads. Second, the sequencing pipeline is bottlenecked at the basecalling step, consuming >40% of genome analysis time. A range of proposals have attempted to address these challenges, with limited success. We propose to address these challenges with a Compute-in-Memory Basecalling Accelerator (CiMBA), the first embedded ($\sim25$mm$^2$) accelerator capable of real-time, on-device basecalling, coupled with AnaLog (AL)-Dorado, a new family of analog focused basecalling DNNs. Our resulting hardware/software co-design greatly reduces data communication overhead, is capable of a throughput of 4.77 million bases per second, 24x that required for real-time operation, and achieves 17x/27x power/area efficiency over the best prior basecalling embedded accelerator while maintaining a high accuracy comparable to state-of-the-art software basecallers.
format Preprint
id arxiv_https___arxiv_org_abs_2504_07298
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CiMBA: Accelerating Genome Sequencing through On-Device Basecalling via Compute-in-Memory
Simon, William Andrew
Boybat, Irem
Kodra, Riselda
Ferro, Elena
Singh, Gagandeep
Alser, Mohammed
Jain, Shubham
Tsai, Hsinyu
Burr, Geoffrey W.
Mutlu, Onur
Sebastian, Abu
Hardware Architecture
Genomics
As genome sequencing is finding utility in a wide variety of domains beyond the confines of traditional medical settings, its computational pipeline faces two significant challenges. First, the creation of up to 0.5 GB of data per minute imposes substantial communication and storage overheads. Second, the sequencing pipeline is bottlenecked at the basecalling step, consuming >40% of genome analysis time. A range of proposals have attempted to address these challenges, with limited success. We propose to address these challenges with a Compute-in-Memory Basecalling Accelerator (CiMBA), the first embedded ($\sim25$mm$^2$) accelerator capable of real-time, on-device basecalling, coupled with AnaLog (AL)-Dorado, a new family of analog focused basecalling DNNs. Our resulting hardware/software co-design greatly reduces data communication overhead, is capable of a throughput of 4.77 million bases per second, 24x that required for real-time operation, and achieves 17x/27x power/area efficiency over the best prior basecalling embedded accelerator while maintaining a high accuracy comparable to state-of-the-art software basecallers.
title CiMBA: Accelerating Genome Sequencing through On-Device Basecalling via Compute-in-Memory
topic Hardware Architecture
Genomics
url https://arxiv.org/abs/2504.07298