Dong, P., Tan, Y., Liu, X., Luo, P., Liu, Y., Pang, D., . . . Cheng, K. (2026). 31.1 A 14.08-to-135.69Token/s ReRAM-on-Logic Stacked Outlier-Free Large-Language-Model Accelerator with Block-Clustered Weight-Compression and Adaptive Parallel-Speculative-Decoding.
Chicago Style (17th ed.) CitationDong, Pingcheng, et al. 31.1 A 14.08-to-135.69Token/s ReRAM-on-Logic Stacked Outlier-Free Large-Language-Model Accelerator with Block-Clustered Weight-Compression and Adaptive Parallel-Speculative-Decoding. 2026.
MLA (9th ed.) CitationDong, Pingcheng, et al. 31.1 A 14.08-to-135.69Token/s ReRAM-on-Logic Stacked Outlier-Free Large-Language-Model Accelerator with Block-Clustered Weight-Compression and Adaptive Parallel-Speculative-Decoding. 2026.
Warning: These citations may not always be 100% accurate.