SmartQuant: CXL-based AI Model Store in Support of Runtime Configurable Weight Quantization
Fuente:
arXiv
Saved in:
| Main Authors: | Xie, Rui, Haq, Asad Ul, Ma, Linsen, Sun, Krystal, Sen, Sanchari, Venkataramani, Swagath, Liu, Liu, Zhang, Tong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Breaking the HBM Bit Cost Barrier: Domain-Specific ECC for AI Inference Infrastructure
by: Xie, Rui, et al.
Published: (2025)
by: Xie, Rui, et al.
Published: (2025)
Making Strong Error-Correcting Codes Work Effectively for HBM in AI Inference
by: Xie, Rui, et al.
Published: (2025)
by: Xie, Rui, et al.
Published: (2025)
TRACE: Unlocking Effective CXL Bandwidth via Lossless Compression and Precision Scaling
by: Xie, Rui, et al.
Published: (2025)
by: Xie, Rui, et al.
Published: (2025)
Reimagining Memory Access for LLM Inference: Compression-Aware Memory Controller Design
by: Xie, Rui, et al.
Published: (2025)
by: Xie, Rui, et al.
Published: (2025)
Accelerating LLM Inference via Dynamic KV Cache Placement in Heterogeneous Memory System
by: Fang, Yunhua, et al.
Published: (2025)
by: Fang, Yunhua, et al.
Published: (2025)
The Case for Persistent CXL switches
by: Hadi, Khan Shaikhul, et al.
Published: (2025)
by: Hadi, Khan Shaikhul, et al.
Published: (2025)
CXL-GPU: Pushing GPU Memory Boundaries with the Integration of CXL Technologies
by: Gouk, Donghyun, et al.
Published: (2025)
by: Gouk, Donghyun, et al.
Published: (2025)
SigmaQuant: Hardware-Aware Heterogeneous Quantization Method for Edge DNN Inference
by: Liu, Qunyou, et al.
Published: (2026)
by: Liu, Qunyou, et al.
Published: (2026)
Salient Store: Enabling Smart Storage for Continuous Learning Edge Servers
by: Mishra, Cyan Subhra, et al.
Published: (2024)
by: Mishra, Cyan Subhra, et al.
Published: (2024)
Performance Characterizations and Usage Guidelines of Samsung CXL Memory Module Hybrid Prototype
by: Zeng, Jianping, et al.
Published: (2025)
by: Zeng, Jianping, et al.
Published: (2025)
CXL-Interference: Analysis and Characterization in Modern Computer Systems
by: Mao, Shunyu, et al.
Published: (2024)
by: Mao, Shunyu, et al.
Published: (2024)
Octopus: Enhancing CXL Memory Pods via Sparse Topology
by: Zhong, Yuhong, et al.
Published: (2025)
by: Zhong, Yuhong, et al.
Published: (2025)
LMB: Augmenting PCIe Devices with CXL-Linked Memory Buffer
by: Wang, Jiapin, et al.
Published: (2024)
by: Wang, Jiapin, et al.
Published: (2024)
A Novel Extensible Simulation Framework for CXL-Enabled Systems
by: An, Yuda, et al.
Published: (2024)
by: An, Yuda, et al.
Published: (2024)
Enabling Efficient Transaction Processing on CXL-Based Memory Sharing
by: Wang, Zhao, et al.
Published: (2025)
by: Wang, Zhao, et al.
Published: (2025)
Bitwise Systolic Array Architecture for Runtime-Reconfigurable Multi-precision Quantized Multiplication on Hardware Accelerators
by: Liu, Yuhao, et al.
Published: (2026)
by: Liu, Yuhao, et al.
Published: (2026)
CXL Topology-Aware and Expander-Driven Prefetching: Unlocking SSD Performance
by: Oh, Dongsuk, et al.
Published: (2025)
by: Oh, Dongsuk, et al.
Published: (2025)
Context-Aware Mixture-of-Experts Inference on CXL-Enabled GPU-NDP Systems
by: Fan, Zehao, et al.
Published: (2025)
by: Fan, Zehao, et al.
Published: (2025)
IBEX: Internal Bandwidth-Efficient Compression Architecture for Scalable CXL Memory Expansion
by: Ko, Younghoon, et al.
Published: (2026)
by: Ko, Younghoon, et al.
Published: (2026)
A Full-System Simulation Framework for CXL-Based SSD Memory System
by: Wang, Yaohui, et al.
Published: (2025)
by: Wang, Yaohui, et al.
Published: (2025)
NeoMem: Hardware/Software Co-Design for CXL-Native Memory Tiering
by: Zhou, Zhe, et al.
Published: (2024)
by: Zhou, Zhe, et al.
Published: (2024)
Low-overhead General-purpose Near-Data Processing in CXL Memory Expanders
by: Ham, Hyungkyu, et al.
Published: (2024)
by: Ham, Hyungkyu, et al.
Published: (2024)
Is Finer Better? The Limits of Microscaling Formats in Large Language Models
by: Fasoli, Andrea, et al.
Published: (2026)
by: Fasoli, Andrea, et al.
Published: (2026)
CXL-DMSim: A Full-System CXL Disaggregated Memory Simulator With Comprehensive Silicon Validation
by: Wang, Yanjing, et al.
Published: (2024)
by: Wang, Yanjing, et al.
Published: (2024)
Pushing the Memory Bandwidth Wall with CXL-enabled Idle I/O Bandwidth Harvesting
by: Kadiyala, Divya Kiran, et al.
Published: (2025)
by: Kadiyala, Divya Kiran, et al.
Published: (2025)
FPGA-based Emulation and Device-Side Management for CXL-based Memory Tiering Systems
by: Chen, Yiqi, et al.
Published: (2025)
by: Chen, Yiqi, et al.
Published: (2025)
From Block to Byte: Transforming PCIe SSDs with CXL Memory Protocol and Instruction Annotation
by: Kwon, Miryeong, et al.
Published: (2025)
by: Kwon, Miryeong, et al.
Published: (2025)
Cosmos: A CXL-Based Full In-Memory System for Approximate Nearest Neighbor Search
by: Ko, Seoyoung, et al.
Published: (2025)
by: Ko, Seoyoung, et al.
Published: (2025)
A Fully-Configurable Open-Source Software-Defined Digital Quantized Spiking Neural Core Architecture
by: Matinizadeh, Shadi, et al.
Published: (2024)
by: Matinizadeh, Shadi, et al.
Published: (2024)
SkyByte: Architecting an Efficient Memory-Semantic CXL-based SSD with OS and Hardware Co-design
by: Zhang, Haoyang, et al.
Published: (2025)
by: Zhang, Haoyang, et al.
Published: (2025)
Cohet: A CXL-Driven Coherent Heterogeneous Computing Framework with Hardware-Calibrated Full-System Simulation
by: Wang, Yanjing, et al.
Published: (2025)
by: Wang, Yanjing, et al.
Published: (2025)
An Introduction to the Compute Express Link (CXL) Interconnect
by: Sharma, Debendra Das, et al.
Published: (2023)
by: Sharma, Debendra Das, et al.
Published: (2023)
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference
by: Gu, Yufeng, et al.
Published: (2025)
by: Gu, Yufeng, et al.
Published: (2025)
Low Complexity Deep Learning Augmented Wireless Channel Estimation for Pilot-Based OFDM on Zynq System on Chip
by: Sharma, Animesh, et al.
Published: (2024)
by: Sharma, Animesh, et al.
Published: (2024)
Architectural and System Implications of CXL-enabled Tiered Memory
by: Yang, Yujie, et al.
Published: (2025)
by: Yang, Yujie, et al.
Published: (2025)
Binary Weight Multi-Bit Activation Quantization for Compute-in-Memory CNN Accelerators
by: Zhou, Wenyong, et al.
Published: (2025)
by: Zhou, Wenyong, et al.
Published: (2025)
Five-Minute Rule 40 Years Later: A First-Principles Revisit for Modern Memory Hierarchy
by: Zhang, Tong, et al.
Published: (2025)
by: Zhang, Tong, et al.
Published: (2025)
Runtime Energy Monitoring for RISC-V Soft-Cores
by: Scionti, Alberto, et al.
Published: (2025)
by: Scionti, Alberto, et al.
Published: (2025)
CXL-ClusterSim: Modeling CXL-based Disaggregated Memory Cluster for Pooling and Sharing using gem5 and SST
by: Goswami, Kaustav, et al.
Published: (2026)
by: Goswami, Kaustav, et al.
Published: (2026)
Toleo: Scaling Freshness to Tera-scale Memory using CXL and PIM
by: Dong, Juechu, et al.
Published: (2024)
by: Dong, Juechu, et al.
Published: (2024)
Similar Items
-
Breaking the HBM Bit Cost Barrier: Domain-Specific ECC for AI Inference Infrastructure
by: Xie, Rui, et al.
Published: (2025) -
Making Strong Error-Correcting Codes Work Effectively for HBM in AI Inference
by: Xie, Rui, et al.
Published: (2025) -
TRACE: Unlocking Effective CXL Bandwidth via Lossless Compression and Precision Scaling
by: Xie, Rui, et al.
Published: (2025) -
Reimagining Memory Access for LLM Inference: Compression-Aware Memory Controller Design
by: Xie, Rui, et al.
Published: (2025) -
Accelerating LLM Inference via Dynamic KV Cache Placement in Heterogeneous Memory System
by: Fang, Yunhua, et al.
Published: (2025)