LLC Intra-set Write Balancing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Krishna, Keshav, Verma, Ayush |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A4: Microarchitecture-Aware LLC Management for Datacenter Servers with Emerging I/O Devices
von: Park, Haneul, et al.
Veröffentlicht: (2025)
von: Park, Haneul, et al.
Veröffentlicht: (2025)
Towards Performance-Aware Allocation for Accelerated Machine Learning on GPU-SSD Systems
von: Gundawar, Ayush, et al.
Veröffentlicht: (2024)
von: Gundawar, Ayush, et al.
Veröffentlicht: (2024)
Design and Analysis of Approximate Hardware Accelerators for VVC Intra Angular Prediction
von: de Fraga, Lucas M. Leipnitz, et al.
Veröffentlicht: (2025)
von: de Fraga, Lucas M. Leipnitz, et al.
Veröffentlicht: (2025)
On the Impact of Intra-node Communication in the Performance of Supercomputer and Data Center Interconnection Networks
von: Tarraga-Moreno, Joaquin, et al.
Veröffentlicht: (2025)
von: Tarraga-Moreno, Joaquin, et al.
Veröffentlicht: (2025)
Scalable and Efficient Intra- and Inter-node Interconnection Networks for Post-Exascale Supercomputers and Data centers
von: Tarraga-Moreno, Joaquin, et al.
Veröffentlicht: (2025)
von: Tarraga-Moreno, Joaquin, et al.
Veröffentlicht: (2025)
Pointer: An Energy-Efficient ReRAM-based Point Cloud Recognition Accelerator with Inter-layer and Intra-layer Optimizations
von: Zhang, Qijun, et al.
Veröffentlicht: (2024)
von: Zhang, Qijun, et al.
Veröffentlicht: (2024)
HARP: Hadamard-Domain Write-and-Verify for Noise-Robust RRAM Programming
von: Choi, Ilhuan, et al.
Veröffentlicht: (2026)
von: Choi, Ilhuan, et al.
Veröffentlicht: (2026)
How to keep pushing ML accelerator performance? Know your rooflines!
von: Verhelst, Marian, et al.
Veröffentlicht: (2025)
von: Verhelst, Marian, et al.
Veröffentlicht: (2025)
An Affordable Experimental Technique for SRAM Write Margin Characterization for Nanometer CMOS Technologies
von: Alorda, Bartomeu, et al.
Veröffentlicht: (2024)
von: Alorda, Bartomeu, et al.
Veröffentlicht: (2024)
BARD: Reducing Write Latency of DDR5 Memory by Exploiting Bank-Parallelism
von: Vittal, Suhas, et al.
Veröffentlicht: (2025)
von: Vittal, Suhas, et al.
Veröffentlicht: (2025)
RCW-CIM: A Digital CIM-based LLM Accelerator with Read-Compute/Write
von: Guo, Yan-Cheng, et al.
Veröffentlicht: (2026)
von: Guo, Yan-Cheng, et al.
Veröffentlicht: (2026)
ML-PCM : Machine Learning Technique for Write Optimization in Phase Change Memory (PCM)
von: Desai, Mahek, et al.
Veröffentlicht: (2025)
von: Desai, Mahek, et al.
Veröffentlicht: (2025)
A Host-SSD Collaborative Write Accelerator for LSM-Tree-Based Key-Value Stores
von: Kim, KiHwan, et al.
Veröffentlicht: (2024)
von: Kim, KiHwan, et al.
Veröffentlicht: (2024)
Limited Read-Write/Set Hardware Transactional Memory without modifying the ISA or the Coherence Protocol
von: Kafousis, Konstantinos
Veröffentlicht: (2025)
von: Kafousis, Konstantinos
Veröffentlicht: (2025)
Nemo: A Low-Write-Amplification Cache for Tiny Objects on Log-Structured Flash Devices
von: Yang, Xufeng, et al.
Veröffentlicht: (2026)
von: Yang, Xufeng, et al.
Veröffentlicht: (2026)
A High-Throughput FPGA Accelerator for Lightweight CNNs With Balanced Dataflow
von: Zhao, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Zhao, Zhiyuan, et al.
Veröffentlicht: (2024)
Striking the Balance: GEMM Performance Optimization Across Generations of Ryzen AI NPUs
von: Taka, Endri, et al.
Veröffentlicht: (2025)
von: Taka, Endri, et al.
Veröffentlicht: (2025)
FEATHER: A Reconfigurable Accelerator with Data Reordering Support for Low-Cost On-Chip Dataflow Switching
von: Tong, Jianming, et al.
Veröffentlicht: (2024)
von: Tong, Jianming, et al.
Veröffentlicht: (2024)
MCMComm: Hardware-Software Co-Optimization for End-to-End Communication in Multi-Chip-Modules
von: Raj, Ritik, et al.
Veröffentlicht: (2025)
von: Raj, Ritik, et al.
Veröffentlicht: (2025)
LEAP: LLM Inference on Scalable PIM-NoC Architecture with Balanced Dataflow and Fine-Grained Parallelism
von: Wang, Yimin, et al.
Veröffentlicht: (2025)
von: Wang, Yimin, et al.
Veröffentlicht: (2025)
CIM-Tuner: Balancing the Compute and Storage Capacity of SRAM-CIM Accelerator via Hardware-mapping Co-exploration
von: Chen, Jinwu, et al.
Veröffentlicht: (2026)
von: Chen, Jinwu, et al.
Veröffentlicht: (2026)
Balancing FP8 Computation Accuracy and Efficiency on Digital CIM via Shift-Aware On-the-fly Aligned-Mantissa Bitwidth Prediction
von: Zhao, Liang, et al.
Veröffentlicht: (2026)
von: Zhao, Liang, et al.
Veröffentlicht: (2026)
PipeOrgan: Efficient Inter-operation Pipelining with Flexible Spatial Organization and Interconnects
von: Garg, Raveesh, et al.
Veröffentlicht: (2024)
von: Garg, Raveesh, et al.
Veröffentlicht: (2024)
Axon: A novel systolic array architecture for improved run time and energy efficient GeMM and Conv operation with on-chip im2col
von: Nayan, Md Mizanur Rahaman, et al.
Veröffentlicht: (2025)
von: Nayan, Md Mizanur Rahaman, et al.
Veröffentlicht: (2025)
ITHICA: Intra-Thread Instruction Checking Approach for Defect-Induced Silent Data Corruptions
von: Vavelidou, Ioanna, et al.
Veröffentlicht: (2026)
von: Vavelidou, Ioanna, et al.
Veröffentlicht: (2026)
CogSys: Efficient and Scalable Neurosymbolic Cognition System via Algorithm-Hardware Co-Design
von: Wan, Zishen, et al.
Veröffentlicht: (2025)
von: Wan, Zishen, et al.
Veröffentlicht: (2025)
Design Conductor: An agent autonomously builds a 1.5 GHz Linux-capable RISC-V CPU
von: The Verkor Team, et al.
Veröffentlicht: (2026)
von: The Verkor Team, et al.
Veröffentlicht: (2026)
Design Conductor 2.0: An agent builds a TurboQuant inference accelerator in 80 hours
von: The Verkor Team, et al.
Veröffentlicht: (2026)
von: The Verkor Team, et al.
Veröffentlicht: (2026)
Cross-Layer Design of Vector-Symbolic Computing: Bridging Cognition and Brain-Inspired Hardware Acceleration
von: Du, Shuting, et al.
Veröffentlicht: (2025)
von: Du, Shuting, et al.
Veröffentlicht: (2025)
OLAF: Programmable Data Plane Acceleration for Asynchronous Distributed Reinforcement Learning
von: Krishna, Nehal Baganal, et al.
Veröffentlicht: (2025)
von: Krishna, Nehal Baganal, et al.
Veröffentlicht: (2025)
SCALE-Sim TPU: Validating and Extending SCALE-Sim for TPUs
von: Dang, Jingtian, et al.
Veröffentlicht: (2026)
von: Dang, Jingtian, et al.
Veröffentlicht: (2026)
U-SWIM: Universal Selective Write-Verify for Computing-in-Memory Neural Accelerators
von: Yan, Zheyu, et al.
Veröffentlicht: (2023)
von: Yan, Zheyu, et al.
Veröffentlicht: (2023)
SMART-WRITE: Adaptive Learning-based Write Energy Optimization for Phase Change Memory
von: Desai, Mahek, et al.
Veröffentlicht: (2025)
von: Desai, Mahek, et al.
Veröffentlicht: (2025)
SCALE-Sim v3: A modular cycle-accurate systolic accelerator simulator for end-to-end system analysis
von: Raj, Ritik, et al.
Veröffentlicht: (2025)
von: Raj, Ritik, et al.
Veröffentlicht: (2025)
A Reconfigurable Multiplier Architecture for Error-Resilient Applications in RISC-V Core
von: Jaswal, Pragun, et al.
Veröffentlicht: (2026)
von: Jaswal, Pragun, et al.
Veröffentlicht: (2026)
Low Power Approximate Multiplier Architecture for Deep Neural Networks
von: Jaswal, Pragun, et al.
Veröffentlicht: (2025)
von: Jaswal, Pragun, et al.
Veröffentlicht: (2025)
Device-Circuit Co-Design of Variation-Resilient Read and Write Drivers for Antiferromagnetic Tunnel Junction (AFMTJ) Memories
von: Choudhary, Yousuf, et al.
Veröffentlicht: (2026)
von: Choudhary, Yousuf, et al.
Veröffentlicht: (2026)
Unsupervised Graph Neural Network Framework for Balanced Multipatterning in Advanced Electronic Design Automation Layouts
von: Helaly, Abdelrahman, et al.
Veröffentlicht: (2025)
von: Helaly, Abdelrahman, et al.
Veröffentlicht: (2025)
EXION: Exploiting Inter- and Intra-Iteration Output Sparsity for Diffusion Models
von: Heo, Jaehoon, et al.
Veröffentlicht: (2025)
von: Heo, Jaehoon, et al.
Veröffentlicht: (2025)
Design high-confidence computers using trusted instructional set architecture and emulators
von: Wang, Shuangbao Paul
Veröffentlicht: (2025)
von: Wang, Shuangbao Paul
Veröffentlicht: (2025)
Ähnliche Einträge
-
A4: Microarchitecture-Aware LLC Management for Datacenter Servers with Emerging I/O Devices
von: Park, Haneul, et al.
Veröffentlicht: (2025) -
Towards Performance-Aware Allocation for Accelerated Machine Learning on GPU-SSD Systems
von: Gundawar, Ayush, et al.
Veröffentlicht: (2024) -
Design and Analysis of Approximate Hardware Accelerators for VVC Intra Angular Prediction
von: de Fraga, Lucas M. Leipnitz, et al.
Veröffentlicht: (2025) -
On the Impact of Intra-node Communication in the Performance of Supercomputer and Data Center Interconnection Networks
von: Tarraga-Moreno, Joaquin, et al.
Veröffentlicht: (2025) -
Scalable and Efficient Intra- and Inter-node Interconnection Networks for Post-Exascale Supercomputers and Data centers
von: Tarraga-Moreno, Joaquin, et al.
Veröffentlicht: (2025)