Saved in:
| Main Authors: | Jiang, Zhuolun, Wang, Songyue, Pei, Xiaokun, Lu, Tianyue, Chen, Mingyu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2511.14990 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Asynchronous Memory Access Unit: Exploiting Massive Parallelism for Far Memory Access
by: Wang, Luming, et al.
Published: (2024)
by: Wang, Luming, et al.
Published: (2024)
Trimma: Trimming Metadata Storage and Latency for Hybrid Memory Systems
by: Li, Yiwei, et al.
Published: (2024)
by: Li, Yiwei, et al.
Published: (2024)
DASICS White Paper: Enhancing Memory Protection with Dynamic Compartmentalization
by: Jin, Yue, et al.
Published: (2023)
by: Jin, Yue, et al.
Published: (2023)
BARD: Reducing Write Latency of DDR5 Memory by Exploiting Bank-Parallelism
by: Vittal, Suhas, et al.
Published: (2025)
by: Vittal, Suhas, et al.
Published: (2025)
TLV-HGNN: Thinking Like a Vertex for Memory-efficient HGNN Inference
by: Han, Dengke, et al.
Published: (2025)
by: Han, Dengke, et al.
Published: (2025)
GDR-HGNN: A Heterogeneous Graph Neural Networks Accelerator Frontend with Graph Decoupling and Recoupling
by: Xue, Runzhen, et al.
Published: (2024)
by: Xue, Runzhen, et al.
Published: (2024)
Ironman: Accelerating Oblivious Transfer Extension for Privacy-Preserving AI with Near-Memory Processing
by: Lin, Chenqi, et al.
Published: (2025)
by: Lin, Chenqi, et al.
Published: (2025)
MERE: Hardware-Software Co-Design for Masking Cache Miss Latency in Embedded Processors
by: You, Dean, et al.
Published: (2025)
by: You, Dean, et al.
Published: (2025)
The Case for Replication-Aware Memory-Error Protection in Disaggregated Memory
by: Volos, Haris
Published: (2023)
by: Volos, Haris
Published: (2023)
A System Architecture for Low Latency Multiprogramming Quantum Computing
by: Zhao, Yilun, et al.
Published: (2026)
by: Zhao, Yilun, et al.
Published: (2026)
Sensitivity-Aware Mixed-Precision Quantization for ReRAM-based Computing-in-Memory
by: Chen, Guan-Cheng, et al.
Published: (2025)
by: Chen, Guan-Cheng, et al.
Published: (2025)
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge
by: Tian, Chunlin, et al.
Published: (2025)
by: Tian, Chunlin, et al.
Published: (2025)
Critical Path Aware Timing-Driven Global Placement for Large-Scale Heterogeneous FPGAs
by: Jiang, He, et al.
Published: (2025)
by: Jiang, He, et al.
Published: (2025)
FusionCIM: Accelerating LLM Inference with Fusion-Driven Computing-in-Memory Architecture
by: Xuan, Zihao, et al.
Published: (2026)
by: Xuan, Zihao, et al.
Published: (2026)
DAE4HLS: Exposing Memory-Level Parallelism for High-Level Synthesis using Explicit Decoupling
by: Metz, David, et al.
Published: (2026)
by: Metz, David, et al.
Published: (2026)
Reimagining Memory Access for LLM Inference: Compression-Aware Memory Controller Design
by: Xie, Rui, et al.
Published: (2025)
by: Xie, Rui, et al.
Published: (2025)
Accelerating GNN Training through Locality-aware Dropout and Merge
by: Sun, Gongjian, et al.
Published: (2025)
by: Sun, Gongjian, et al.
Published: (2025)
EONSim: An NPU Simulator for On-Chip Memory and Embedding Vector Operations
by: Choi, Sangun, et al.
Published: (2025)
by: Choi, Sangun, et al.
Published: (2025)
LMB: Augmenting PCIe Devices with CXL-Linked Memory Buffer
by: Wang, Jiapin, et al.
Published: (2024)
by: Wang, Jiapin, et al.
Published: (2024)
DataMaestro: A Versatile and Efficient Data Streaming Engine Bringing Decoupled Memory Access To Dataflow Accelerators
by: Yi, Xiaoling, et al.
Published: (2025)
by: Yi, Xiaoling, et al.
Published: (2025)
Computing-In-Memory Aware Model Adaption For Edge Devices
by: Lin, Ming-Han, et al.
Published: (2025)
by: Lin, Ming-Han, et al.
Published: (2025)
ADE-HGNN: Accelerating HGNNs through Attention Disparity Exploitation
by: Han, Dengke, et al.
Published: (2024)
by: Han, Dengke, et al.
Published: (2024)
AssertGen: Enhancement of LLM-aided Assertion Generation through Cross-Layer Signal Bridging
by: Lyu, Hongqin, et al.
Published: (2025)
by: Lyu, Hongqin, et al.
Published: (2025)
Customized Retrieval Augmented Generation and Benchmarking for EDA Tool Documentation QA
by: Pu, Yuan, et al.
Published: (2024)
by: Pu, Yuan, et al.
Published: (2024)
A Scalable FPGA Architecture With Adaptive Memory Utilization for GEMM-Based Operations
by: Petropoulos, Anastasios, et al.
Published: (2025)
by: Petropoulos, Anastasios, et al.
Published: (2025)
In-Memory Computing Enabled Deep MIMO Detection to Support Ultra-Low-Latency Communications
by: Ding, Tingyu, et al.
Published: (2025)
by: Ding, Tingyu, et al.
Published: (2025)
SynDCIM: A Performance-Aware Digital Computing-in-Memory Compiler with Multi-Spec-Oriented Subcircuit Synthesis
by: Shao, Kunming, et al.
Published: (2024)
by: Shao, Kunming, et al.
Published: (2024)
PIM-GPT: A Hybrid Process-in-Memory Accelerator for Autoregressive Transformers
by: Wu, Yuting, et al.
Published: (2023)
by: Wu, Yuting, et al.
Published: (2023)
AutoPower: Automated Few-Shot Architecture-Level Power Modeling by Power Group Decoupling
by: Zhang, Qijun, et al.
Published: (2025)
by: Zhang, Qijun, et al.
Published: (2025)
An Event-Driven Spiking Compute-In-Memory Macro based on SOT-MRAM
by: Yu, Deyang, et al.
Published: (2025)
by: Yu, Deyang, et al.
Published: (2025)
CoverAssert: Iterative LLM Assertion Generation Driven by Functional Coverage via Syntax-Semantic Representations
by: Wang, Yonghao, et al.
Published: (2026)
by: Wang, Yonghao, et al.
Published: (2026)
AgileWatts: An Energy-Efficient CPU Core Idle-State Architecture for Latency-Sensitive Server Applications
by: Yahya, Jawad Haj, et al.
Published: (2022)
by: Yahya, Jawad Haj, et al.
Published: (2022)
MIREDO: MIP-Driven Resource-Efficient Dataflow Optimization for Computing-in-Memory Accelerator
by: He, Xiaolin, et al.
Published: (2025)
by: He, Xiaolin, et al.
Published: (2025)
HiHGNN: Accelerating HGNNs through Parallelism and Data Reusability Exploitation
by: Xue, Runzhen, et al.
Published: (2023)
by: Xue, Runzhen, et al.
Published: (2023)
HSCO-Bench: An Agent-Driven End-to-End Hardware-Software Co-design Benchmark for Systems-on-Chip
by: Tsai, Pei-Huan, et al.
Published: (2026)
by: Tsai, Pei-Huan, et al.
Published: (2026)
PolyLUT: Ultra-low Latency Polynomial Inference with Hardware-Aware Structured Pruning
by: Andronic, Marta, et al.
Published: (2025)
by: Andronic, Marta, et al.
Published: (2025)
Fast Cross-Operator Optimization of Attention Dataflow
by: Chang, Haodong, et al.
Published: (2026)
by: Chang, Haodong, et al.
Published: (2026)
CXL Topology-Aware and Expander-Driven Prefetching: Unlocking SSD Performance
by: Oh, Dongsuk, et al.
Published: (2025)
by: Oh, Dongsuk, et al.
Published: (2025)
ADOR: A Design Exploration Framework for LLM Serving with Enhanced Latency and Throughput
by: Kim, Junsoo, et al.
Published: (2025)
by: Kim, Junsoo, et al.
Published: (2025)
Allspark: Workload Orchestration for Visual Transformers on Processing In-Memory Systems
by: Ge, Mengke, et al.
Published: (2024)
by: Ge, Mengke, et al.
Published: (2024)
Similar Items
-
Asynchronous Memory Access Unit: Exploiting Massive Parallelism for Far Memory Access
by: Wang, Luming, et al.
Published: (2024) -
Trimma: Trimming Metadata Storage and Latency for Hybrid Memory Systems
by: Li, Yiwei, et al.
Published: (2024) -
DASICS White Paper: Enhancing Memory Protection with Dynamic Compartmentalization
by: Jin, Yue, et al.
Published: (2023) -
BARD: Reducing Write Latency of DDR5 Memory by Exploiting Bank-Parallelism
by: Vittal, Suhas, et al.
Published: (2025) -
TLV-HGNN: Thinking Like a Vertex for Memory-efficient HGNN Inference
by: Han, Dengke, et al.
Published: (2025)