MERE: Hardware-Software Co-Design for Masking Cache Miss Latency in Embedded Processors
Fuente:
arXiv
Saved in:
| Main Authors: | You, Dean, Jiang, Jieyu, Wang, Xiaoxuan, Du, Yushu, Tan, Zhihang, Xu, Wenbo, Wang, Hui, Guan, Jiapeng, Wang, Zhenyuan, Wei, Ran, Zhao, Shuai, Jiang, Zhe |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MESC: Re-thinking Algorithmic Priority and/or Criticality Inversions for Heterogeneous MCSs
by: Guan, Jiapeng, et al.
Published: (2024)
by: Guan, Jiapeng, et al.
Published: (2024)
MEEK: Re-thinking Heterogeneous Parallel Error Detection Architecture for Real-World OoO Superscalar Processors
by: Jiang, Zhe, et al.
Published: (2025)
by: Jiang, Zhe, et al.
Published: (2025)
Strix: Re-thinking NPU Reliability from a System Perspective
by: Guan, Jiapeng, et al.
Published: (2026)
by: Guan, Jiapeng, et al.
Published: (2026)
ISAAC: Intelligent, Scalable, Agile, and Accelerated CPU Verification via LLM-aided FPGA Parallelism
by: Sun, Jialin, et al.
Published: (2025)
by: Sun, Jialin, et al.
Published: (2025)
NVR: Vector Runahead on NPUs for Sparse Memory Access
by: Wang, Hui, et al.
Published: (2025)
by: Wang, Hui, et al.
Published: (2025)
From Characterization to Microarchitecture: Designing an Elegant and Reliable BFP-Based NPU
by: Zhang, Jie, et al.
Published: (2026)
by: Zhang, Jie, et al.
Published: (2026)
QiMeng: Fully Automated Hardware and Software Design for Processor Chip
by: Zhang, Rui, et al.
Published: (2025)
by: Zhang, Rui, et al.
Published: (2025)
Hardware/Algorithm Co-design for Real-Time I/O Control with Improved Timing Accuracy and Robustness
by: Jiang, Zhe, et al.
Published: (2024)
by: Jiang, Zhe, et al.
Published: (2024)
METRO: A Software-Hardware Co-Design of Interconnections for Spatial DNN Accelerators
by: Wang, Zhao, et al.
Published: (2021)
by: Wang, Zhao, et al.
Published: (2021)
NeoMem: Hardware/Software Co-Design for CXL-Native Memory Tiering
by: Zhou, Zhe, et al.
Published: (2024)
by: Zhou, Zhe, et al.
Published: (2024)
Harmonia: Algorithm-Hardware Co-Design for Memory- and Compute-Efficient BFP-based LLM Inference
by: Wang, Xinyu, et al.
Published: (2026)
by: Wang, Xinyu, et al.
Published: (2026)
TurboFuzz: FPGA Accelerated Hardware Fuzzing for Processor Agile Verification
by: Zhong, Yang, et al.
Published: (2025)
by: Zhong, Yang, et al.
Published: (2025)
FlexStep: Enabling Flexible Error Detection in Multi/Many-core Real-time Systems
by: Wang, Tinglue, et al.
Published: (2025)
by: Wang, Tinglue, et al.
Published: (2025)
Titanus: Enabling KV Cache Pruning and Quantization On-the-Fly for LLM Acceleration
by: Chen, Peilin, et al.
Published: (2025)
by: Chen, Peilin, et al.
Published: (2025)
Area-Efficient In-Memory Computing for Mixture-of-Experts via Multiplexing and Caching
by: Gao, Hanyuan, et al.
Published: (2026)
by: Gao, Hanyuan, et al.
Published: (2026)
SoftmAP: Software-Hardware Co-design for Integer-Only Softmax on Associative Processors
by: Rakka, Mariam, et al.
Published: (2024)
by: Rakka, Mariam, et al.
Published: (2024)
On Reducing the Execution Latency of Superconducting Quantum Processors via Quantum Job Scheduling
by: Wu, Wenjie, et al.
Published: (2024)
by: Wu, Wenjie, et al.
Published: (2024)
On the Impact of ISA Extension on Energy Consumption of I-Cache in Extensible Processors
by: Behboudi, Noushin, et al.
Published: (2024)
by: Behboudi, Noushin, et al.
Published: (2024)
Lyra: A Hardware-Accelerated RISC-V Verification Framework with Generative Model-Based Processor Fuzzing
by: Huo, Juncheng, et al.
Published: (2025)
by: Huo, Juncheng, et al.
Published: (2025)
Hardware-Software Co-design for 3D-DRAM-based LLM Serving Accelerator
by: Li, Cong, et al.
Published: (2026)
by: Li, Cong, et al.
Published: (2026)
LPU: A Latency-Optimized and Highly Scalable Processor for Large Language Model Inference
by: Moon, Seungjae, et al.
Published: (2024)
by: Moon, Seungjae, et al.
Published: (2024)
Implementation of Compute Intensive Algorithms on Software Configurable Processor
by: Ganesha, et al.
Published: (2025)
by: Ganesha, et al.
Published: (2025)
CIM-Tuner: Balancing the Compute and Storage Capacity of SRAM-CIM Accelerator via Hardware-mapping Co-exploration
by: Chen, Jinwu, et al.
Published: (2026)
by: Chen, Jinwu, et al.
Published: (2026)
Duet: Creating Harmony between Processors and Embedded FPGAs
by: Li, Ang, et al.
Published: (2023)
by: Li, Ang, et al.
Published: (2023)
PCG: Mitigating Conflict-based Cache Side-channel Attacks with Prefetching
by: Jiang, Fang, et al.
Published: (2024)
by: Jiang, Fang, et al.
Published: (2024)
CoroAMU: Unleashing Memory-Driven Coroutines through Latency-Aware Decoupled Operations
by: Jiang, Zhuolun, et al.
Published: (2025)
by: Jiang, Zhuolun, et al.
Published: (2025)
Image processing Application Development on Software Configurable Processor Array
by: Prabhu, Ganesh, et al.
Published: (2025)
by: Prabhu, Ganesh, et al.
Published: (2025)
VeriDebug: A Unified LLM for Verilog Debugging via Contrastive Embedding and Guided Correction
by: Wang, Ning, et al.
Published: (2025)
by: Wang, Ning, et al.
Published: (2025)
Parameterized Hardware Design with Latency-Abstract Interfaces
by: Nigam, Rachit, et al.
Published: (2024)
by: Nigam, Rachit, et al.
Published: (2024)
SeDA: Secure and Efficient DNN Accelerators with Hardware/Software Synergy
by: Xuan, Wei, et al.
Published: (2025)
by: Xuan, Wei, et al.
Published: (2025)
FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs
by: Xie, Xilong, et al.
Published: (2025)
by: Xie, Xilong, et al.
Published: (2025)
AIM: Software and Hardware Co-design for Architecture-level IR-drop Mitigation in High-performance PIM
by: Zhang, Yuanpeng, et al.
Published: (2025)
by: Zhang, Yuanpeng, et al.
Published: (2025)
Large Processor Chip Model
by: Chang, Kaiyan, et al.
Published: (2025)
by: Chang, Kaiyan, et al.
Published: (2025)
VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator
by: Wang, Zhican, et al.
Published: (2025)
by: Wang, Zhican, et al.
Published: (2025)
Large Language Model for Verilog Generation with Code-Structure-Guided Reinforcement Learning
by: Wang, Ning, et al.
Published: (2024)
by: Wang, Ning, et al.
Published: (2024)
Insights from Rights and Wrongs: A Large Language Model for Solving Assertion Failures in RTL Design
by: Zhou, Jie, et al.
Published: (2025)
by: Zhou, Jie, et al.
Published: (2025)
A Joint Learning Approach to Hardware Caching and Prefetching
by: Yuan, Samuel, et al.
Published: (2025)
by: Yuan, Samuel, et al.
Published: (2025)
I/O Transit Caching for PMem-based Block Device
by: Xu, Qing, et al.
Published: (2024)
by: Xu, Qing, et al.
Published: (2024)
aLEAKator: HDL Mixed-Domain Simulation for Masked Hardware \& Software Formal Verification
by: Amiot, Noé, et al.
Published: (2025)
by: Amiot, Noé, et al.
Published: (2025)
BackCache: Mitigating Contention-Based Cache Timing Attacks by Hiding Cache Line Evictions
by: Wang, Quancheng, et al.
Published: (2023)
by: Wang, Quancheng, et al.
Published: (2023)
Similar Items
-
MESC: Re-thinking Algorithmic Priority and/or Criticality Inversions for Heterogeneous MCSs
by: Guan, Jiapeng, et al.
Published: (2024) -
MEEK: Re-thinking Heterogeneous Parallel Error Detection Architecture for Real-World OoO Superscalar Processors
by: Jiang, Zhe, et al.
Published: (2025) -
Strix: Re-thinking NPU Reliability from a System Perspective
by: Guan, Jiapeng, et al.
Published: (2026) -
ISAAC: Intelligent, Scalable, Agile, and Accelerated CPU Verification via LLM-aided FPGA Parallelism
by: Sun, Jialin, et al.
Published: (2025) -
NVR: Vector Runahead on NPUs for Sparse Memory Access
by: Wang, Hui, et al.
Published: (2025)