Gespeichert in:
| 1. Verfasser: | Singh, Shubham Kumar |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2603.10032 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Hybrid JIT-CUDA Graph Optimization for Low-Latency Large Language Model Inference
von: Yadav, Divakar Kumar, et al.
Veröffentlicht: (2026)
von: Yadav, Divakar Kumar, et al.
Veröffentlicht: (2026)
MEMHD: Memory-Efficient Multi-Centroid Hyperdimensional Computing for Fully-Utilized In-Memory Computing Architectures
von: Kang, Do Yeong, et al.
Veröffentlicht: (2025)
von: Kang, Do Yeong, et al.
Veröffentlicht: (2025)
Hierarchical Source-to-Post-Route QoR Prediction in High-Level Synthesis with GNNs
von: Gao, Mingzhe, et al.
Veröffentlicht: (2024)
von: Gao, Mingzhe, et al.
Veröffentlicht: (2024)
AQPIM: Breaking the PIM Capacity Wall for LLMs with In-Memory Activation Quantization
von: Matsushima, Kosuke, et al.
Veröffentlicht: (2026)
von: Matsushima, Kosuke, et al.
Veröffentlicht: (2026)
Efficient Calibration for RRAM-based In-Memory Computing using DoRA
von: Dong, Weirong, et al.
Veröffentlicht: (2025)
von: Dong, Weirong, et al.
Veröffentlicht: (2025)
SwiftRL: Towards Efficient Reinforcement Learning on Real Processing-In-Memory Systems
von: Gogineni, Kailash, et al.
Veröffentlicht: (2024)
von: Gogineni, Kailash, et al.
Veröffentlicht: (2024)
MEADOW: Memory-efficient Dataflow and Data Packing for Low Power Edge LLMs
von: Moitra, Abhishek, et al.
Veröffentlicht: (2025)
von: Moitra, Abhishek, et al.
Veröffentlicht: (2025)
Column-wise Quantization of Weights and Partial Sums for Accurate and Efficient Compute-In-Memory Accelerators
von: Kim, Jiyoon, et al.
Veröffentlicht: (2025)
von: Kim, Jiyoon, et al.
Veröffentlicht: (2025)
NeFT: Negative Feedback Training to Improve Robustness of Compute-In-Memory DNN Accelerators
von: Qin, Yifan, et al.
Veröffentlicht: (2023)
von: Qin, Yifan, et al.
Veröffentlicht: (2023)
Differentiable Initialization-Accelerated CPU-GPU Hybrid Combinatorial Scheduling
von: Liu, Mingju, et al.
Veröffentlicht: (2026)
von: Liu, Mingju, et al.
Veröffentlicht: (2026)
Evaluating CUDA Tile for AI Workloads on Hopper and Blackwell GPUs
von: Yadav, Divakar Kumar, et al.
Veröffentlicht: (2026)
von: Yadav, Divakar Kumar, et al.
Veröffentlicht: (2026)
FlexLLM: Composable HLS Library for Flexible Hybrid LLM Accelerator Design
von: Zhang, Jiahao, et al.
Veröffentlicht: (2026)
von: Zhang, Jiahao, et al.
Veröffentlicht: (2026)
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator
von: Ramachandran, Akshat, et al.
Veröffentlicht: (2025)
von: Ramachandran, Akshat, et al.
Veröffentlicht: (2025)
HPD: Hybrid Projection Decomposition for Robust State Space Models on Analog CIM Hardware
von: Feng, Yuannuo, et al.
Veröffentlicht: (2025)
von: Feng, Yuannuo, et al.
Veröffentlicht: (2025)
AttentionLego: An Open-Source Building Block For Spatially-Scalable Large Language Model Accelerator With Processing-In-Memory Technology
von: Cong, Rongqing, et al.
Veröffentlicht: (2024)
von: Cong, Rongqing, et al.
Veröffentlicht: (2024)
NeuroSim V1.5: Improved Software Backbone for Benchmarking Compute-in-Memory Accelerators with Device and Circuit-level Non-idealities
von: Read, James, et al.
Veröffentlicht: (2025)
von: Read, James, et al.
Veröffentlicht: (2025)
A Hybrid Edge Classifier: Combining TinyML-Optimised CNN with RRAM-CMOS ACAM for Energy-Efficient Inference
von: Woodward, Kieran, et al.
Veröffentlicht: (2025)
von: Woodward, Kieran, et al.
Veröffentlicht: (2025)
HALO: Memory-Centric Heterogeneous Accelerator with 2.5D Integration for Low-Batch LLM Inference
von: Negi, Shubham, et al.
Veröffentlicht: (2025)
von: Negi, Shubham, et al.
Veröffentlicht: (2025)
VerilogDB: The Largest, Highest-Quality Dataset with a Preprocessing Framework for LLM-based RTL Generation
von: Calzada, Paul E., et al.
Veröffentlicht: (2025)
von: Calzada, Paul E., et al.
Veröffentlicht: (2025)
PGR-DRC: Pre-Global Routing DRC Violation Prediction Using Unsupervised Learning
von: Islam, Riadul, et al.
Veröffentlicht: (2025)
von: Islam, Riadul, et al.
Veröffentlicht: (2025)
A Unified Memory Perspective for Probabilistic Trustworthy AI
von: Zhao, Xueji, et al.
Veröffentlicht: (2026)
von: Zhao, Xueji, et al.
Veröffentlicht: (2026)
In-Memory Learning Automata Architecture using Y-Flash Cell
von: Ghazal, Omar, et al.
Veröffentlicht: (2024)
von: Ghazal, Omar, et al.
Veröffentlicht: (2024)
Sorted Weight Sectioning for Energy-Efficient Unstructured Sparse DNNs on Compute-in-Memory Crossbars
von: Farias, Matheus, et al.
Veröffentlicht: (2024)
von: Farias, Matheus, et al.
Veröffentlicht: (2024)
Causal AI For AMS Circuit Design: Interpretable Parameter Effects Analysis
von: Hussain, Mohyeu, et al.
Veröffentlicht: (2026)
von: Hussain, Mohyeu, et al.
Veröffentlicht: (2026)
Design Rules for Extreme-Edge Scientific Computing on AI Engines
von: Ma, Zhenghua, et al.
Veröffentlicht: (2026)
von: Ma, Zhenghua, et al.
Veröffentlicht: (2026)
Agent Factories for High Level Synthesis: How Far Can General-Purpose Coding Agents Go in Hardware Optimization?
von: Bhandwaldar, Abhishek, et al.
Veröffentlicht: (2026)
von: Bhandwaldar, Abhishek, et al.
Veröffentlicht: (2026)
TRAM: Training Approximate Multiplier Structures for Low-Power AI Accelerators
von: Meng, Chang, et al.
Veröffentlicht: (2026)
von: Meng, Chang, et al.
Veröffentlicht: (2026)
Position Paper: From Edge AI to Adaptive Edge AI
von: Pittorino, Fabrizio, et al.
Veröffentlicht: (2026)
von: Pittorino, Fabrizio, et al.
Veröffentlicht: (2026)
Graph Computation Meets Circuit Algebra: A Task-Aligned Analysis of Graph Neural Networks for Electronic Design Automation
von: Kim, Hyunmog
Veröffentlicht: (2026)
von: Kim, Hyunmog
Veröffentlicht: (2026)
CacheMind: From Miss Rates to Why -- Natural-Language, Trace-Grounded Reasoning for Cache Replacement
von: Mhapsekar, Kaushal, et al.
Veröffentlicht: (2026)
von: Mhapsekar, Kaushal, et al.
Veröffentlicht: (2026)
From Fuzzy to Exact: The Halo Architecture for Infinite-Depth Reasoning via Rational Arithmetic
von: Ren, Hansheng
Veröffentlicht: (2026)
von: Ren, Hansheng
Veröffentlicht: (2026)
ALADIN: Accuracy-Latency-Aware Design-space Inference Analysis for Embedded AI Accelerators
von: Baldi, T., et al.
Veröffentlicht: (2026)
von: Baldi, T., et al.
Veröffentlicht: (2026)
Dynamic Sparse Attention: Access Patterns and Architecture
von: Levy, Noam
Veröffentlicht: (2026)
von: Levy, Noam
Veröffentlicht: (2026)
Challenges and Research Directions for Large Language Model Inference Hardware
von: Ma, Xiaoyu, et al.
Veröffentlicht: (2026)
von: Ma, Xiaoyu, et al.
Veröffentlicht: (2026)
Improving the Performance and Learning Stability of Parallelizable RNNs Designed for Ultra-Low Power Applications
von: Brandoit, Julien, et al.
Veröffentlicht: (2026)
von: Brandoit, Julien, et al.
Veröffentlicht: (2026)
Hardware Efficient Approximate Convolution with Tunable Error Tolerance for CNNs
von: Shashidhar, Vishal, et al.
Veröffentlicht: (2026)
von: Shashidhar, Vishal, et al.
Veröffentlicht: (2026)
FASQ: Flexible Accelerated Subspace Quantization for Calibration-Free LLM Compression
von: Qiao, Ye, et al.
Veröffentlicht: (2026)
von: Qiao, Ye, et al.
Veröffentlicht: (2026)
SPARQ: Spiking Early-Exit Neural Networks for Energy-Efficient Edge AI
von: Patne, Parth, et al.
Veröffentlicht: (2026)
von: Patne, Parth, et al.
Veröffentlicht: (2026)
RESQ: A Unified Framework for REliability- and Security Enhancement of Quantized Deep Neural Networks
von: Mohammadi, Ali Soltan, et al.
Veröffentlicht: (2026)
von: Mohammadi, Ali Soltan, et al.
Veröffentlicht: (2026)
Continuous-Flow Data-Rate-Aware CNN Inference on FPGA
von: Habermann, Tobias, et al.
Veröffentlicht: (2026)
von: Habermann, Tobias, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Hybrid JIT-CUDA Graph Optimization for Low-Latency Large Language Model Inference
von: Yadav, Divakar Kumar, et al.
Veröffentlicht: (2026) -
MEMHD: Memory-Efficient Multi-Centroid Hyperdimensional Computing for Fully-Utilized In-Memory Computing Architectures
von: Kang, Do Yeong, et al.
Veröffentlicht: (2025) -
Hierarchical Source-to-Post-Route QoR Prediction in High-Level Synthesis with GNNs
von: Gao, Mingzhe, et al.
Veröffentlicht: (2024) -
AQPIM: Breaking the PIM Capacity Wall for LLMs with In-Memory Activation Quantization
von: Matsushima, Kosuke, et al.
Veröffentlicht: (2026) -
Efficient Calibration for RRAM-based In-Memory Computing using DoRA
von: Dong, Weirong, et al.
Veröffentlicht: (2025)