NeuroScalar: A Deep Learning Framework for Fast, Accurate, and In-the-Wild Cycle-Level Performance Prediction
Fuente:
arXiv
Guardado en:
| Autores principales: | Wadle, Shayne, Zhang, Yanxin, Singh, Vikas, Sankaralingam, Karthikeyan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SAHM: State-Aware Heterogeneous Multicore for Single-Thread Performance
por: Wadle, Shayne, et al.
Publicado: (2025)
por: Wadle, Shayne, et al.
Publicado: (2025)
Beyond Static Policies: Exploring Dynamic Policy Selection for Single-Thread Performance Optimization
por: Zhang, Yanxin, et al.
Publicado: (2026)
por: Zhang, Yanxin, et al.
Publicado: (2026)
Computer Architecture's AlphaZero Moment: Automated Discovery in an Encircled World
por: Sankaralingam, Karthikeyan
Publicado: (2026)
por: Sankaralingam, Karthikeyan
Publicado: (2026)
IPU: Flexible Hardware Introspection Units
por: McDougall, Ian, et al.
Publicado: (2023)
por: McDougall, Ian, et al.
Publicado: (2023)
Privacy-Preserving Performance Profiling of In-The-Wild GPUs
por: McDougall, Ian, et al.
Publicado: (2025)
por: McDougall, Ian, et al.
Publicado: (2025)
LIMINAL: Exploring The Frontiers of LLM Decode Performance
por: Davies, Michael, et al.
Publicado: (2025)
por: Davies, Michael, et al.
Publicado: (2025)
MGS: Markov Greedy Sums for Accurate Low-Bitwidth Floating-Point Accumulation
por: Natesh, Vikas, et al.
Publicado: (2025)
por: Natesh, Vikas, et al.
Publicado: (2025)
Concorde: Fast and Accurate CPU Performance Modeling with Compositional Analytical-ML Fusion
por: Nasr-Esfahany, Arash, et al.
Publicado: (2025)
por: Nasr-Esfahany, Arash, et al.
Publicado: (2025)
The Impact Market to Save Conference Peer Review: Decoupling Dissemination and Credentialing
por: Sankaralingam, Karthikeyan
Publicado: (2025)
por: Sankaralingam, Karthikeyan
Publicado: (2025)
Automatic Generation of Fast and Accurate Performance Models for Deep Neural Network Accelerators
por: Lübeck, Konstantin, et al.
Publicado: (2024)
por: Lübeck, Konstantin, et al.
Publicado: (2024)
Introducing Instruction-Accurate Simulators for Performance Estimation of Autotuning Workloads
por: Pelke, Rebecca, et al.
Publicado: (2025)
por: Pelke, Rebecca, et al.
Publicado: (2025)
Pedagogically Motivated and Composable Open-Source RISC-V Processors for Computer Science Education
por: McDougall, Ian, et al.
Publicado: (2025)
por: McDougall, Ian, et al.
Publicado: (2025)
Deep Inverse Design for High-Level Synthesis
por: Chang, Ping, et al.
Publicado: (2024)
por: Chang, Ping, et al.
Publicado: (2024)
HLSFactory: A Framework Empowering High-Level Synthesis Datasets for Machine Learning and Beyond
por: Abi-Karam, Stefan, et al.
Publicado: (2024)
por: Abi-Karam, Stefan, et al.
Publicado: (2024)
AP-DRL: A Synergistic Algorithm-Hardware Framework for Automatic Task Partitioning of Deep Reinforcement Learning on Versal ACAP
por: Li, Enlai, et al.
Publicado: (2026)
por: Li, Enlai, et al.
Publicado: (2026)
EvolveGen: Algorithmic Level Hardware Model Checking Benchmark Generation through Reinforcement Learning
por: Hu, Guangyu, et al.
Publicado: (2026)
por: Hu, Guangyu, et al.
Publicado: (2026)
TCL: Enabling Fast and Efficient Cross-Hardware Tensor Program Optimization via Continual Learning
por: Shen, Chaoyao, et al.
Publicado: (2026)
por: Shen, Chaoyao, et al.
Publicado: (2026)
RLPlanner: Reinforcement Learning based Floorplanning for Chiplets with Fast Thermal Analysis
por: Duan, Yuanyuan, et al.
Publicado: (2023)
por: Duan, Yuanyuan, et al.
Publicado: (2023)
Energy-Aware Deep Learning on Resource-Constrained Hardware
por: Millar, Josh, et al.
Publicado: (2025)
por: Millar, Josh, et al.
Publicado: (2025)
Learning Generalizable Program and Architecture Representations for Performance Modeling
por: Li, Lingda, et al.
Publicado: (2023)
por: Li, Lingda, et al.
Publicado: (2023)
DeepVigor+: Scalable and Accurate Semi-Analytical Fault Resilience Analysis for Deep Neural Network
por: Ahmadilivani, Mohammad Hasan, et al.
Publicado: (2024)
por: Ahmadilivani, Mohammad Hasan, et al.
Publicado: (2024)
BBS: Bi-directional Bit-level Sparsity for Deep Learning Acceleration
por: Chen, Yuzong, et al.
Publicado: (2024)
por: Chen, Yuzong, et al.
Publicado: (2024)
DeepSeq2: Enhanced Sequential Circuit Learning with Disentangled Representations
por: Khan, Sadaf, et al.
Publicado: (2024)
por: Khan, Sadaf, et al.
Publicado: (2024)
Compact Yet Highly Accurate Printed Classifiers Using Sequential Support Vector Machine Circuits
por: Sertaridis, Ilias, et al.
Publicado: (2025)
por: Sertaridis, Ilias, et al.
Publicado: (2025)
DeepGate4: Efficient and Effective Representation Learning for Circuit Design at Scale
por: Zheng, Ziyang, et al.
Publicado: (2025)
por: Zheng, Ziyang, et al.
Publicado: (2025)
Schrödinger's FP: Dynamic Adaptation of Floating-Point Containers for Deep Learning Training
por: Nikolić, Miloš, et al.
Publicado: (2022)
por: Nikolić, Miloš, et al.
Publicado: (2022)
TimingLLM: A Two-Stage Retrieval-Augmented Framework for Pre-Synthesis Timing Prediction from Verilog
por: Abdollahi, Armin, et al.
Publicado: (2026)
por: Abdollahi, Armin, et al.
Publicado: (2026)
Rescaling-Aware Training for Efficient Deployment of Deep Learning Models on Full-Integer Hardware
por: Mueller, Lion, et al.
Publicado: (2025)
por: Mueller, Lion, et al.
Publicado: (2025)
CuAsmRL: Optimizing GPU SASS Schedules via Deep Reinforcement Learning
por: He, Guoliang, et al.
Publicado: (2025)
por: He, Guoliang, et al.
Publicado: (2025)
KLLM: Fast LLM Inference with K-Means Quantization
por: Wu, Xueying, et al.
Publicado: (2025)
por: Wu, Xueying, et al.
Publicado: (2025)
MetaML-Pro: Cross-Stage Design Flow Automation for Efficient Deep Learning Acceleration
por: Que, Zhiqiang, et al.
Publicado: (2025)
por: Que, Zhiqiang, et al.
Publicado: (2025)
Mugi: Value Level Parallelism For Efficient LLMs
por: Price, Daniel, et al.
Publicado: (2026)
por: Price, Daniel, et al.
Publicado: (2026)
Hardware Software Optimizations for Fast Model Recovery on Reconfigurable Architectures
por: Xu, Bin, et al.
Publicado: (2025)
por: Xu, Bin, et al.
Publicado: (2025)
DS-CIM: Digital Stochastic Computing-In-Memory Featuring Accurate OR-Accumulation via Sample Region Remapping for Edge AI Models
por: Shao, Kunming, et al.
Publicado: (2026)
por: Shao, Kunming, et al.
Publicado: (2026)
RL-MUL 2.0: Multiplier Design Optimization with Parallel Deep Reinforcement Learning and Space Reduction
por: Zuo, Dongsheng, et al.
Publicado: (2024)
por: Zuo, Dongsheng, et al.
Publicado: (2024)
SafeCiM: Investigating Resilience of Hybrid Floating-Point Compute-in-Memory Deep Learning Accelerators
por: Bhattacharya, Swastik, et al.
Publicado: (2025)
por: Bhattacharya, Swastik, et al.
Publicado: (2025)
NAS-Cap: Deep-Learning Driven 3-D Capacitance Extraction with Neural Architecture Search and Data Augmentation
por: Li, Haoyuan, et al.
Publicado: (2024)
por: Li, Haoyuan, et al.
Publicado: (2024)
Taming the Exponential: A Fast Softmax Surrogate for Integer-Native Edge Inference
por: Danopoulos, Dimitrios, et al.
Publicado: (2026)
por: Danopoulos, Dimitrios, et al.
Publicado: (2026)
FuseFlow: A Fusion-Centric Compilation Framework for Sparse Deep Learning on Streaming Dataflow
por: Lacouture, Rubens, et al.
Publicado: (2025)
por: Lacouture, Rubens, et al.
Publicado: (2025)
Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache Quantization
por: Kim, Minsu, et al.
Publicado: (2025)
por: Kim, Minsu, et al.
Publicado: (2025)
Ejemplares similares
-
SAHM: State-Aware Heterogeneous Multicore for Single-Thread Performance
por: Wadle, Shayne, et al.
Publicado: (2025) -
Beyond Static Policies: Exploring Dynamic Policy Selection for Single-Thread Performance Optimization
por: Zhang, Yanxin, et al.
Publicado: (2026) -
Computer Architecture's AlphaZero Moment: Automated Discovery in an Encircled World
por: Sankaralingam, Karthikeyan
Publicado: (2026) -
IPU: Flexible Hardware Introspection Units
por: McDougall, Ian, et al.
Publicado: (2023) -
Privacy-Preserving Performance Profiling of In-The-Wild GPUs
por: McDougall, Ian, et al.
Publicado: (2025)