On Latency Predictors for Neural Architecture Search
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Akhauri, Yash, Abdelfattah, Mohamed S. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Neural Architecture Search of Hybrid Models for NPU-CIM Heterogeneous AR/VR Devices
von: Zhao, Yiwei, et al.
Veröffentlicht: (2024)
von: Zhao, Yiwei, et al.
Veröffentlicht: (2024)
Encodings for Prediction-based Neural Architecture Search
von: Akhauri, Yash, et al.
Veröffentlicht: (2024)
von: Akhauri, Yash, et al.
Veröffentlicht: (2024)
SparAMX: Accelerating Compressed LLMs Token Generation on AMX-powered CPUs
von: AbouElhamayed, Ahmed F., et al.
Veröffentlicht: (2025)
von: AbouElhamayed, Ahmed F., et al.
Veröffentlicht: (2025)
Model Quantization and Hardware Acceleration for Vision Transformers: A Comprehensive Survey
von: Du, Dayou, et al.
Veröffentlicht: (2024)
von: Du, Dayou, et al.
Veröffentlicht: (2024)
Shedding the Bits: Pushing the Boundaries of Quantization with Minifloats on FPGAs
von: Aggarwal, Shivam, et al.
Veröffentlicht: (2023)
von: Aggarwal, Shivam, et al.
Veröffentlicht: (2023)
SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
USEFUSE: Uniform Stride for Enhanced Performance in Fused Layer Architecture of Deep Neural Networks
von: Ibrahim, Muhammad Sohail, et al.
Veröffentlicht: (2024)
von: Ibrahim, Muhammad Sohail, et al.
Veröffentlicht: (2024)
Non-Monotonic Latency in Apple MPS Decoding: KV Cache Interactions and Execution Regimes
von: Hendria, Willy Fitra
Veröffentlicht: (2026)
von: Hendria, Willy Fitra
Veröffentlicht: (2026)
HARFLOW3D: A Latency-Oriented 3D-CNN Accelerator Toolflow for HAR on FPGA Devices
von: Toupas, Petros, et al.
Veröffentlicht: (2023)
von: Toupas, Petros, et al.
Veröffentlicht: (2023)
Search Your Block Floating Point Scales!
von: Gupta, Tanmaey, et al.
Veröffentlicht: (2026)
von: Gupta, Tanmaey, et al.
Veröffentlicht: (2026)
Smaller, Faster, Cheaper: Architectural Designs for Efficient Machine Learning
von: Walton, Steven
Veröffentlicht: (2025)
von: Walton, Steven
Veröffentlicht: (2025)
Energy Efficient Exact and Approximate Systolic Array Architecture for Matrix Multiplication
von: Jaswal, Pragun, et al.
Veröffentlicht: (2025)
von: Jaswal, Pragun, et al.
Veröffentlicht: (2025)
QUILL: An Algorithm-Architecture Co-Design for Cache-Local Deformable Attention
von: Oh, Hyunwoo, et al.
Veröffentlicht: (2025)
von: Oh, Hyunwoo, et al.
Veröffentlicht: (2025)
Neuro-Channel Networks: A Multiplication-Free Architecture by Biological Signal Transmission
von: Mete, Emrah, et al.
Veröffentlicht: (2026)
von: Mete, Emrah, et al.
Veröffentlicht: (2026)
NeuralFuse: Learning to Recover the Accuracy of Access-Limited Neural Network Inference in Low-Voltage Regimes
von: Sun, Hao-Lun, et al.
Veröffentlicht: (2023)
von: Sun, Hao-Lun, et al.
Veröffentlicht: (2023)
Design Insights and Comparative Evaluation of a Hardware-Based Cooperative Perception Architecture for Lane Change Prediction
von: Manzour, Mohamed, et al.
Veröffentlicht: (2025)
von: Manzour, Mohamed, et al.
Veröffentlicht: (2025)
PyGim: An Efficient Graph Neural Network Library for Real Processing-In-Memory Architectures
von: Giannoula, Christina, et al.
Veröffentlicht: (2024)
von: Giannoula, Christina, et al.
Veröffentlicht: (2024)
Accelerating 3D Gaussian Splatting with Neural Sorting and Axis-Oriented Rasterization
von: Wang, Zhican, et al.
Veröffentlicht: (2025)
von: Wang, Zhican, et al.
Veröffentlicht: (2025)
RaGNNarok: A Light-Weight Graph Neural Network for Enhancing Radar Point Clouds on Unmanned Ground Vehicles
von: Hunt, David, et al.
Veröffentlicht: (2025)
von: Hunt, David, et al.
Veröffentlicht: (2025)
Benchmarking Deep Learning Models on NVIDIA Jetson Nano for Real-Time Systems: An Empirical Investigation
von: Swaminathan, Tushar Prasanna, et al.
Veröffentlicht: (2024)
von: Swaminathan, Tushar Prasanna, et al.
Veröffentlicht: (2024)
SMOF: Streaming Modern CNNs on FPGAs with Smart Off-Chip Eviction
von: Toupas, Petros, et al.
Veröffentlicht: (2024)
von: Toupas, Petros, et al.
Veröffentlicht: (2024)
ViM-Q: Scalable Algorithm-Hardware Co-Design for Vision Mamba Model Inference on FPGA
von: Lyu, Shengzhe, et al.
Veröffentlicht: (2026)
von: Lyu, Shengzhe, et al.
Veröffentlicht: (2026)
Stella Nera: A Differentiable Maddness-Based Hardware Accelerator for Efficient Approximate Matrix Multiplication
von: Schönleber, Jannis, et al.
Veröffentlicht: (2023)
von: Schönleber, Jannis, et al.
Veröffentlicht: (2023)
AHCQ-SAM: Toward Accurate and Hardware-Compatible Post-Training Segment Anything Model Quantization
von: Zhang, Wenlun, et al.
Veröffentlicht: (2025)
von: Zhang, Wenlun, et al.
Veröffentlicht: (2025)
Ditto: Accelerating Diffusion Model via Temporal Value Similarity
von: Kim, Sungbin, et al.
Veröffentlicht: (2025)
von: Kim, Sungbin, et al.
Veröffentlicht: (2025)
CRISP: Hybrid Structured Sparsity for Class-aware Model Pruning
von: Aggarwal, Shivam, et al.
Veröffentlicht: (2023)
von: Aggarwal, Shivam, et al.
Veröffentlicht: (2023)
Performance Analysis of Edge and In-Sensor AI Processors: A Comparative Review
von: Capogrosso, Luigi, et al.
Veröffentlicht: (2026)
von: Capogrosso, Luigi, et al.
Veröffentlicht: (2026)
ViTCoD: Vision Transformer Acceleration via Dedicated Algorithm and Accelerator Co-Design
von: You, Haoran, et al.
Veröffentlicht: (2022)
von: You, Haoran, et al.
Veröffentlicht: (2022)
Mix-and-Match Pruning: Globally Guided Layer-Wise Sparsification of DNNs
von: Monachan, Danial, et al.
Veröffentlicht: (2026)
von: Monachan, Danial, et al.
Veröffentlicht: (2026)
A2Q+: Improving Accumulator-Aware Weight Quantization
von: Colbert, Ian, et al.
Veröffentlicht: (2024)
von: Colbert, Ian, et al.
Veröffentlicht: (2024)
Understanding the Performance Horizon of the Latest ML Workloads with NonGEMM Workloads
von: Karami, Rachid, et al.
Veröffentlicht: (2024)
von: Karami, Rachid, et al.
Veröffentlicht: (2024)
Graph neural networks with configuration cross-attention for tensor compilers
von: Khizbullin, Dmitrii, et al.
Veröffentlicht: (2024)
von: Khizbullin, Dmitrii, et al.
Veröffentlicht: (2024)
PoTAcc: A Pipeline for End-to-End Acceleration of Power-of-Two Quantized DNNs
von: Saha, Rappy, et al.
Veröffentlicht: (2026)
von: Saha, Rappy, et al.
Veröffentlicht: (2026)
Toward A Formalized Approach for Spike Sorting Algorithms and Hardware Evaluation
von: Zhang, Tim, et al.
Veröffentlicht: (2022)
von: Zhang, Tim, et al.
Veröffentlicht: (2022)
SONIQ: System-Optimized Noise-Injected Ultra-Low-Precision Quantization with Full-Precision Parity
von: Zhou, Cyrus, et al.
Veröffentlicht: (2023)
von: Zhou, Cyrus, et al.
Veröffentlicht: (2023)
GCL-Sampler: Discovering Kernel Similarity for Sampled GPU Simulation via Graph Contrastive Learning
von: Wang, Jiaqi, et al.
Veröffentlicht: (2026)
von: Wang, Jiaqi, et al.
Veröffentlicht: (2026)
Concorde: Fast and Accurate CPU Performance Modeling with Compositional Analytical-ML Fusion
von: Nasr-Esfahany, Arash, et al.
Veröffentlicht: (2025)
von: Nasr-Esfahany, Arash, et al.
Veröffentlicht: (2025)
Prefill vs. Decode Bottlenecks: SRAM-Frequency Tradeoffs and the Memory-Bandwidth Ceiling
von: Atmer, Hannah, et al.
Veröffentlicht: (2025)
von: Atmer, Hannah, et al.
Veröffentlicht: (2025)
HERCULES: Hardware-Efficient, Robust, Continual Learning Neural Architecture Search
von: Gambella, Matteo, et al.
Veröffentlicht: (2026)
von: Gambella, Matteo, et al.
Veröffentlicht: (2026)
LowRA: Accurate and Efficient LoRA Fine-Tuning of LLMs under 2 Bits
von: Zhou, Zikai, et al.
Veröffentlicht: (2025)
von: Zhou, Zikai, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Neural Architecture Search of Hybrid Models for NPU-CIM Heterogeneous AR/VR Devices
von: Zhao, Yiwei, et al.
Veröffentlicht: (2024) -
Encodings for Prediction-based Neural Architecture Search
von: Akhauri, Yash, et al.
Veröffentlicht: (2024) -
SparAMX: Accelerating Compressed LLMs Token Generation on AMX-powered CPUs
von: AbouElhamayed, Ahmed F., et al.
Veröffentlicht: (2025) -
Model Quantization and Hardware Acceleration for Vision Transformers: A Comprehensive Survey
von: Du, Dayou, et al.
Veröffentlicht: (2024) -
Shedding the Bits: Pushing the Boundaries of Quantization with Minifloats on FPGAs
von: Aggarwal, Shivam, et al.
Veröffentlicht: (2023)