PIVOT- Input-aware Path Selection for Energy-efficient ViT Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Moitra, Abhishek, Bhattacharjee, Abhiroop, Panda, Priyadarshini |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TReX- Reusing Vision Transformer's Attention for Efficient Xbar-based Computing
von: Moitra, Abhishek, et al.
Veröffentlicht: (2024)
von: Moitra, Abhishek, et al.
Veröffentlicht: (2024)
DiffAxE: Diffusion-driven Hardware Accelerator Generation and Design Space Exploration
von: Ghosh, Arkapravo, et al.
Veröffentlicht: (2025)
von: Ghosh, Arkapravo, et al.
Veröffentlicht: (2025)
When In-memory Computing Meets Spiking Neural Networks -- A Perspective on Device-Circuit-System-and-Algorithm Co-design
von: Moitra, Abhishek, et al.
Veröffentlicht: (2024)
von: Moitra, Abhishek, et al.
Veröffentlicht: (2024)
MEADOW: Memory-efficient Dataflow and Data Packing for Low Power Edge LLMs
von: Moitra, Abhishek, et al.
Veröffentlicht: (2025)
von: Moitra, Abhishek, et al.
Veröffentlicht: (2025)
Low-power Spike-based Wearable Analytics on RRAM Crossbars
von: Bhattacharjee, Abhiroop, et al.
Veröffentlicht: (2025)
von: Bhattacharjee, Abhiroop, et al.
Veröffentlicht: (2025)
Refining Datapath for Microscaling ViTs
von: Xiao, Can, et al.
Veröffentlicht: (2025)
von: Xiao, Can, et al.
Veröffentlicht: (2025)
PacQ: A SIMT Microarchitecture for Efficient Dataflow in Hyper-asymmetric GEMMs
von: Yin, Ruokai, et al.
Veröffentlicht: (2025)
von: Yin, Ruokai, et al.
Veröffentlicht: (2025)
Enabling Efficient Hardware Acceleration of Hybrid Vision Transformer (ViT) Networks at the Edge
von: Dumoulin, Joren, et al.
Veröffentlicht: (2025)
von: Dumoulin, Joren, et al.
Veröffentlicht: (2025)
Opto-ViT: Architecting a Near-Sensor Region of Interest-Aware Vision Transformer Accelerator with Silicon Photonics
von: Morsali, Mehrdad, et al.
Veröffentlicht: (2025)
von: Morsali, Mehrdad, et al.
Veröffentlicht: (2025)
Towards Forever Access for Implanted Brain-Computer Interfaces
von: Ugur, Muhammed, et al.
Veröffentlicht: (2024)
von: Ugur, Muhammed, et al.
Veröffentlicht: (2024)
Swapping-Centric Neural Recording Systems
von: Ugur, Muhammed, et al.
Veröffentlicht: (2024)
von: Ugur, Muhammed, et al.
Veröffentlicht: (2024)
M$^2$-ViT: Accelerating Hybrid Vision Transformers with Two-Level Mixed Quantization
von: Liang, Yanbiao, et al.
Veröffentlicht: (2024)
von: Liang, Yanbiao, et al.
Veröffentlicht: (2024)
Accelerating ViT Inference on FPGA through Static and Dynamic Pruning
von: Parikh, Dhruv, et al.
Veröffentlicht: (2024)
von: Parikh, Dhruv, et al.
Veröffentlicht: (2024)
ViPSN 2.0: A Reconfigurable Battery-free IoT Platform for Vibration Energy Harvesting
von: Li, Xin, et al.
Veröffentlicht: (2025)
von: Li, Xin, et al.
Veröffentlicht: (2025)
TaiBai: A fully programmable brain-inspired processor with topology-aware efficiency
von: Li, Qianpeng, et al.
Veröffentlicht: (2025)
von: Li, Qianpeng, et al.
Veröffentlicht: (2025)
PolyThrottle: Energy-efficient Neural Network Inference on Edge Devices
von: Yan, Minghao, et al.
Veröffentlicht: (2023)
von: Yan, Minghao, et al.
Veröffentlicht: (2023)
ClipFormer: Key-Value Clipping of Transformers on Memristive Crossbars for Write Noise Mitigation
von: Bhattacharjee, Abhiroop, et al.
Veröffentlicht: (2024)
von: Bhattacharjee, Abhiroop, et al.
Veröffentlicht: (2024)
The Interplay of Computing, Ethics, and Policy in Brain-Computer Interface Design
von: Ugur, Muhammed, et al.
Veröffentlicht: (2024)
von: Ugur, Muhammed, et al.
Veröffentlicht: (2024)
TLV-HGNN: Thinking Like a Vertex for Memory-efficient HGNN Inference
von: Han, Dengke, et al.
Veröffentlicht: (2025)
von: Han, Dengke, et al.
Veröffentlicht: (2025)
Tasa: Thermal-aware 3D-Stacked Architecture Design with Bandwidth Sharing for LLM Inference
von: He, Siyuan, et al.
Veröffentlicht: (2025)
von: He, Siyuan, et al.
Veröffentlicht: (2025)
LoAS: Fully Temporal-Parallel Dataflow for Dual-Sparse Spiking Neural Networks
von: Yin, Ruokai, et al.
Veröffentlicht: (2024)
von: Yin, Ruokai, et al.
Veröffentlicht: (2024)
RED: Energy Optimization Framework for eDRAM-based PIM with Reconfigurable Voltage Swing and Retention-aware Scheduling
von: Kim, Jae-Young, et al.
Veröffentlicht: (2025)
von: Kim, Jae-Young, et al.
Veröffentlicht: (2025)
High Utilization Energy-Aware Real-Time Inference Deep Convolutional Neural Network Accelerator
von: Lin, Kuan-Ting, et al.
Veröffentlicht: (2025)
von: Lin, Kuan-Ting, et al.
Veröffentlicht: (2025)
The Immutable Tensor Architecture: A Pure Dataflow Approach for Secure, Energy-Efficient AI Inference
von: Li, Fang
Veröffentlicht: (2025)
von: Li, Fang
Veröffentlicht: (2025)
HyDRA: Deadline and Reuse-Aware Cacheability for Hardware Accelerators
von: Agarwal, Ayushi, et al.
Veröffentlicht: (2026)
von: Agarwal, Ayushi, et al.
Veröffentlicht: (2026)
Correct Wrong Path
von: Godala, Bhargav Reddy, et al.
Veröffentlicht: (2024)
von: Godala, Bhargav Reddy, et al.
Veröffentlicht: (2024)
CounterPoint: Using Hardware Event Counters to Refute and Refine Microarchitectural Assumptions (Extended Version)
von: Lindsay, Nick, et al.
Veröffentlicht: (2026)
von: Lindsay, Nick, et al.
Veröffentlicht: (2026)
Vulnerabilities in Partial TEE-Shielded LLM Inference with Precomputed Noise
von: Saini, Abhishek, et al.
Veröffentlicht: (2026)
von: Saini, Abhishek, et al.
Veröffentlicht: (2026)
DAG-aware Synthesis Orchestration
von: Li, Yingjie, et al.
Veröffentlicht: (2023)
von: Li, Yingjie, et al.
Veröffentlicht: (2023)
A Hybrid Delay Model for Interconnected Multi-Input Gates
von: Ferdowsi, Arman, et al.
Veröffentlicht: (2024)
von: Ferdowsi, Arman, et al.
Veröffentlicht: (2024)
ACS: Concurrent Kernel Execution on Irregular, Input-Dependent Computational Graphs
von: Durvasula, Sankeerth, et al.
Veröffentlicht: (2024)
von: Durvasula, Sankeerth, et al.
Veröffentlicht: (2024)
LLMulator: Generalizable Cost Modeling for Dataflow Accelerators with Input-Adaptive Control Flow
von: Chang, Kaiyan, et al.
Veröffentlicht: (2025)
von: Chang, Kaiyan, et al.
Veröffentlicht: (2025)
NASiC: 3D NAND-based CAM-Selected Multibit CIM Architecture for Efficient On-Device Mixture-of-Experts LLM Inference
von: Xu, Weikai, et al.
Veröffentlicht: (2026)
von: Xu, Weikai, et al.
Veröffentlicht: (2026)
Optimising GPGPU Execution Through Runtime Micro-Architecture Parameter Analysis
von: Sarda, Giuseppe M., et al.
Veröffentlicht: (2024)
von: Sarda, Giuseppe M., et al.
Veröffentlicht: (2024)
Decoupled Control Flow and Data Access in RISC-V GPGPUs
von: Sarda, Giuseppe M., et al.
Veröffentlicht: (2025)
von: Sarda, Giuseppe M., et al.
Veröffentlicht: (2025)
Accelerating GNN Training through Locality-aware Dropout and Merge
von: Sun, Gongjian, et al.
Veröffentlicht: (2025)
von: Sun, Gongjian, et al.
Veröffentlicht: (2025)
Addressing memory bandwidth scalability in vector processors for streaming applications
von: Altayo, Jordi, et al.
Veröffentlicht: (2025)
von: Altayo, Jordi, et al.
Veröffentlicht: (2025)
RARO: Reliability-aware Conversion with Enhanced Read Performance for QLC SSDs
von: Wang, Yanyun, et al.
Veröffentlicht: (2025)
von: Wang, Yanyun, et al.
Veröffentlicht: (2025)
Educating for Hardware Specialization in the Chiplet Era: A Path for the HPC Community
von: Yoshii, Kazutomo, et al.
Veröffentlicht: (2024)
von: Yoshii, Kazutomo, et al.
Veröffentlicht: (2024)
Be CIM or Be Memory: A Dual-mode-aware DNN Compiler for CIM Accelerators
von: Zhao, Shixin, et al.
Veröffentlicht: (2025)
von: Zhao, Shixin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
TReX- Reusing Vision Transformer's Attention for Efficient Xbar-based Computing
von: Moitra, Abhishek, et al.
Veröffentlicht: (2024) -
DiffAxE: Diffusion-driven Hardware Accelerator Generation and Design Space Exploration
von: Ghosh, Arkapravo, et al.
Veröffentlicht: (2025) -
When In-memory Computing Meets Spiking Neural Networks -- A Perspective on Device-Circuit-System-and-Algorithm Co-design
von: Moitra, Abhishek, et al.
Veröffentlicht: (2024) -
MEADOW: Memory-efficient Dataflow and Data Packing for Low Power Edge LLMs
von: Moitra, Abhishek, et al.
Veröffentlicht: (2025) -
Low-power Spike-based Wearable Analytics on RRAM Crossbars
von: Bhattacharjee, Abhiroop, et al.
Veröffentlicht: (2025)