TCL: Enabling Fast and Efficient Cross-Hardware Tensor Program Optimization via Continual Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Shen, Chaoyao, Jiang, Linfeng, Shen, Yixian, Xu, Tao, Li, Guoqing, Pathania, Anuj, Pimentel, Andy D., Zhang, Meng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Active Imitation Learning for Thermal- and Kernel-Aware LFM Inference on 3D S-NUCA Many-Cores
by: Shen, Yixian, et al.
Published: (2026)
by: Shen, Yixian, et al.
Published: (2026)
Energy-Efficient QoS-Aware Scheduling for S-NUCA Many-Cores
by: Wasala, Sudam M., et al.
Published: (2025)
by: Wasala, Sudam M., et al.
Published: (2025)
Workload-Aware Early-Stage Power Delivery Network Optimization via Architectural Power Traces
by: Hayes, Oran, et al.
Published: (2026)
by: Hayes, Oran, et al.
Published: (2026)
Multi-Partner Project: COIN-3D -- Collaborative Innovation in 3D VLSI Reliability
by: Gourdoumanis, George Rafael, et al.
Published: (2026)
by: Gourdoumanis, George Rafael, et al.
Published: (2026)
FETTA: Flexible and Efficient Hardware Accelerator for Tensorized Neural Network Training
by: Lu, Jinming, et al.
Published: (2025)
by: Lu, Jinming, et al.
Published: (2025)
Fast Graph Vector Search via Hardware Acceleration and Delayed-Synchronization Traversal
by: Jiang, Wenqi, et al.
Published: (2024)
by: Jiang, Wenqi, et al.
Published: (2024)
HERO: Hardware-Efficient RL-based Optimization Framework for NeRF Quantization
by: Zhang, Yipu, et al.
Published: (2025)
by: Zhang, Yipu, et al.
Published: (2025)
Fast Cross-Operator Optimization of Attention Dataflow
by: Chang, Haodong, et al.
Published: (2026)
by: Chang, Haodong, et al.
Published: (2026)
Enabling Efficient Hardware Acceleration of Hybrid Vision Transformer (ViT) Networks at the Edge
by: Dumoulin, Joren, et al.
Published: (2025)
by: Dumoulin, Joren, et al.
Published: (2025)
FPGA-Optimized Hardware Accelerator for Fast Fourier Transform and Singular Value Decomposition in AI
by: Ding, Hong, et al.
Published: (2025)
by: Ding, Hong, et al.
Published: (2025)
A Tensor-Train Decomposition based Compression of LLMs on Group Vector Systolic Accelerator
by: Huang, Sixiao, et al.
Published: (2025)
by: Huang, Sixiao, et al.
Published: (2025)
Hardware Software Optimizations for Fast Model Recovery on Reconfigurable Architectures
by: Xu, Bin, et al.
Published: (2025)
by: Xu, Bin, et al.
Published: (2025)
ATiM: Autotuning Tensor Programs for Processing-in-DRAM
by: Shin, Yongwon, et al.
Published: (2024)
by: Shin, Yongwon, et al.
Published: (2024)
In-Memory Computing Architecture for Efficient Hardware Security
by: Ajmi, Hala, et al.
Published: (2024)
by: Ajmi, Hala, et al.
Published: (2024)
EN-T: Optimizing Tensor Computing Engines Performance via Encoder-Based Methodology
by: Wu, Qizhe, et al.
Published: (2024)
by: Wu, Qizhe, et al.
Published: (2024)
CogSys: Efficient and Scalable Neurosymbolic Cognition System via Algorithm-Hardware Co-Design
by: Wan, Zishen, et al.
Published: (2025)
by: Wan, Zishen, et al.
Published: (2025)
CREATE: Cross-Layer Resilience Characterization and Optimization for Efficient yet Reliable Embodied AI Systems
by: Xie, Tong, et al.
Published: (2026)
by: Xie, Tong, et al.
Published: (2026)
A Power-Efficient Hardware Implementation of L-Mul
by: Chen, Ruiqi, et al.
Published: (2024)
by: Chen, Ruiqi, et al.
Published: (2024)
An Efficient Sparse Hardware Accelerator for Spike-Driven Transformer
by: Li, Zhengke, et al.
Published: (2025)
by: Li, Zhengke, et al.
Published: (2025)
The Quest for Reliable AI Accelerators: Cross-Layer Evaluation and Design Optimization
by: Li, Meng, et al.
Published: (2026)
by: Li, Meng, et al.
Published: (2026)
Energy-Efficient Hardware Acceleration of Whisper ASR on a CGLA
by: Ando, Takuto, et al.
Published: (2025)
by: Ando, Takuto, et al.
Published: (2025)
SemanticBBV: A Semantic Signature for Cross-Program Knowledge Reuse in Microarchitecture Simulation
by: Liu, Zhenguo, et al.
Published: (2025)
by: Liu, Zhenguo, et al.
Published: (2025)
DRACO: Co-design for DSP-Efficient Rigid Body Dynamics Accelerator
by: Liu, Xingyu, et al.
Published: (2025)
by: Liu, Xingyu, et al.
Published: (2025)
TTP: A Hardware-Efficient Design for Precise Prefetching in Ray Tracing
by: Tozlu, Yavuz Selim, et al.
Published: (2026)
by: Tozlu, Yavuz Selim, et al.
Published: (2026)
SeDA: Secure and Efficient DNN Accelerators with Hardware/Software Synergy
by: Xuan, Wei, et al.
Published: (2025)
by: Xuan, Wei, et al.
Published: (2025)
Automatic Hardware Pragma Insertion in High-Level Synthesis: A Non-Linear Programming Approach
by: Pouget, Stéphane, et al.
Published: (2024)
by: Pouget, Stéphane, et al.
Published: (2024)
N-TORC: Native Tensor Optimizer for Real-time Constraints
by: Singh, Suyash Vardhan, et al.
Published: (2025)
by: Singh, Suyash Vardhan, et al.
Published: (2025)
Cross-Layer Design of Vector-Symbolic Computing: Bridging Cognition and Brain-Inspired Hardware Acceleration
by: Du, Shuting, et al.
Published: (2025)
by: Du, Shuting, et al.
Published: (2025)
Cocco: Hardware-Mapping Co-Exploration towards Memory Capacity-Communication Optimization
by: Tan, Zhanhong, et al.
Published: (2024)
by: Tan, Zhanhong, et al.
Published: (2024)
Hardware-Efficient CNNs: Interleaved Approximate FP32 Multipliers for Kernel Computation
by: Gowda, Bindu G, et al.
Published: (2025)
by: Gowda, Bindu G, et al.
Published: (2025)
StruM: Structured Mixed Precision for Efficient Deep Learning Hardware Codesign
by: Wu, Michael, et al.
Published: (2025)
by: Wu, Michael, et al.
Published: (2025)
Hardware-Efficient Accurate 4-bit Multiplier for Xilinx 7 Series FPGAs
by: Kida, Misaki, et al.
Published: (2025)
by: Kida, Misaki, et al.
Published: (2025)
Hardware Efficient Accelerator for Spiking Transformer With Reconfigurable Parallel Time Step Computing
by: Chen, Bo-Yu, et al.
Published: (2025)
by: Chen, Bo-Yu, et al.
Published: (2025)
TEMP: A Memory Efficient Physical-aware Tensor Partition-Mapping Framework on Wafer-scale Chips
by: Wang, Huizheng, et al.
Published: (2025)
by: Wang, Huizheng, et al.
Published: (2025)
CIM-Tuner: Balancing the Compute and Storage Capacity of SRAM-CIM Accelerator via Hardware-mapping Co-exploration
by: Chen, Jinwu, et al.
Published: (2026)
by: Chen, Jinwu, et al.
Published: (2026)
FADiff: Fusion-Aware Differentiable Optimization for DNN Scheduling on Tensor Accelerators
by: Jia, Shuao, et al.
Published: (2025)
by: Jia, Shuao, et al.
Published: (2025)
Comprehensive Design Space Exploration for Tensorized Neural Network Hardware Accelerators
by: Zhang, Jinsong, et al.
Published: (2025)
by: Zhang, Jinsong, et al.
Published: (2025)
Realizing Hardware-Optimized General Tree-Based Data Structures for Heterogeneous System Classes
by: Biebert, Daniel, et al.
Published: (2025)
by: Biebert, Daniel, et al.
Published: (2025)
MCMComm: Hardware-Software Co-Optimization for End-to-End Communication in Multi-Chip-Modules
by: Raj, Ritik, et al.
Published: (2025)
by: Raj, Ritik, et al.
Published: (2025)
Hardware Acceleration of Kolmogorov-Arnold Network (KAN) for Lightweight Edge Inference
by: Huang, Wei-Hsing, et al.
Published: (2024)
by: Huang, Wei-Hsing, et al.
Published: (2024)
Similar Items
-
Active Imitation Learning for Thermal- and Kernel-Aware LFM Inference on 3D S-NUCA Many-Cores
by: Shen, Yixian, et al.
Published: (2026) -
Energy-Efficient QoS-Aware Scheduling for S-NUCA Many-Cores
by: Wasala, Sudam M., et al.
Published: (2025) -
Workload-Aware Early-Stage Power Delivery Network Optimization via Architectural Power Traces
by: Hayes, Oran, et al.
Published: (2026) -
Multi-Partner Project: COIN-3D -- Collaborative Innovation in 3D VLSI Reliability
by: Gourdoumanis, George Rafael, et al.
Published: (2026) -
FETTA: Flexible and Efficient Hardware Accelerator for Tensorized Neural Network Training
by: Lu, Jinming, et al.
Published: (2025)