TLX: Hardware-Native, Evolvable MIMW GPU Compiler for Large-scale Production Environments
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Guan, Yue, Yu, Hongtao, Chen, Peng, Shi, Daohang, Manivannan, Karthik, Riasanovsky, Nicholas J, Ren, Manman, Wang, Lei, Nay, Shane, Kanuparthy, Partha, Pan, Zaifeng, Hu, Zhengding, Ding, Yufei |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Real Time Evolvable Hardware for Optimal Reconfiguration of Cusp-Like Pulse Shapers
par: Lanchares, Juan, et autres
Publié: (2024)
par: Lanchares, Juan, et autres
Publié: (2024)
Bombyx: OpenCilk Compilation for FPGA Hardware Acceleration
par: Shahawy, Mohamed, et autres
Publié: (2025)
par: Shahawy, Mohamed, et autres
Publié: (2025)
Generation of Compiler Backends from Formal Models of Hardware
par: Smith, Gus Henry
Publié: (2024)
par: Smith, Gus Henry
Publié: (2024)
NeoMem: Hardware/Software Co-Design for CXL-Native Memory Tiering
par: Zhou, Zhe, et autres
Publié: (2024)
par: Zhou, Zhe, et autres
Publié: (2024)
HERO-Sign: Hierarchical Tuning and Efficient Compiler-Time GPU Optimizations for SPHINCS+ Signature Generation
par: Zhou, Yaoyun, et autres
Publié: (2025)
par: Zhou, Yaoyun, et autres
Publié: (2025)
Hardware-Aware Neural Network Compilation with Learned Optimization: A RISC-V Accelerator Approach
par: Ganti, Ravindra, et autres
Publié: (2025)
par: Ganti, Ravindra, et autres
Publié: (2025)
Hardware vs. Software Implementation of Warp-Level Features in Vortex RISC-V GPU
par: Pu, Huanzhi, et autres
Publié: (2025)
par: Pu, Huanzhi, et autres
Publié: (2025)
Look-Up Table based Neural Network Hardware
par: Sen, Ovishake, et autres
Publié: (2024)
par: Sen, Ovishake, et autres
Publié: (2024)
SIP: Autotuning GPU Native Schedules via Stochastic Instruction Perturbation
par: He, Guoliang, et autres
Publié: (2024)
par: He, Guoliang, et autres
Publié: (2024)
Vorion: A RISC-V GPU with Hardware-Accelerated 3D Gaussian Rendering and Training
par: Wang, Yipeng, et autres
Publié: (2025)
par: Wang, Yipeng, et autres
Publié: (2025)
Design Environment of Quantization-Aware Edge AI Hardware for Few-Shot Learning
par: Kanda, R., et autres
Publié: (2026)
par: Kanda, R., et autres
Publié: (2026)
Bit-Width-Aware Design Environment for Few-Shot Learning on Edge AI Hardware
par: Kanda, R., et autres
Publié: (2026)
par: Kanda, R., et autres
Publié: (2026)
ChipBench: A Next-Step Benchmark for Evaluating LLM Performance in AI-Aided Chip Design
par: Yu, Zhongkai, et autres
Publié: (2026)
par: Yu, Zhongkai, et autres
Publié: (2026)
OneAdapt: Adaptive Compilation for Resource-Constrained Photonic One-Way Quantum Computing
par: Zhang, Hezi, et autres
Publié: (2025)
par: Zhang, Hezi, et autres
Publié: (2025)
Parallax: A Compiler for Neutral Atom Quantum Computers under Hardware Constraints
par: Ludmir, Jason, et autres
Publié: (2024)
par: Ludmir, Jason, et autres
Publié: (2024)
CXL-GPU: Pushing GPU Memory Boundaries with the Integration of CXL Technologies
par: Gouk, Donghyun, et autres
Publié: (2025)
par: Gouk, Donghyun, et autres
Publié: (2025)
AMMA: A Multi-Chiplet Memory-Centric Architecture for Low-Latency 1M Context Attention Serving
par: Yu, Zhongkai, et autres
Publié: (2026)
par: Yu, Zhongkai, et autres
Publié: (2026)
ChipMATE: Multi-Agent Training via Reinforcement Learning for Enhanced RTL Generation
par: Yu, Zhongkai, et autres
Publié: (2026)
par: Yu, Zhongkai, et autres
Publié: (2026)
Sustainable Hardware Specialization
par: Dangi, Pranav, et autres
Publié: (2024)
par: Dangi, Pranav, et autres
Publié: (2024)
RoboGPU: Accelerating GPU Collision Detection for Robotics
par: Liu, Lufei, et autres
Publié: (2026)
par: Liu, Lufei, et autres
Publié: (2026)
An Evaluation and Comparison of GPU Hardware and Solver Libraries for Accelerating the OPM Flow Reservoir Simulator
par: Qiu, Tong Dong, et autres
Publié: (2023)
par: Qiu, Tong Dong, et autres
Publié: (2023)
CHOSEN: Compilation to Hardware Optimization Stack for Efficient Vision Transformer Inference
par: Sadeghi, Mohammad Erfan, et autres
Publié: (2024)
par: Sadeghi, Mohammad Erfan, et autres
Publié: (2024)
MERE: Hardware-Software Co-Design for Masking Cache Miss Latency in Embedded Processors
par: You, Dean, et autres
Publié: (2025)
par: You, Dean, et autres
Publié: (2025)
CODO: An Automated Compiler for Comprehensive Dataflow Optimization
par: Zhang, Weichuang, et autres
Publié: (2026)
par: Zhang, Weichuang, et autres
Publié: (2026)
Compilation and Execution of an Embeddable YOLO-NAS on the VTA
par: Faure-Gignoux, Anthony, et autres
Publié: (2026)
par: Faure-Gignoux, Anthony, et autres
Publié: (2026)
Revet: A Language and Compiler for Dataflow Threads
par: Rucker, Alexander, et autres
Publié: (2023)
par: Rucker, Alexander, et autres
Publié: (2023)
Analyzing Modern NVIDIA GPU cores
par: Huerta, Rodrigo, et autres
Publié: (2025)
par: Huerta, Rodrigo, et autres
Publié: (2025)
An FPGA Compiler for On-the-Fly Adaptive CNN Deployment and Reconfiguration
par: Mazouz, Alaa, et autres
Publié: (2025)
par: Mazouz, Alaa, et autres
Publié: (2025)
SynDCIM: A Performance-Aware Digital Computing-in-Memory Compiler with Multi-Spec-Oriented Subcircuit Synthesis
par: Shao, Kunming, et autres
Publié: (2024)
par: Shao, Kunming, et autres
Publié: (2024)
Patterns behind Chaos: Forecasting Data Movement for Efficient Large-Scale MoE LLM Inference
par: Yu, Zhongkai, et autres
Publié: (2025)
par: Yu, Zhongkai, et autres
Publié: (2025)
Design of a GPU with Heterogeneous Cores for Graphics
par: Tomás, Aurora, et autres
Publié: (2026)
par: Tomás, Aurora, et autres
Publié: (2026)
Benchmarking and Dissecting the Nvidia Hopper GPU Architecture
par: Luo, Weile, et autres
Publié: (2024)
par: Luo, Weile, et autres
Publié: (2024)
COOK Access Control on an embedded Volta GPU
par: Lesage, Benjamin, et autres
Publié: (2024)
par: Lesage, Benjamin, et autres
Publié: (2024)
PIMCOMP: An End-to-End DNN Compiler for Processing-In-Memory Accelerators
par: Sun, Xiaotian, et autres
Publié: (2024)
par: Sun, Xiaotian, et autres
Publié: (2024)
Analyzing and Improving Hardware Modeling of Accel-Sim
par: Huerta, Rodrigo, et autres
Publié: (2024)
par: Huerta, Rodrigo, et autres
Publié: (2024)
Hardware and software build flow with SoCMake
par: Pejašinović, Risto, et autres
Publié: (2025)
par: Pejašinović, Risto, et autres
Publié: (2025)
QED: Scalable Verification of Hardware Memory Consistency
par: Ravi, Gokulan, et autres
Publié: (2024)
par: Ravi, Gokulan, et autres
Publié: (2024)
NeuroVM: Dynamic Neuromorphic Hardware Virtualization
par: Isik, Murat, et autres
Publié: (2024)
par: Isik, Murat, et autres
Publié: (2024)
In-Memory Computing Architecture for Efficient Hardware Security
par: Ajmi, Hala, et autres
Publié: (2024)
par: Ajmi, Hala, et autres
Publié: (2024)
DRCY: Agentic Hardware Design Reviews
par: Dumont, Kyle, et autres
Publié: (2026)
par: Dumont, Kyle, et autres
Publié: (2026)
Documents similaires
-
Real Time Evolvable Hardware for Optimal Reconfiguration of Cusp-Like Pulse Shapers
par: Lanchares, Juan, et autres
Publié: (2024) -
Bombyx: OpenCilk Compilation for FPGA Hardware Acceleration
par: Shahawy, Mohamed, et autres
Publié: (2025) -
Generation of Compiler Backends from Formal Models of Hardware
par: Smith, Gus Henry
Publié: (2024) -
NeoMem: Hardware/Software Co-Design for CXL-Native Memory Tiering
par: Zhou, Zhe, et autres
Publié: (2024) -
HERO-Sign: Hierarchical Tuning and Efficient Compiler-Time GPU Optimizations for SPHINCS+ Signature Generation
par: Zhou, Yaoyun, et autres
Publié: (2025)