CuLifter: Lifting GPU Binaries to Typed IR
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhao, Jisheng, Pu, Huanzhi, Jeong, Shinnung, Ahn, Chihyo, Kim, Hyesoon |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Hardware vs. Software Implementation of Warp-Level Features in Vortex RISC-V GPU
par: Pu, Huanzhi, et autres
Publié: (2025)
par: Pu, Huanzhi, et autres
Publié: (2025)
Inside VOLT: Designing an Open-Source GPU Compiler
par: Jeong, Shinnung, et autres
Publié: (2025)
par: Jeong, Shinnung, et autres
Publié: (2025)
Towards Performance-Aware Allocation for Accelerated Machine Learning on GPU-SSD Systems
par: Gundawar, Ayush, et autres
Publié: (2024)
par: Gundawar, Ayush, et autres
Publié: (2024)
Characterizing CPU-Induced Slowdowns in Multi-GPU LLM Inference
par: Chung, Euijun, et autres
Publié: (2026)
par: Chung, Euijun, et autres
Publié: (2026)
CXL-GPU: Pushing GPU Memory Boundaries with the Integration of CXL Technologies
par: Gouk, Donghyun, et autres
Publié: (2025)
par: Gouk, Donghyun, et autres
Publié: (2025)
RoboGPU: Accelerating GPU Collision Detection for Robotics
par: Liu, Lufei, et autres
Publié: (2026)
par: Liu, Lufei, et autres
Publié: (2026)
GAP-LA: GPU-Accelerated Performance-Driven Layer Assignment
par: Zhao, Chunyuan, et autres
Publié: (2025)
par: Zhao, Chunyuan, et autres
Publié: (2025)
CMD: A Cache-assisted GPU Memory Deduplication Architecture
par: Zhao, Wei, et autres
Publié: (2024)
par: Zhao, Wei, et autres
Publié: (2024)
Analyzing Modern NVIDIA GPU cores
par: Huerta, Rodrigo, et autres
Publié: (2025)
par: Huerta, Rodrigo, et autres
Publié: (2025)
Design of a GPU with Heterogeneous Cores for Graphics
par: Tomás, Aurora, et autres
Publié: (2026)
par: Tomás, Aurora, et autres
Publié: (2026)
Benchmarking and Dissecting the Nvidia Hopper GPU Architecture
par: Luo, Weile, et autres
Publié: (2024)
par: Luo, Weile, et autres
Publié: (2024)
COOK Access Control on an embedded Volta GPU
par: Lesage, Benjamin, et autres
Publié: (2024)
par: Lesage, Benjamin, et autres
Publié: (2024)
PipeRTL: Timing-Aware Pipeline Optimization at IR-Level for RTL Generation
par: Yin, Shuo, et autres
Publié: (2026)
par: Yin, Shuo, et autres
Publié: (2026)
RoMe: Row Granularity Access Memory System for Large Language Models
par: Nam, Hwayong, et autres
Publié: (2025)
par: Nam, Hwayong, et autres
Publié: (2025)
CuAsmRL: Optimizing GPU SASS Schedules via Deep Reinforcement Learning
par: He, Guoliang, et autres
Publié: (2025)
par: He, Guoliang, et autres
Publié: (2025)
Multiport Support for Vortex OpenGPU Memory Hierarchy
par: Shin, Injae, et autres
Publié: (2025)
par: Shin, Injae, et autres
Publié: (2025)
Exploiting Control-flow Enforcement Technology for Sound and Precise Static Binary Disassembly
par: Zhao, Brian, et autres
Publié: (2025)
par: Zhao, Brian, et autres
Publié: (2025)
Thermal Analysis for NVIDIA GTX480 Fermi GPU Architecture
par: Nagendra, Savinay
Publié: (2024)
par: Nagendra, Savinay
Publié: (2024)
The Anatomy of Silent Data Corruption: GPU Error Pattern Study and Modeling Guidance
par: Tung, Chung-Hsuan, et autres
Publié: (2026)
par: Tung, Chung-Hsuan, et autres
Publié: (2026)
EnergAIzer: Fast and Accurate GPU Power Estimation Framework for AI Workloads
par: Lee, Kyungmi, et autres
Publié: (2026)
par: Lee, Kyungmi, et autres
Publié: (2026)
Empirical Measurements of AI Training Power Demand on a GPU-Accelerated Node
par: Latif, Imran, et autres
Publié: (2024)
par: Latif, Imran, et autres
Publié: (2024)
TLX: Hardware-Native, Evolvable MIMW GPU Compiler for Large-scale Production Environments
par: Guan, Yue, et autres
Publié: (2026)
par: Guan, Yue, et autres
Publié: (2026)
Edge GPU Aware Multiple AI Model Pipeline for Accelerated MRI Reconstruction and Analysis
par: Majeed, Ashiyana Abdul, et autres
Publié: (2025)
par: Majeed, Ashiyana Abdul, et autres
Publié: (2025)
GPU-Accelerated Simulated Oscillator Ising/Potts Machine Solving Combinatorial Optimization Problems
par: Gonul, Yilmaz Ege, et autres
Publié: (2025)
par: Gonul, Yilmaz Ege, et autres
Publié: (2025)
LLM-PRISM: Characterizing Silent Data Corruption from Permanent GPU Faults in LLM Training
par: Tyagi, Abhishek, et autres
Publié: (2026)
par: Tyagi, Abhishek, et autres
Publié: (2026)
Evaluation of GPU Video Encoder for Low-Latency Real-Time 4K UHD Encoding
par: Arunruangsirilert, Kasidis, et autres
Publié: (2025)
par: Arunruangsirilert, Kasidis, et autres
Publié: (2025)
HERO-Sign: Hierarchical Tuning and Efficient Compiler-Time GPU Optimizations for SPHINCS+ Signature Generation
par: Zhou, Yaoyun, et autres
Publié: (2025)
par: Zhou, Yaoyun, et autres
Publié: (2025)
EMSpice 3: Full-chip Temperature-Aware Multiphysics Electromigration and IR-Drop Analysis
par: Lu, Haotian, et autres
Publié: (2026)
par: Lu, Haotian, et autres
Publié: (2026)
Global and Local Attention-based Inception U-Net for Static IR Drop Prediction
par: Chen, Yilu, et autres
Publié: (2024)
par: Chen, Yilu, et autres
Publié: (2024)
A Host-SSD Collaborative Write Accelerator for LSM-Tree-Based Key-Value Stores
par: Kim, KiHwan, et autres
Publié: (2024)
par: Kim, KiHwan, et autres
Publié: (2024)
SSD Offloading for LLM Mixture-of-Experts Weights Considered Harmful in Energy Efficiency
par: Kyung, Kwanhee, et autres
Publié: (2025)
par: Kyung, Kwanhee, et autres
Publié: (2025)
e-GPU: An Open-Source and Configurable RISC-V Graphic Processing Unit for TinyAI Applications
par: Machetti, Simone, et autres
Publié: (2025)
par: Machetti, Simone, et autres
Publié: (2025)
A Comparison of the Cerebras Wafer-Scale Integration Technology with Nvidia GPU-based Systems for Artificial Intelligence
par: Kundu, Yudhishthira, et autres
Publié: (2025)
par: Kundu, Yudhishthira, et autres
Publié: (2025)
Scalable Processing-Near-Memory for 1M-Token LLM Inference: CXL-Enabled KV-Cache Management Beyond GPU Limits
par: Kim, Dowon, et autres
Publié: (2025)
par: Kim, Dowon, et autres
Publié: (2025)
Binary Neural Network Implementation for Handwritten Digit Recognition on FPGA
par: Ertörer, Emir Devlet, et autres
Publié: (2025)
par: Ertörer, Emir Devlet, et autres
Publié: (2025)
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference
par: Gu, Yufeng, et autres
Publié: (2025)
par: Gu, Yufeng, et autres
Publié: (2025)
R-HLS: An IR for Dynamic High-Level Synthesis and Memory Disambiguation based on Regions and State Edges
par: Metz, David, et autres
Publié: (2024)
par: Metz, David, et autres
Publié: (2024)
Guac: Energy-Aware and SSA-Based Generation of Coarse-Grained Merged Accelerators from LLVM-IR
par: Brumar, Iulian, et autres
Publié: (2024)
par: Brumar, Iulian, et autres
Publié: (2024)
A RISC-V Multicore and GPU SoC Platform with a Qualifiable Software Stack for Safety Critical Systems
par: Bonet, Marc Solé i, et autres
Publié: (2025)
par: Bonet, Marc Solé i, et autres
Publié: (2025)
Converting Binary Floating-Point Numbers to Shortest Decimal Strings: An Experimental Review
par: Gareau, Jaël Champagne, et autres
Publié: (2026)
par: Gareau, Jaël Champagne, et autres
Publié: (2026)
Documents similaires
-
Hardware vs. Software Implementation of Warp-Level Features in Vortex RISC-V GPU
par: Pu, Huanzhi, et autres
Publié: (2025) -
Inside VOLT: Designing an Open-Source GPU Compiler
par: Jeong, Shinnung, et autres
Publié: (2025) -
Towards Performance-Aware Allocation for Accelerated Machine Learning on GPU-SSD Systems
par: Gundawar, Ayush, et autres
Publié: (2024) -
Characterizing CPU-Induced Slowdowns in Multi-GPU LLM Inference
par: Chung, Euijun, et autres
Publié: (2026) -
CXL-GPU: Pushing GPU Memory Boundaries with the Integration of CXL Technologies
par: Gouk, Donghyun, et autres
Publié: (2025)