Portable Targeted Sampling Framework Using LLVM
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Qiu, Zhantong, Samani, Mahyar, Lowe-Power, Jason |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Toward Reproducible and Standardized Computer Architecture Simulation with gem5
von: Pai, Kunal, et al.
Veröffentlicht: (2025)
von: Pai, Kunal, et al.
Veröffentlicht: (2025)
Choreographer: A Full-System Framework for Fine-Grained Tasks in Cache Hierarchies
von: Nguyen, Hoa, et al.
Veröffentlicht: (2025)
von: Nguyen, Hoa, et al.
Veröffentlicht: (2025)
Potential and Limitation of High-Frequency Cores and Caches
von: Pai, Kunal, et al.
Veröffentlicht: (2024)
von: Pai, Kunal, et al.
Veröffentlicht: (2024)
Pickle Prefetcher: Programmable and Scalable Last-Level Cache Prefetcher
von: Nguyen, Hoa, et al.
Veröffentlicht: (2025)
von: Nguyen, Hoa, et al.
Veröffentlicht: (2025)
LP-Spec: Leveraging LPDDR PIM for Efficient LLM Mobile Speculative Inference with Architecture-Dataflow Co-Optimization
von: He, Siyuan, et al.
Veröffentlicht: (2025)
von: He, Siyuan, et al.
Veröffentlicht: (2025)
ChipLight: Cross-Layer Optimization of Chiplet Design with Optical Interconnects for LLM Training
von: Bai, Kangbo, et al.
Veröffentlicht: (2026)
von: Bai, Kangbo, et al.
Veröffentlicht: (2026)
HammerSim: A System-Level Tool to Model RowHammer
von: Goswami, Kaustav, et al.
Veröffentlicht: (2026)
von: Goswami, Kaustav, et al.
Veröffentlicht: (2026)
Guac: Energy-Aware and SSA-Based Generation of Coarse-Grained Merged Accelerators from LLVM-IR
von: Brumar, Iulian, et al.
Veröffentlicht: (2024)
von: Brumar, Iulian, et al.
Veröffentlicht: (2024)
SILVIA: Automated Superword-Level Parallelism Exploitation via HLS-Specific LLVM Passes for Compute-Intensive FPGA Accelerators
von: Brignone, Giovanni, et al.
Veröffentlicht: (2024)
von: Brignone, Giovanni, et al.
Veröffentlicht: (2024)
TDRAM: Tag-enhanced DRAM for Efficient Caching
von: Babaie, Maryam, et al.
Veröffentlicht: (2024)
von: Babaie, Maryam, et al.
Veröffentlicht: (2024)
Space-Control: Process-Level Isolation for Sharing CXL-based Disaggregated Memory
von: Goswami, Kaustav, et al.
Veröffentlicht: (2026)
von: Goswami, Kaustav, et al.
Veröffentlicht: (2026)
CXL-ClusterSim: Modeling CXL-based Disaggregated Memory Cluster for Pooling and Sharing using gem5 and SST
von: Goswami, Kaustav, et al.
Veröffentlicht: (2026)
von: Goswami, Kaustav, et al.
Veröffentlicht: (2026)
Retrofitting Control Flow Graphs in LLVM IR for Auto Vectorization
von: Fang, Shihan, et al.
Veröffentlicht: (2025)
von: Fang, Shihan, et al.
Veröffentlicht: (2025)
Hardware Acceleration in Portable MRIs: State of the Art and Future Prospects
von: Habsi, Omar Al, et al.
Veröffentlicht: (2025)
von: Habsi, Omar Al, et al.
Veröffentlicht: (2025)
Spec2Cov: An Agentic Framework for Code Coverage Closure of Digital Hardware Designs
von: Lowe, Sean, et al.
Veröffentlicht: (2026)
von: Lowe, Sean, et al.
Veröffentlicht: (2026)
Leveraging Compute-in-Memory for Efficient Generative Model Inference in TPUs
von: Zhu, Zhantong, et al.
Veröffentlicht: (2025)
von: Zhu, Zhantong, et al.
Veröffentlicht: (2025)
Holistic Optimization Framework for FPGA Accelerators
von: Pouget, Stéphane, et al.
Veröffentlicht: (2025)
von: Pouget, Stéphane, et al.
Veröffentlicht: (2025)
FIFOAdvisor: A DSE Framework for Automated FIFO Sizing of High-Level Synthesis Designs
von: Abi-Karam, Stefan, et al.
Veröffentlicht: (2025)
von: Abi-Karam, Stefan, et al.
Veröffentlicht: (2025)
Manticore: Hardware-Accelerated RTL Simulation with Static Bulk-Synchronous Parallelism
von: Emami, Mahyar, et al.
Veröffentlicht: (2023)
von: Emami, Mahyar, et al.
Veröffentlicht: (2023)
WebAssembly on Resource-Constrained IoT Devices: Performance, Efficiency, and Portability
von: Has, Mislav, et al.
Veröffentlicht: (2025)
von: Has, Mislav, et al.
Veröffentlicht: (2025)
Stream-HLS: Towards Automatic Dataflow Acceleration
von: Basalama, Suhail, et al.
Veröffentlicht: (2025)
von: Basalama, Suhail, et al.
Veröffentlicht: (2025)
Branch Target Buffer Reverse Engineering on Arm
von: Wan, Junpeng
Veröffentlicht: (2024)
von: Wan, Junpeng
Veröffentlicht: (2024)
Compromising the Intelligence of Modern DNNs: On the Effectiveness of Targeted RowPress
von: Zhou, Ranyang, et al.
Veröffentlicht: (2024)
von: Zhou, Ranyang, et al.
Veröffentlicht: (2024)
A Lightweight Algorithm for Classifying Ex Vivo Tissues Samples
von: Li, Tzu-Hao, et al.
Veröffentlicht: (2024)
von: Li, Tzu-Hao, et al.
Veröffentlicht: (2024)
FTTN: Feature-Targeted Testing for Numerical Properties of NVIDIA & AMD Matrix Accelerators
von: Li, Xinyi, et al.
Veröffentlicht: (2024)
von: Li, Xinyi, et al.
Veröffentlicht: (2024)
Automatic Hardware Pragma Insertion in High-Level Synthesis: A Non-Linear Programming Approach
von: Pouget, Stéphane, et al.
Veröffentlicht: (2024)
von: Pouget, Stéphane, et al.
Veröffentlicht: (2024)
Mapping Space Exploration for Multi-Chiplet Accelerators Targeting LLM Inference Serving Workloads
von: Li, Boyu, et al.
Veröffentlicht: (2025)
von: Li, Boyu, et al.
Veröffentlicht: (2025)
HWE-Bench: Can Language Models Perform Board-level Schematic Designs?
von: Qiu, Weibo, et al.
Veröffentlicht: (2026)
von: Qiu, Weibo, et al.
Veröffentlicht: (2026)
Demystifying FPGA Hard NoC Performance
von: Liu, Sihao, et al.
Veröffentlicht: (2025)
von: Liu, Sihao, et al.
Veröffentlicht: (2025)
Real-time Object Detection and Associated Hardware Accelerators Targeting Autonomous Vehicles: A Review
von: Sali, Safa, et al.
Veröffentlicht: (2025)
von: Sali, Safa, et al.
Veröffentlicht: (2025)
BlissCam: Boosting Eye Tracking Efficiency with Learned In-Sensor Sparse Sampling
von: Feng, Yu, et al.
Veröffentlicht: (2024)
von: Feng, Yu, et al.
Veröffentlicht: (2024)
GPU Performance Portability needs Autotuning
von: Ringlein, Burkhard, et al.
Veröffentlicht: (2025)
von: Ringlein, Burkhard, et al.
Veröffentlicht: (2025)
LASANA: Large-Scale Surrogate Modeling for Analog Neuromorphic Architecture Exploration
von: Ho, Jason, et al.
Veröffentlicht: (2025)
von: Ho, Jason, et al.
Veröffentlicht: (2025)
Reconfigurable Stream Network Architecture
von: Wang, Chengyue, et al.
Veröffentlicht: (2024)
von: Wang, Chengyue, et al.
Veröffentlicht: (2024)
CPU Simulation Using Two-Phase Stratified Sampling
von: Ekman, Magnus
Veröffentlicht: (2026)
von: Ekman, Magnus
Veröffentlicht: (2026)
From Indiscriminate to Targeted: Efficient RTL Verification via Functionally Key Signal-Driven LLM Assertion Generation
von: Wang, Yonghao, et al.
Veröffentlicht: (2026)
von: Wang, Yonghao, et al.
Veröffentlicht: (2026)
SD-Acc: Accelerating Stable Diffusion through Phase-aware Sampling and Hardware Co-Optimizations
von: Wang, Zhican, et al.
Veröffentlicht: (2025)
von: Wang, Zhican, et al.
Veröffentlicht: (2025)
FuseFPS: Accelerating Farthest Point Sampling with Fusing KD-tree Construction for Point Clouds
von: Han, Meng, et al.
Veröffentlicht: (2023)
von: Han, Meng, et al.
Veröffentlicht: (2023)
QED: Scalable Verification of Hardware Memory Consistency
von: Ravi, Gokulan, et al.
Veröffentlicht: (2024)
von: Ravi, Gokulan, et al.
Veröffentlicht: (2024)
DEFA: Efficient Deformable Attention Acceleration via Pruning-Assisted Grid-Sampling and Multi-Scale Parallel Processing
von: Xu, Yansong, et al.
Veröffentlicht: (2024)
von: Xu, Yansong, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Toward Reproducible and Standardized Computer Architecture Simulation with gem5
von: Pai, Kunal, et al.
Veröffentlicht: (2025) -
Choreographer: A Full-System Framework for Fine-Grained Tasks in Cache Hierarchies
von: Nguyen, Hoa, et al.
Veröffentlicht: (2025) -
Potential and Limitation of High-Frequency Cores and Caches
von: Pai, Kunal, et al.
Veröffentlicht: (2024) -
Pickle Prefetcher: Programmable and Scalable Last-Level Cache Prefetcher
von: Nguyen, Hoa, et al.
Veröffentlicht: (2025) -
LP-Spec: Leveraging LPDDR PIM for Efficient LLM Mobile Speculative Inference with Architecture-Dataflow Co-Optimization
von: He, Siyuan, et al.
Veröffentlicht: (2025)