A Review on Proprietary Accelerators for Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Park, Sihyeong, Lee, Jemin, Kim, Byung-Soo, Jeon, Seokhun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Quantitative Analysis and Guidelines of Data Streaming Accelerator in Modern Intel Xeon Scalable Processors
von: Kuper, Reese, et al.
Veröffentlicht: (2023)
von: Kuper, Reese, et al.
Veröffentlicht: (2023)
ONNXim: A Fast, Cycle-level Multi-core NPU Simulator
von: Ham, Hyungkyu, et al.
Veröffentlicht: (2024)
von: Ham, Hyungkyu, et al.
Veröffentlicht: (2024)
An Analytical Cost Model for Fast Evaluation of Multiple Compute-Engine CNN Accelerators
von: Qararyah, Fareed, et al.
Veröffentlicht: (2025)
von: Qararyah, Fareed, et al.
Veröffentlicht: (2025)
Adaptive Cache Pollution Control for Large Language Model Inference Workloads Using Temporal CNN-Based Prediction and Priority-Aware Replacement
von: Liu, Songze, et al.
Veröffentlicht: (2025)
von: Liu, Songze, et al.
Veröffentlicht: (2025)
Accelerating Transistor-Level Simulation of Integrated Circuits via Equivalence of RC Long-Chain Structures
von: Tang, Ruibai, et al.
Veröffentlicht: (2025)
von: Tang, Ruibai, et al.
Veröffentlicht: (2025)
Gem5-AcceSys: Enabling System-Level Exploration of Standard Interconnects for Novel Accelerators
von: Liu, Qunyou, et al.
Veröffentlicht: (2025)
von: Liu, Qunyou, et al.
Veröffentlicht: (2025)
JSPIM: A Skew-Aware PIM Accelerator for High-Performance Databases Join and Select Operations
von: Tajdari, Sabiha, et al.
Veröffentlicht: (2025)
von: Tajdari, Sabiha, et al.
Veröffentlicht: (2025)
Regular-Dead on Arrival: Characterizing and Protecting Against Dead-Entry TLB Misses in GPU Microarchitectures
von: Anik, Shafayat Mowla, et al.
Veröffentlicht: (2026)
von: Anik, Shafayat Mowla, et al.
Veröffentlicht: (2026)
ACALSim: A Scalable Parallel Simulation Framework for High-Performance System Design Space Exploration
von: Lin, Wei-Fen, et al.
Veröffentlicht: (2026)
von: Lin, Wei-Fen, et al.
Veröffentlicht: (2026)
ChatNeuroSim: An LLM Agent Framework for Automated Compute-in-Memory Accelerator Deployment and Optimization
von: Lee, Ming-Yen, et al.
Veröffentlicht: (2026)
von: Lee, Ming-Yen, et al.
Veröffentlicht: (2026)
Data-Driven Power Modeling and Monitoring via Hardware Performance Counter Tracking
von: Mazzola, Sergio, et al.
Veröffentlicht: (2025)
von: Mazzola, Sergio, et al.
Veröffentlicht: (2025)
PipeWeave: Synergizing Analytical and Learning Models for Unified GPU Performance Prediction
von: Zhang, Kaixuan, et al.
Veröffentlicht: (2026)
von: Zhang, Kaixuan, et al.
Veröffentlicht: (2026)
WaveTune: Wave-aware Bilinear Modeling for Efficient GPU Kernel Auto-tuning
von: Zhang, Kaixuan, et al.
Veröffentlicht: (2026)
von: Zhang, Kaixuan, et al.
Veröffentlicht: (2026)
Cleaning up the Mess: Re-Evaluating the Real-System Modeling Accuracy of Ramulator 2.0
von: Bostanci, F. Nisa, et al.
Veröffentlicht: (2025)
von: Bostanci, F. Nisa, et al.
Veröffentlicht: (2025)
AI Load Dynamics--A Power Electronics Perspective
von: Li, Yuzhuo, et al.
Veröffentlicht: (2025)
von: Li, Yuzhuo, et al.
Veröffentlicht: (2025)
A$^3$PIM: An Automated, Analytic and Accurate Processing-in-Memory Offloader
von: Jiang, Qingcai, et al.
Veröffentlicht: (2024)
von: Jiang, Qingcai, et al.
Veröffentlicht: (2024)
SCALE-Sim v3: A modular cycle-accurate systolic accelerator simulator for end-to-end system analysis
von: Raj, Ritik, et al.
Veröffentlicht: (2025)
von: Raj, Ritik, et al.
Veröffentlicht: (2025)
Heterogeneous Memory Benchmarking Toolkit
von: Ghaemi, Golsana, et al.
Veröffentlicht: (2025)
von: Ghaemi, Golsana, et al.
Veröffentlicht: (2025)
Simulation-Driven Evaluation of Chiplet-Based Architectures Using VisualSim
von: Ali, Wajid, et al.
Veröffentlicht: (2025)
von: Ali, Wajid, et al.
Veröffentlicht: (2025)
Recurrent CircuitSAT Sampling for Sequential Circuits
von: Ardakani, Arash, et al.
Veröffentlicht: (2025)
von: Ardakani, Arash, et al.
Veröffentlicht: (2025)
Introducing the Arm-membench Throughput Benchmark
von: Burth, Cyrill, et al.
Veröffentlicht: (2025)
von: Burth, Cyrill, et al.
Veröffentlicht: (2025)
Enhancing software-hardware co-design for HEP by low-overhead profiling of single- and multi-threaded programs on diverse architectures with Adaptyst
von: Graczyk, Maksymilian, et al.
Veröffentlicht: (2025)
von: Graczyk, Maksymilian, et al.
Veröffentlicht: (2025)
SAHM: State-Aware Heterogeneous Multicore for Single-Thread Performance
von: Wadle, Shayne, et al.
Veröffentlicht: (2025)
von: Wadle, Shayne, et al.
Veröffentlicht: (2025)
Characterizing and Optimizing Realistic Workloads on a Commercial Compute-in-SRAM Device
von: Zhang, Niansong, et al.
Veröffentlicht: (2025)
von: Zhang, Niansong, et al.
Veröffentlicht: (2025)
OmniSim: Simulating Hardware with C Speed and RTL Accuracy for High-Level Synthesis Designs
von: Sarkar, Rishov, et al.
Veröffentlicht: (2025)
von: Sarkar, Rishov, et al.
Veröffentlicht: (2025)
GeneTEK: Low-power, high-performance and scalable FPGA architecture for exact unit-cost edit distance
von: Espinosa, Elena, et al.
Veröffentlicht: (2025)
von: Espinosa, Elena, et al.
Veröffentlicht: (2025)
Análisis de rendimiento y eficiencia energética en el cluster Raspberry Pi Cronos
von: Semken, Martha, et al.
Veröffentlicht: (2025)
von: Semken, Martha, et al.
Veröffentlicht: (2025)
How to Increase Energy Efficiency with a Single Linux Command
von: Jelvani, Alborz, et al.
Veröffentlicht: (2025)
von: Jelvani, Alborz, et al.
Veröffentlicht: (2025)
DEER: Deep Runahead for Instruction Prefetching on Modern Mobile Workloads
von: Vahdatniya, Parmida, et al.
Veröffentlicht: (2025)
von: Vahdatniya, Parmida, et al.
Veröffentlicht: (2025)
Enhancing Instruction Prefetching via Cache and TLB Management
von: Jamet, Alexandre Valentin, et al.
Veröffentlicht: (2026)
von: Jamet, Alexandre Valentin, et al.
Veröffentlicht: (2026)
ETM2: Empowering Traditional Memory Bandwidth Regulation using ETM
von: Zuepke, Alexander, et al.
Veröffentlicht: (2026)
von: Zuepke, Alexander, et al.
Veröffentlicht: (2026)
Towards CPU Performance Prediction: New Challenge Benchmark Dataset and Novel Approach
von: Liu, Xiaoman
Veröffentlicht: (2024)
von: Liu, Xiaoman
Veröffentlicht: (2024)
SPEC CPU2026: Characterization, Representativeness, and Cross-Suite Comparison
von: Li, Ruihao, et al.
Veröffentlicht: (2026)
von: Li, Ruihao, et al.
Veröffentlicht: (2026)
Single 32-bit Sub-Channel DDR5 DIMMs: Architecture, Performance Bounds, and Standardisation
von: Ke, Chih-Hua
Veröffentlicht: (2026)
von: Ke, Chih-Hua
Veröffentlicht: (2026)
Makinote: An FPGA-Based HW/SW Platform for Pre-Silicon Emulation of RISC-V Designs
von: Perdomo, Elias, et al.
Veröffentlicht: (2024)
von: Perdomo, Elias, et al.
Veröffentlicht: (2024)
LightningSimV2: Faster and Scalable Simulation for High-Level Synthesis via Graph Compilation and Optimization
von: Sarkar, Rishov, et al.
Veröffentlicht: (2024)
von: Sarkar, Rishov, et al.
Veröffentlicht: (2024)
Range, Not Precision: Block-Floating-Point Half-Precision FFT and SAR Imaging on Apple Silicon
von: Bergach, Mohamed Amine
Veröffentlicht: (2026)
von: Bergach, Mohamed Amine
Veröffentlicht: (2026)
CXL-Interference: Analysis and Characterization in Modern Computer Systems
von: Mao, Shunyu, et al.
Veröffentlicht: (2024)
von: Mao, Shunyu, et al.
Veröffentlicht: (2024)
OPTIMA: Design-Space Exploration of Discharge-Based In-SRAM Computing: Quantifying Energy-Accuracy Trade-Offs
von: Seyedfaraji, Saeed, et al.
Veröffentlicht: (2024)
von: Seyedfaraji, Saeed, et al.
Veröffentlicht: (2024)
The Bicameral Cache: a split cache for vector architectures
von: Rebolledo, Susana, et al.
Veröffentlicht: (2024)
von: Rebolledo, Susana, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
A Quantitative Analysis and Guidelines of Data Streaming Accelerator in Modern Intel Xeon Scalable Processors
von: Kuper, Reese, et al.
Veröffentlicht: (2023) -
ONNXim: A Fast, Cycle-level Multi-core NPU Simulator
von: Ham, Hyungkyu, et al.
Veröffentlicht: (2024) -
An Analytical Cost Model for Fast Evaluation of Multiple Compute-Engine CNN Accelerators
von: Qararyah, Fareed, et al.
Veröffentlicht: (2025) -
Adaptive Cache Pollution Control for Large Language Model Inference Workloads Using Temporal CNN-Based Prediction and Priority-Aware Replacement
von: Liu, Songze, et al.
Veröffentlicht: (2025) -
Accelerating Transistor-Level Simulation of Integrated Circuits via Equivalence of RC Long-Chain Structures
von: Tang, Ruibai, et al.
Veröffentlicht: (2025)