Gespeichert in:
| Hauptverfasser: | Wu, Yanran, Hua, Inez, Ding, Yi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2502.11256 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Not All Water Consumption Is Equal: A Water Stress Weighted Metric for Sustainable Computing
von: Wu, Yanran, et al.
Veröffentlicht: (2025)
von: Wu, Yanran, et al.
Veröffentlicht: (2025)
HDLxGraph: Bridging Large Language Models and HDL Repositories via HDL Graph Databases
von: Zheng, Pingqing, et al.
Veröffentlicht: (2025)
von: Zheng, Pingqing, et al.
Veröffentlicht: (2025)
GreenLLM: Disaggregating Large Language Model Serving on Heterogeneous GPUs for Lower Carbon Emissions
von: Shi, Tianyao, et al.
Veröffentlicht: (2024)
von: Shi, Tianyao, et al.
Veröffentlicht: (2024)
OPAL: Outlier-Preserved Microscaling Quantization Accelerator for Generative Large Language Models
von: Koo, Jahyun, et al.
Veröffentlicht: (2024)
von: Koo, Jahyun, et al.
Veröffentlicht: (2024)
Basis Selection: Low-Rank Decomposition of Pretrained Large Language Models for Target Applications
von: Li, Yang, et al.
Veröffentlicht: (2024)
von: Li, Yang, et al.
Veröffentlicht: (2024)
DeepRTL2: A Versatile Model for RTL-Related Tasks
von: Liu, Yi, et al.
Veröffentlicht: (2025)
von: Liu, Yi, et al.
Veröffentlicht: (2025)
DeepRTL: Bridging Verilog Understanding and Generation with a Unified Representation Model
von: Liu, Yi, et al.
Veröffentlicht: (2025)
von: Liu, Yi, et al.
Veröffentlicht: (2025)
Pimba: A Processing-in-Memory Acceleration for Post-Transformer Large Language Model Serving
von: Kim, Wonung, et al.
Veröffentlicht: (2025)
von: Kim, Wonung, et al.
Veröffentlicht: (2025)
Llumnix: Dynamic Scheduling for Large Language Model Serving
von: Sun, Biao, et al.
Veröffentlicht: (2024)
von: Sun, Biao, et al.
Veröffentlicht: (2024)
Understanding the Potential of FPGA-Based Spatial Acceleration for Large Language Model Inference
von: Chen, Hongzheng, et al.
Veröffentlicht: (2023)
von: Chen, Hongzheng, et al.
Veröffentlicht: (2023)
MX+: Pushing the Limits of Microscaling Formats for Efficient Large Language Model Serving
von: Lee, Jungi, et al.
Veröffentlicht: (2025)
von: Lee, Jungi, et al.
Veröffentlicht: (2025)
PASCAL: A Phase-Aware Scheduling Algorithm for Serving Reasoning-based Large Language Models
von: Cho, Eunyeong, et al.
Veröffentlicht: (2026)
von: Cho, Eunyeong, et al.
Veröffentlicht: (2026)
Speculative Decoding for Verilog: Speed and Quality, All in One
von: Xu, Changran, et al.
Veröffentlicht: (2025)
von: Xu, Changran, et al.
Veröffentlicht: (2025)
Characterizing the Behavior of Training Mamba-based State Space Models on GPUs
von: Baruah, Trinayan, et al.
Veröffentlicht: (2025)
von: Baruah, Trinayan, et al.
Veröffentlicht: (2025)
Is Finer Better? The Limits of Microscaling Formats in Large Language Models
von: Fasoli, Andrea, et al.
Veröffentlicht: (2026)
von: Fasoli, Andrea, et al.
Veröffentlicht: (2026)
D2S-FLOW: Automated Parameter Extraction from Datasheets for SPICE Model Generation Using Large Language Models
von: Chen, Hong Cai, et al.
Veröffentlicht: (2025)
von: Chen, Hong Cai, et al.
Veröffentlicht: (2025)
Serving Large Language Models on Huawei CloudMatrix384
von: Zuo, Pengfei, et al.
Veröffentlicht: (2025)
von: Zuo, Pengfei, et al.
Veröffentlicht: (2025)
Scaling Laws for Floating Point Quantization Training
von: Sun, Xingwu, et al.
Veröffentlicht: (2025)
von: Sun, Xingwu, et al.
Veröffentlicht: (2025)
Leveraging High-Level Synthesis and Large Language Models to Generate, Simulate, and Deploy a Uniform Random Number Generator Hardware Design
von: Meech, James T.
Veröffentlicht: (2023)
von: Meech, James T.
Veröffentlicht: (2023)
Chameleon: a Heterogeneous and Disaggregated Accelerator System for Retrieval-Augmented Language Models
von: Jiang, Wenqi, et al.
Veröffentlicht: (2023)
von: Jiang, Wenqi, et al.
Veröffentlicht: (2023)
Ultra Memory-Efficient On-FPGA Training of Transformers via Tensor-Compressed Optimization
von: Tian, Jiayi, et al.
Veröffentlicht: (2025)
von: Tian, Jiayi, et al.
Veröffentlicht: (2025)
Understanding and Mitigating Errors of LLM-Generated RTL Code
von: Zhang, Jiazheng, et al.
Veröffentlicht: (2025)
von: Zhang, Jiazheng, et al.
Veröffentlicht: (2025)
From Loop Nests to Silicon: Mapping AI Workloads onto AMD NPUs with MLIR-AIR
von: Wang, Erwei, et al.
Veröffentlicht: (2025)
von: Wang, Erwei, et al.
Veröffentlicht: (2025)
Orion: Characterizing and Programming Apple's Neural Engine for LLM Training and Inference
von: Kumaresan, Ramchand
Veröffentlicht: (2026)
von: Kumaresan, Ramchand
Veröffentlicht: (2026)
GRPO with State Mutations: Improving LLM-Based Hardware Test Plan Generation
von: Kochar, Dimple Vijay, et al.
Veröffentlicht: (2026)
von: Kochar, Dimple Vijay, et al.
Veröffentlicht: (2026)
Guaranteed Guess: A Language Modeling Approach for CISC-to-RISC Transpilation with Testing Guarantees
von: Heakl, Ahmed, et al.
Veröffentlicht: (2025)
von: Heakl, Ahmed, et al.
Veröffentlicht: (2025)
When Servers Meet Species: A Fab-to-Grave Lens on Computing's Biodiversity Impact
von: Shi, Tianyao, et al.
Veröffentlicht: (2025)
von: Shi, Tianyao, et al.
Veröffentlicht: (2025)
Hardwired-Neurons Language Processing Units as General-Purpose Cognitive Substrates
von: Liu, Yang, et al.
Veröffentlicht: (2025)
von: Liu, Yang, et al.
Veröffentlicht: (2025)
Lorecast: Layout-Aware Performance and Power Forecasting from Natural Language
von: Wang, Runzhi, et al.
Veröffentlicht: (2025)
von: Wang, Runzhi, et al.
Veröffentlicht: (2025)
Memory Access Characterization of Large Language Models in CPU Environment and its Potential Impacts
von: Banasik, Spencer
Veröffentlicht: (2025)
von: Banasik, Spencer
Veröffentlicht: (2025)
Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective
von: Li, Jinhao, et al.
Veröffentlicht: (2024)
von: Li, Jinhao, et al.
Veröffentlicht: (2024)
GFormer: Accelerating Large Language Models with Optimized Transformers on Gaudi Processors
von: Zhang, Chengming, et al.
Veröffentlicht: (2024)
von: Zhang, Chengming, et al.
Veröffentlicht: (2024)
HaLoRA: Hardware-aware Low-Rank Adaptation for Large Language Models Based on Hybrid Compute-in-Memory Architecture
von: Wu, Taiqiang, et al.
Veröffentlicht: (2025)
von: Wu, Taiqiang, et al.
Veröffentlicht: (2025)
LowRA: Accurate and Efficient LoRA Fine-Tuning of LLMs under 2 Bits
von: Zhou, Zikai, et al.
Veröffentlicht: (2025)
von: Zhou, Zikai, et al.
Veröffentlicht: (2025)
Non-Monotonic Latency in Apple MPS Decoding: KV Cache Interactions and Execution Regimes
von: Hendria, Willy Fitra
Veröffentlicht: (2026)
von: Hendria, Willy Fitra
Veröffentlicht: (2026)
A Survey on Hardware Accelerators for Large Language Models
von: Kachris, Christoforos
Veröffentlicht: (2024)
von: Kachris, Christoforos
Veröffentlicht: (2024)
Observation, Not Prediction: Conversation-Level Disaggregated Scheduling for Agentic Serving
von: Ding, Jianru, et al.
Veröffentlicht: (2026)
von: Ding, Jianru, et al.
Veröffentlicht: (2026)
FP6-LLM: Efficiently Serving Large Language Models Through FP6-Centric Algorithm-System Co-Design
von: Xia, Haojun, et al.
Veröffentlicht: (2024)
von: Xia, Haojun, et al.
Veröffentlicht: (2024)
Fine-Tuning Small Language Models for Domain-Specific AI: An Edge AI Perspective
von: Aralimatti, Rakshit, et al.
Veröffentlicht: (2025)
von: Aralimatti, Rakshit, et al.
Veröffentlicht: (2025)
Allo: A Programming Model for Composable Accelerator Design
von: Chen, Hongzheng, et al.
Veröffentlicht: (2024)
von: Chen, Hongzheng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Not All Water Consumption Is Equal: A Water Stress Weighted Metric for Sustainable Computing
von: Wu, Yanran, et al.
Veröffentlicht: (2025) -
HDLxGraph: Bridging Large Language Models and HDL Repositories via HDL Graph Databases
von: Zheng, Pingqing, et al.
Veröffentlicht: (2025) -
GreenLLM: Disaggregating Large Language Model Serving on Heterogeneous GPUs for Lower Carbon Emissions
von: Shi, Tianyao, et al.
Veröffentlicht: (2024) -
OPAL: Outlier-Preserved Microscaling Quantization Accelerator for Generative Large Language Models
von: Koo, Jahyun, et al.
Veröffentlicht: (2024) -
Basis Selection: Low-Rank Decomposition of Pretrained Large Language Models for Target Applications
von: Li, Yang, et al.
Veröffentlicht: (2024)