Guardado en:
| Autores principales: | Pan, Xiurui, Li, Endian, Li, Qiao, Liang, Shengwen, Shan, Yizhou, Zhou, Ke, Luo, Yingwei, Wang, Xiaolin, Zhang, Jie |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2409.04992 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A Cost-Effective Near-Storage Processing Solution for Offline Inference of Long-Context LLMs
por: Jang, Hongsun, et al.
Publicado: (2025)
por: Jang, Hongsun, et al.
Publicado: (2025)
HillInfer: Efficient Long-Context LLM Inference on the Edge with Hierarchical KV Eviction using SmartSSD
por: Sun, He, et al.
Publicado: (2026)
por: Sun, He, et al.
Publicado: (2026)
GPU Acceleration of TFHE-Based High-Precision Nonlinear Layers for Encrypted LLM Inference
por: Chen, Guoci, et al.
Publicado: (2026)
por: Chen, Guoci, et al.
Publicado: (2026)
LLMulator: Generalizable Cost Modeling for Dataflow Accelerators with Input-Adaptive Control Flow
por: Chang, Kaiyan, et al.
Publicado: (2025)
por: Chang, Kaiyan, et al.
Publicado: (2025)
A Novel Extensible Simulation Framework for CXL-Enabled Systems
por: An, Yuda, et al.
Publicado: (2024)
por: An, Yuda, et al.
Publicado: (2024)
FAST-Prefill: FPGA Accelerated Sparse Attention for Long Context LLM Prefill
por: Jayanth, Rakshith, et al.
Publicado: (2026)
por: Jayanth, Rakshith, et al.
Publicado: (2026)
Combating the Memory Walls: Optimization Pathways for Long-Context Agentic LLM Inference
por: Wu, Haoran, et al.
Publicado: (2025)
por: Wu, Haoran, et al.
Publicado: (2025)
An RDMA-First Object Storage System with SmartNIC Offload
por: Zhu, Yu, et al.
Publicado: (2025)
por: Zhu, Yu, et al.
Publicado: (2025)
UniCAIM: A Unified CAM/CIM Architecture with Static-Dynamic KV Cache Pruning for Efficient Long-Context LLM Inference
por: Xu, Weikai, et al.
Publicado: (2025)
por: Xu, Weikai, et al.
Publicado: (2025)
Cambricon-LLM: A Chiplet-Based Hybrid Architecture for On-Device Inference of 70B LLM
por: Yu, Zhongkai, et al.
Publicado: (2024)
por: Yu, Zhongkai, et al.
Publicado: (2024)
Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
por: Lin, Bin, et al.
Publicado: (2024)
por: Lin, Bin, et al.
Publicado: (2024)
Knowledge-Guided Attention-Inspired Learning for Task Offloading in Vehicle Edge Computing
por: Ma, Ke, et al.
Publicado: (2025)
por: Ma, Ke, et al.
Publicado: (2025)
Understanding Bottlenecks for Efficiently Serving LLM Inference With KV Offloading
por: Meng, William, et al.
Publicado: (2025)
por: Meng, William, et al.
Publicado: (2025)
SSD Offloading for LLM Mixture-of-Experts Weights Considered Harmful in Energy Efficiency
por: Kyung, Kwanhee, et al.
Publicado: (2025)
por: Kyung, Kwanhee, et al.
Publicado: (2025)
Graphitron: A Domain Specific Language for FPGA-based Graph Processing Accelerator Generation
por: Zhang, Xinmiao, et al.
Publicado: (2024)
por: Zhang, Xinmiao, et al.
Publicado: (2024)
PIMphony: Overcoming Bandwidth and Capacity Inefficiency in PIM-based Long-Context LLM Inference System
por: Kwon, Hyucksung, et al.
Publicado: (2024)
por: Kwon, Hyucksung, et al.
Publicado: (2024)
L3: DIMM-PIM Integrated Architecture and Coordination for Scalable Long-Context LLM Inference
por: Liu, Qingyuan, et al.
Publicado: (2025)
por: Liu, Qingyuan, et al.
Publicado: (2025)
Lifecycle Cost-Effectiveness Modeling for Redundancy-Enhanced Multi-Chiplet Architectures
por: Liu, Zizhen, et al.
Publicado: (2026)
por: Liu, Zizhen, et al.
Publicado: (2026)
PD-Swap: Prefill-Decode Logic Swapping for End-to-End LLM Inference on Edge FPGAs via Dynamic Partial Reconfiguration
por: Zhang, Yifan, et al.
Publicado: (2025)
por: Zhang, Yifan, et al.
Publicado: (2025)
NVLLM: A 3D NAND-Centric Architecture Enabling Edge on-Device LLM Inference
por: Hao, Mingbo, et al.
Publicado: (2026)
por: Hao, Mingbo, et al.
Publicado: (2026)
A Systematic Characterization of LLM Inference on GPUs
por: Wang, Haonan, et al.
Publicado: (2025)
por: Wang, Haonan, et al.
Publicado: (2025)
Convolutions Predictable Offloading to an Accelerator: Formalization and Optimization
por: Husson, Benjamin, et al.
Publicado: (2026)
por: Husson, Benjamin, et al.
Publicado: (2026)
TerEffic: Highly Efficient Ternary LLM Inference on FPGA
por: Yin, Chenyang, et al.
Publicado: (2025)
por: Yin, Chenyang, et al.
Publicado: (2025)
Bandwidth-Effective DRAM Cache for GPUs with Storage-Class Memory
por: Hong, Jeongmin, et al.
Publicado: (2024)
por: Hong, Jeongmin, et al.
Publicado: (2024)
AccLLM: Accelerating Long-Context LLM Inference Via Algorithm-Hardware Co-Design
por: Liang, Yanbiao, et al.
Publicado: (2025)
por: Liang, Yanbiao, et al.
Publicado: (2025)
Mapping Space Exploration for Multi-Chiplet Accelerators Targeting LLM Inference Serving Workloads
por: Li, Boyu, et al.
Publicado: (2025)
por: Li, Boyu, et al.
Publicado: (2025)
Resilient and Secure Programmable System-on-Chip Accelerator Offload
por: Gouveia, Inês Pinto, et al.
Publicado: (2024)
por: Gouveia, Inês Pinto, et al.
Publicado: (2024)
OffRAC: Offloading Through Remote Accelerator Calls
por: Yang, Ziyi, et al.
Publicado: (2025)
por: Yang, Ziyi, et al.
Publicado: (2025)
Harmonia: Algorithm-Hardware Co-Design for Memory- and Compute-Efficient BFP-based LLM Inference
por: Wang, Xinyu, et al.
Publicado: (2026)
por: Wang, Xinyu, et al.
Publicado: (2026)
FusionCIM: Accelerating LLM Inference with Fusion-Driven Computing-in-Memory Architecture
por: Xuan, Zihao, et al.
Publicado: (2026)
por: Xuan, Zihao, et al.
Publicado: (2026)
PermuteV: A Performant Side-channel-Resistant RISC-V Core Securing Edge AI Inference
por: Narkthong, Nuntipat, et al.
Publicado: (2025)
por: Narkthong, Nuntipat, et al.
Publicado: (2025)
A Switch-Centric In-Network Architecture for Accelerating LLM Inference in Shared-Memory Network
por: Jiang, Aojie, et al.
Publicado: (2026)
por: Jiang, Aojie, et al.
Publicado: (2026)
CHIME: Chiplet-based Heterogeneous Near-Memory Acceleration for Edge Multimodal LLM Inference
por: Chen, Yanru, et al.
Publicado: (2025)
por: Chen, Yanru, et al.
Publicado: (2025)
Rethinking LLM Inference Bottlenecks: Insights from Latent Attention and Mixture-of-Experts
por: Yun, Sungmin, et al.
Publicado: (2025)
por: Yun, Sungmin, et al.
Publicado: (2025)
Salca: A Sparsity-Aware Hardware Accelerator for Efficient Long-Context Attention Decoding
por: Fan, Wang, et al.
Publicado: (2026)
por: Fan, Wang, et al.
Publicado: (2026)
Breaking the HBM Bit Cost Barrier: Domain-Specific ECC for AI Inference Infrastructure
por: Xie, Rui, et al.
Publicado: (2025)
por: Xie, Rui, et al.
Publicado: (2025)
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Large Attention-Based Model Inference on Tile-Based Accelerators
por: Zhang, Chi, et al.
Publicado: (2026)
por: Zhang, Chi, et al.
Publicado: (2026)
Developing Cost-Effective Drones for 5G Non-Terrestrial Network Research and Experimentation
por: Cáceres, Carlos de Quinto, et al.
Publicado: (2024)
por: Cáceres, Carlos de Quinto, et al.
Publicado: (2024)
BitROM: Weight Reload-Free CiROM Architecture Towards Billion-Parameter 1.58-bit LLM Inference
por: Zhang, Wenlun, et al.
Publicado: (2025)
por: Zhang, Wenlun, et al.
Publicado: (2025)
T-MAN: Enabling End-to-End Low-Bit LLM Inference on NPUs via Unified Table Lookup
por: Wei, Jianyu, et al.
Publicado: (2025)
por: Wei, Jianyu, et al.
Publicado: (2025)
Ejemplares similares
-
A Cost-Effective Near-Storage Processing Solution for Offline Inference of Long-Context LLMs
por: Jang, Hongsun, et al.
Publicado: (2025) -
HillInfer: Efficient Long-Context LLM Inference on the Edge with Hierarchical KV Eviction using SmartSSD
por: Sun, He, et al.
Publicado: (2026) -
GPU Acceleration of TFHE-Based High-Precision Nonlinear Layers for Encrypted LLM Inference
por: Chen, Guoci, et al.
Publicado: (2026) -
LLMulator: Generalizable Cost Modeling for Dataflow Accelerators with Input-Adaptive Control Flow
por: Chang, Kaiyan, et al.
Publicado: (2025) -
A Novel Extensible Simulation Framework for CXL-Enabled Systems
por: An, Yuda, et al.
Publicado: (2024)