MIST: A Co-Design Framework for Heterogeneous, Multi-Stage LLM Inference
Fuente:
arXiv
Salvato in:
| Autori principali: | Bambhaniya, Abhimanyu Rajeshkumar, Wu, Hanjiang, Subramanian, Suvinay, Srinivasan, Sudarshan, Kundu, Souvik, Yazdanbakhsh, Amir, Elavazhagan, Midhilesh, Kumar, Madhu, Yu, Minlan, Raychowdhury, Arijit, Krishna, Tushar |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models
di: Bambhaniya, Abhimanyu, et al.
Pubblicazione: (2024)
di: Bambhaniya, Abhimanyu, et al.
Pubblicazione: (2024)
How Far Can Disaggregation Go? A Design-Space Exploration of Attention-FFN Disaggregation for Efficient MoE LLM Serving
di: Wu, Hanjiang, et al.
Pubblicazione: (2026)
di: Wu, Hanjiang, et al.
Pubblicazione: (2026)
Progressive Gradient Flow for Robust N:M Sparsity Training in Transformers
di: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Pubblicazione: (2024)
di: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Pubblicazione: (2024)
TACOS: Topology-Aware Collective Algorithm Synthesizer for Distributed Machine Learning
di: Won, William, et al.
Pubblicazione: (2023)
di: Won, William, et al.
Pubblicazione: (2023)
H3DFact: Heterogeneous 3D Integrated CIM for Factorization with Holographic Perceptual Representations
di: Wan, Zishen, et al.
Pubblicazione: (2024)
di: Wan, Zishen, et al.
Pubblicazione: (2024)
MicroScopiQ: Accelerating Foundational Models through Outlier-Aware Microscaling Quantization
di: Ramachandran, Akshat, et al.
Pubblicazione: (2024)
di: Ramachandran, Akshat, et al.
Pubblicazione: (2024)
Enabling Unstructured Sparse Acceleration on Structured Sparse Accelerators
di: Jeong, Geonhwa, et al.
Pubblicazione: (2024)
di: Jeong, Geonhwa, et al.
Pubblicazione: (2024)
CogSys: Efficient and Scalable Neurosymbolic Cognition System via Algorithm-Hardware Co-Design
di: Wan, Zishen, et al.
Pubblicazione: (2025)
di: Wan, Zishen, et al.
Pubblicazione: (2025)
REASON: Accelerating Probabilistic Logical Reasoning for Scalable Neuro-Symbolic Intelligence
di: Wan, Zishen, et al.
Pubblicazione: (2026)
di: Wan, Zishen, et al.
Pubblicazione: (2026)
Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns
di: Bambhaniya, Abhimanyu, et al.
Pubblicazione: (2026)
di: Bambhaniya, Abhimanyu, et al.
Pubblicazione: (2026)
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator
di: Ramachandran, Akshat, et al.
Pubblicazione: (2025)
di: Ramachandran, Akshat, et al.
Pubblicazione: (2025)
FRED: Flexible REduction-Distribution Interconnect and Communication Implementation for Wafer-Scale Distributed Training of DNN Models
di: Rashidi, Saeed, et al.
Pubblicazione: (2024)
di: Rashidi, Saeed, et al.
Pubblicazione: (2024)
Beyond Moore's Law: Harnessing the Redshift of Generative AI with Effective Hardware-Software Co-Design
di: Yazdanbakhsh, Amir
Pubblicazione: (2025)
di: Yazdanbakhsh, Amir
Pubblicazione: (2025)
Characterizing VLA Models: Identifying the Action Generation Bottleneck for Edge AI Architectures
di: Vishwanathan, Manoj, et al.
Pubblicazione: (2026)
di: Vishwanathan, Manoj, et al.
Pubblicazione: (2026)
Cross-Layer Design of Vector-Symbolic Computing: Bridging Cognition and Brain-Inspired Hardware Acceleration
di: Du, Shuting, et al.
Pubblicazione: (2025)
di: Du, Shuting, et al.
Pubblicazione: (2025)
NSFlow: An End-to-End FPGA Framework with Scalable Dataflow Architecture for Neuro-Symbolic AI
di: Yang, Hanchen, et al.
Pubblicazione: (2025)
di: Yang, Hanchen, et al.
Pubblicazione: (2025)
A 28nm 1.80Mb/mm2 Digital/Analog Hybrid SRAM-CIM Macro Using 2D-Weighted Capacitor Array for Complex Number Mac Operations
di: Konno, Shota, et al.
Pubblicazione: (2025)
di: Konno, Shota, et al.
Pubblicazione: (2025)
Towards Cognitive AI Systems: a Survey and Prospective on Neuro-Symbolic AI
di: Wan, Zishen, et al.
Pubblicazione: (2024)
di: Wan, Zishen, et al.
Pubblicazione: (2024)
SPARQLe: Sub-Precision Activation Representation for Quantized LLM Inference
di: Parvathy, Aradhana Mohan, et al.
Pubblicazione: (2026)
di: Parvathy, Aradhana Mohan, et al.
Pubblicazione: (2026)
Tao: Re-Thinking DL-based Microarchitecture Simulation
di: Pandey, Santosh, et al.
Pubblicazione: (2024)
di: Pandey, Santosh, et al.
Pubblicazione: (2024)
HARP: A Taxonomy for Heterogeneous and Hierarchical Processors for Mixed-reuse Workloads
di: Garg, Raveesh, et al.
Pubblicazione: (2025)
di: Garg, Raveesh, et al.
Pubblicazione: (2025)
MINISA: Minimal Instruction Set Architecture for Next-gen Reconfigurable Inference Accelerator
di: Tong, Jianming, et al.
Pubblicazione: (2026)
di: Tong, Jianming, et al.
Pubblicazione: (2026)
SATA: Sparsity-Aware Scheduling for Selective Token Attention
di: Fan, Zhenkun, et al.
Pubblicazione: (2026)
di: Fan, Zhenkun, et al.
Pubblicazione: (2026)
FEATHER: A Reconfigurable Accelerator with Data Reordering Support for Low-Cost On-Chip Dataflow Switching
di: Tong, Jianming, et al.
Pubblicazione: (2024)
di: Tong, Jianming, et al.
Pubblicazione: (2024)
MCMComm: Hardware-Software Co-Optimization for End-to-End Communication in Multi-Chip-Modules
di: Raj, Ritik, et al.
Pubblicazione: (2025)
di: Raj, Ritik, et al.
Pubblicazione: (2025)
Towards Efficient Neuro-Symbolic AI: From Workload Characterization to Hardware Architecture
di: Wan, Zishen, et al.
Pubblicazione: (2024)
di: Wan, Zishen, et al.
Pubblicazione: (2024)
FG-Attn: Leveraging Fine-Grained Sparsity In Diffusion Transformers
di: Durvasula, Sankeerth, et al.
Pubblicazione: (2025)
di: Durvasula, Sankeerth, et al.
Pubblicazione: (2025)
SCALE-Sim TPU: Validating and Extending SCALE-Sim for TPUs
di: Dang, Jingtian, et al.
Pubblicazione: (2026)
di: Dang, Jingtian, et al.
Pubblicazione: (2026)
CiMNet: Towards Joint Optimization for DNN Architecture and Configuration for Compute-In-Memory Hardware
di: Kundu, Souvik, et al.
Pubblicazione: (2024)
di: Kundu, Souvik, et al.
Pubblicazione: (2024)
PipeOrgan: Efficient Inter-operation Pipelining with Flexible Spatial Organization and Interconnects
di: Garg, Raveesh, et al.
Pubblicazione: (2024)
di: Garg, Raveesh, et al.
Pubblicazione: (2024)
3DGauCIM: Accelerating Static/Dynamic 3D Gaussian Splatting via Digital CIM for High Frame Rate Real-Time Edge Rendering
di: Huang, Wei-Hsing, et al.
Pubblicazione: (2025)
di: Huang, Wei-Hsing, et al.
Pubblicazione: (2025)
Axon: A novel systolic array architecture for improved run time and energy efficient GeMM and Conv operation with on-chip im2col
di: Nayan, Md Mizanur Rahaman, et al.
Pubblicazione: (2025)
di: Nayan, Md Mizanur Rahaman, et al.
Pubblicazione: (2025)
Instruction-Based Coordination of Heterogeneous Processing Units for Acceleration of DNN Inference
di: Petropoulos, Anastasios, et al.
Pubblicazione: (2025)
di: Petropoulos, Anastasios, et al.
Pubblicazione: (2025)
Hardware-based Heterogeneous Memory Management for Large Language Model Inference
di: Hwang, Soojin, et al.
Pubblicazione: (2025)
di: Hwang, Soojin, et al.
Pubblicazione: (2025)
NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing
di: Heo, Guseul, et al.
Pubblicazione: (2024)
di: Heo, Guseul, et al.
Pubblicazione: (2024)
OneDSE: A Unified Microprocessor Metric Prediction and Design Space Exploration Framework
di: Raj, Ritik, et al.
Pubblicazione: (2025)
di: Raj, Ritik, et al.
Pubblicazione: (2025)
GreenFPGA: Evaluating FPGAs as Environmentally Sustainable Computing Solutions
di: Sudarshan, Chetan Choppali, et al.
Pubblicazione: (2023)
di: Sudarshan, Chetan Choppali, et al.
Pubblicazione: (2023)
HPIM: Heterogeneous Processing-In-Memory-based Accelerator for Large Language Models Inference
di: Duan, Cenlin, et al.
Pubblicazione: (2025)
di: Duan, Cenlin, et al.
Pubblicazione: (2025)
MemExplorer: Navigating the Heterogeneous Memory Design Space for Agentic Inference NPUs
di: Wu, Haoran, et al.
Pubblicazione: (2026)
di: Wu, Haoran, et al.
Pubblicazione: (2026)
Algorithm-Hardware Co-Design of Distribution-Aware Logarithmic-Posit Encodings for Efficient DNN Inference
di: Ramachandran, Akshat, et al.
Pubblicazione: (2024)
di: Ramachandran, Akshat, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models
di: Bambhaniya, Abhimanyu, et al.
Pubblicazione: (2024) -
How Far Can Disaggregation Go? A Design-Space Exploration of Attention-FFN Disaggregation for Efficient MoE LLM Serving
di: Wu, Hanjiang, et al.
Pubblicazione: (2026) -
Progressive Gradient Flow for Robust N:M Sparsity Training in Transformers
di: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Pubblicazione: (2024) -
TACOS: Topology-Aware Collective Algorithm Synthesizer for Distributed Machine Learning
di: Won, William, et al.
Pubblicazione: (2023) -
H3DFact: Heterogeneous 3D Integrated CIM for Factorization with Holographic Perceptual Representations
di: Wan, Zishen, et al.
Pubblicazione: (2024)