AMMA: A Multi-Chiplet Memory-Centric Architecture for Low-Latency 1M Context Attention Serving
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yu, Zhongkai, Ye, Haotian, Zhou, Chenyang, Venkatachalam, Ohm Rishabh, Pan, Zaifeng, Hu, Zhengding, Kim, Junsung, Ro, Won Woo, Tsai, Po-An, Pei, Shuyi, Kang, Yangwook, Ding, Yufei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Patterns behind Chaos: Forecasting Data Movement for Efficient Large-Scale MoE LLM Inference
von: Yu, Zhongkai, et al.
Veröffentlicht: (2025)
von: Yu, Zhongkai, et al.
Veröffentlicht: (2025)
ChipBench: A Next-Step Benchmark for Evaluating LLM Performance in AI-Aided Chip Design
von: Yu, Zhongkai, et al.
Veröffentlicht: (2026)
von: Yu, Zhongkai, et al.
Veröffentlicht: (2026)
Chiplet Actuary: A Quantitative Cost Model and Multi-Chiplet Architecture Exploration
von: Feng, Yinxiao, et al.
Veröffentlicht: (2022)
von: Feng, Yinxiao, et al.
Veröffentlicht: (2022)
RapidChiplet: A Toolchain for Rapid Design Space Exploration of Chiplet Architectures
von: Iff, Patrick, et al.
Veröffentlicht: (2023)
von: Iff, Patrick, et al.
Veröffentlicht: (2023)
zhongkaiyu/moe_exp_placement: Case Study2 for ISCA 2026 AE
von: Yu, Zhongkai, et al.
Veröffentlicht: (2026)
von: Yu, Zhongkai, et al.
Veröffentlicht: (2026)
Cambricon-LLM: A Chiplet-Based Hybrid Architecture for On-Device Inference of 70B LLM
von: Yu, Zhongkai, et al.
Veröffentlicht: (2024)
von: Yu, Zhongkai, et al.
Veröffentlicht: (2024)
Chiplet Cloud: Building AI Supercomputers for Serving Large Generative Language Models
von: Peng, Huwan, et al.
Veröffentlicht: (2023)
von: Peng, Huwan, et al.
Veröffentlicht: (2023)
ChipMATE: Multi-Agent Training via Reinforcement Learning for Enhanced RTL Generation
von: Yu, Zhongkai, et al.
Veröffentlicht: (2026)
von: Yu, Zhongkai, et al.
Veröffentlicht: (2026)
The Survey of Chiplet-based Integrated Architecture: An EDA perspective
von: Chen, Shixin, et al.
Veröffentlicht: (2024)
von: Chen, Shixin, et al.
Veröffentlicht: (2024)
Mapping Space Exploration for Multi-Chiplet Accelerators Targeting LLM Inference Serving Workloads
von: Li, Boyu, et al.
Veröffentlicht: (2025)
von: Li, Boyu, et al.
Veröffentlicht: (2025)
Lifecycle Cost-Effectiveness Modeling for Redundancy-Enhanced Multi-Chiplet Architectures
von: Liu, Zizhen, et al.
Veröffentlicht: (2026)
von: Liu, Zizhen, et al.
Veröffentlicht: (2026)
Chiplets on Wheels: Review Paper on Holistic Chiplet Solutions for Autonomous Vehicles
von: Narashiman, Swathi, et al.
Veröffentlicht: (2024)
von: Narashiman, Swathi, et al.
Veröffentlicht: (2024)
ECO-CHIP: Estimation of Carbon Footprint of Chiplet-based Architectures for Sustainable VLSI
von: Sudarshan, Chetan Choppali, et al.
Veröffentlicht: (2023)
von: Sudarshan, Chetan Choppali, et al.
Veröffentlicht: (2023)
Chiplet-Gym: Optimizing Chiplet-based AI Accelerator Design with Reinforcement Learning
von: Mishty, Kaniz, et al.
Veröffentlicht: (2024)
von: Mishty, Kaniz, et al.
Veröffentlicht: (2024)
Simulation-Driven Evaluation of Chiplet-Based Architectures Using VisualSim
von: Ali, Wajid, et al.
Veröffentlicht: (2025)
von: Ali, Wajid, et al.
Veröffentlicht: (2025)
TLX: Hardware-Native, Evolvable MIMW GPU Compiler for Large-scale Production Environments
von: Guan, Yue, et al.
Veröffentlicht: (2026)
von: Guan, Yue, et al.
Veröffentlicht: (2026)
Hemlet: A Heterogeneous Compute-in-Memory Chiplet Architecture for Vision Transformers with Group-Level Parallelism
von: Wang, Cong, et al.
Veröffentlicht: (2025)
von: Wang, Cong, et al.
Veröffentlicht: (2025)
FoldedHexaTorus: An Inter-Chiplet Interconnect Topology for Chiplet-based Systems using Organic and Glass Substrates
von: Iff, Patrick, et al.
Veröffentlicht: (2025)
von: Iff, Patrick, et al.
Veröffentlicht: (2025)
THERMOS: Thermally-Aware Multi-Objective Scheduling of AI Workloads on Heterogeneous Multi-Chiplet PIM Architectures
von: Kanani, Alish, et al.
Veröffentlicht: (2025)
von: Kanani, Alish, et al.
Veröffentlicht: (2025)
MFIT: Multi-Fidelity Thermal Modeling for 2.5D and 3D Multi-Chiplet Architectures
von: Pfromm, Lukas, et al.
Veröffentlicht: (2024)
von: Pfromm, Lukas, et al.
Veröffentlicht: (2024)
Mozart: Modularized and Efficient MoE Training on 3.5D Wafer-Scale Chiplet Architectures
von: Luo, Shuqing, et al.
Veröffentlicht: (2026)
von: Luo, Shuqing, et al.
Veröffentlicht: (2026)
Designing High-Performance and Thermally Feasible Multi-Chiplet Architectures enabled by Non-bendable Glass Interposer
von: Sharma, Harsh, et al.
Veröffentlicht: (2025)
von: Sharma, Harsh, et al.
Veröffentlicht: (2025)
Garibaldi: A Pairwise Instruction-Data Management for Enhancing Shared Last-Level Cache Performance in Server Workloads
von: Kwon, Jaewon, et al.
Veröffentlicht: (2025)
von: Kwon, Jaewon, et al.
Veröffentlicht: (2025)
ChipletQuake: On-die Digital Impedance Sensing for Chiplet and Interposer Verification
von: Monfared, Saleh Khalaj, et al.
Veröffentlicht: (2025)
von: Monfared, Saleh Khalaj, et al.
Veröffentlicht: (2025)
ADOR: A Design Exploration Framework for LLM Serving with Enhanced Latency and Throughput
von: Kim, Junsoo, et al.
Veröffentlicht: (2025)
von: Kim, Junsoo, et al.
Veröffentlicht: (2025)
Link Quality Aware Pathfinding for Chiplet Interconnects
von: Yen, Aaron, et al.
Veröffentlicht: (2026)
von: Yen, Aaron, et al.
Veröffentlicht: (2026)
PlaceIT: Placement-based Inter-Chiplet Interconnect Topologies
von: Iff, Patrick, et al.
Veröffentlicht: (2025)
von: Iff, Patrick, et al.
Veröffentlicht: (2025)
Designing Secure Interconnects for Modern Microelectronics: From SoCs to Emerging Chiplet-Based Architectures
von: Halder, Dipal
Veröffentlicht: (2023)
von: Halder, Dipal
Veröffentlicht: (2023)
DCRA: A Distributed Chiplet-based Reconfigurable Architecture for Irregular Applications
von: Orenes-Vera, Marcelo, et al.
Veröffentlicht: (2023)
von: Orenes-Vera, Marcelo, et al.
Veröffentlicht: (2023)
A Heterogeneous Chiplet Architecture for Accelerating End-to-End Transformer Models
von: Sharma, Harsh, et al.
Veröffentlicht: (2023)
von: Sharma, Harsh, et al.
Veröffentlicht: (2023)
Toward Open-Source Chiplets for HPC and AI: Occamy and Beyond
von: Scheffler, Paul, et al.
Veröffentlicht: (2025)
von: Scheffler, Paul, et al.
Veröffentlicht: (2025)
Hecaton: Training Large Language Models with Scalable Chiplet Systems
von: Huang, Zongle, et al.
Veröffentlicht: (2024)
von: Huang, Zongle, et al.
Veröffentlicht: (2024)
A System Architecture for Low Latency Multiprogramming Quantum Computing
von: Zhao, Yilun, et al.
Veröffentlicht: (2026)
von: Zhao, Yilun, et al.
Veröffentlicht: (2026)
Monad: Towards Cost-effective Specialization for Chiplet-based Spatial Accelerators
von: Hao, Xiaochen, et al.
Veröffentlicht: (2023)
von: Hao, Xiaochen, et al.
Veröffentlicht: (2023)
Designing Reconfigurable Interconnection Network of Heterogeneous Chiplets Using Kalman Filter
von: Biglari, Siamak, et al.
Veröffentlicht: (2024)
von: Biglari, Siamak, et al.
Veröffentlicht: (2024)
Educating for Hardware Specialization in the Chiplet Era: A Path for the HPC Community
von: Yoshii, Kazutomo, et al.
Veröffentlicht: (2024)
von: Yoshii, Kazutomo, et al.
Veröffentlicht: (2024)
Challenges and Opportunities to Enable Large-Scale Computing via Heterogeneous Chiplets
von: Yang, Zhuoping, et al.
Veröffentlicht: (2023)
von: Yang, Zhuoping, et al.
Veröffentlicht: (2023)
ChipletPart: Cost-Aware Partitioning for 2.5D Systems
von: Graening, Alexander, et al.
Veröffentlicht: (2025)
von: Graening, Alexander, et al.
Veröffentlicht: (2025)
Tiny Chiplets Enabled by Packaging Scaling: Opportunities in ESD Protection and Signal Integrity
von: Haque, Emad, et al.
Veröffentlicht: (2025)
von: Haque, Emad, et al.
Veröffentlicht: (2025)
LEXI: Lossless Exponent Coding for Efficient Inter-Chiplet Communication in Hybrid LLMs
von: Sun, Miao, et al.
Veröffentlicht: (2026)
von: Sun, Miao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Patterns behind Chaos: Forecasting Data Movement for Efficient Large-Scale MoE LLM Inference
von: Yu, Zhongkai, et al.
Veröffentlicht: (2025) -
ChipBench: A Next-Step Benchmark for Evaluating LLM Performance in AI-Aided Chip Design
von: Yu, Zhongkai, et al.
Veröffentlicht: (2026) -
Chiplet Actuary: A Quantitative Cost Model and Multi-Chiplet Architecture Exploration
von: Feng, Yinxiao, et al.
Veröffentlicht: (2022) -
RapidChiplet: A Toolchain for Rapid Design Space Exploration of Chiplet Architectures
von: Iff, Patrick, et al.
Veröffentlicht: (2023) -
zhongkaiyu/moe_exp_placement: Case Study2 for ISCA 2026 AE
von: Yu, Zhongkai, et al.
Veröffentlicht: (2026)