Move the Query, Not the Cache: Characterizing Cross-Instance Latent Attention Redistribution Across GPU Fabrics
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Bole, Eitzinger, Jan, Köstler, Harald, Wellein, Gerhard |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Irminsul: MLA-Native Position-Independent Caching for Agentic LLM Serving
by: Ma, Bole, et al.
Published: (2026)
by: Ma, Bole, et al.
Published: (2026)
Diagnosing Overhead in Dispatch Operations: Cross-architecture Observatory
by: Ma, Bole, et al.
Published: (2026)
by: Ma, Bole, et al.
Published: (2026)
The Illusion of Power Capping in LLM Decode: A Phase-Aware Energy Characterisation Across Attention Architectures
by: Ma, Bole, et al.
Published: (2026)
by: Ma, Bole, et al.
Published: (2026)
Leyline: KV Cache Directives for Agentic Inference
by: Ma, Bole, et al.
Published: (2026)
by: Ma, Bole, et al.
Published: (2026)
GORGO: Maximizing KV-Cache Reuse While Minimizing Network Latency in Cross-Region LLM Load Balancing
by: Toniolo, Alessio Ricci, et al.
Published: (2026)
by: Toniolo, Alessio Ricci, et al.
Published: (2026)
TINC: Trusted Intelligent NetChain
by: Xia, Qi, et al.
Published: (2025)
by: Xia, Qi, et al.
Published: (2025)
FAST: An Efficient Scheduler for All-to-All GPU Communication
by: Lei, Yiran, et al.
Published: (2025)
by: Lei, Yiran, et al.
Published: (2025)
Diving into 3D Parallelism with Heterogeneous Spot Instance GPUs: Design and Implications
by: Wang, Yuxiao, et al.
Published: (2025)
by: Wang, Yuxiao, et al.
Published: (2025)
LOAM: Low-latency Communication, Caching, and Computation Placement in Data-Intensive Computing Networks
by: Zhang, Jinkun, et al.
Published: (2024)
by: Zhang, Jinkun, et al.
Published: (2024)
Dynamic Hierarchical Birkhoff-von Neumann Decomposition for All-to-All GPU Communication
by: Wu, Yen-Chieh, et al.
Published: (2026)
by: Wu, Yen-Chieh, et al.
Published: (2026)
An Online Fragmentation-Aware GPU Scheduler for Multi-Tenant MIG-based Clouds
by: Zambianco, Marco, et al.
Published: (2025)
by: Zambianco, Marco, et al.
Published: (2025)
A Hybrid Approach to Monitor Context Parameters for Optimising Caching for Context-Aware IoT Applications
by: Manchanda, Ashish, et al.
Published: (2024)
by: Manchanda, Ashish, et al.
Published: (2024)
Temporal-Aware GPU Resource Allocation for Distributed LLM Inference via Reinforcement Learning
by: Du, Chengze, et al.
Published: (2025)
by: Du, Chengze, et al.
Published: (2025)
From Skew to Symmetry: Node-Interconnect Multi-Path Balancing with Execution-time Planning for Modern GPU Clusters
by: Yao, Jinghan, et al.
Published: (2026)
by: Yao, Jinghan, et al.
Published: (2026)
Carbon-Aware Temporal Data Transfer Scheduling Across Cloud Datacenters
by: Rodrigues, Elvis, et al.
Published: (2025)
by: Rodrigues, Elvis, et al.
Published: (2025)
Whack-a-Mole: Deterministic Packet Spraying Across Multiple Network Paths
by: Luby, Michael, et al.
Published: (2025)
by: Luby, Michael, et al.
Published: (2025)
Palladium: A DPU-enabled Multi-Tenant Serverless Cloud over Zero-copy Multi-node RDMA Fabrics
by: Qi, Shixiong, et al.
Published: (2025)
by: Qi, Shixiong, et al.
Published: (2025)
SlimCaching: Edge Caching of Mixture-of-Experts for Distributed Inference
by: Chen, Qian, et al.
Published: (2025)
by: Chen, Qian, et al.
Published: (2025)
Local Rendezvous Hashing: Bounded Loads and Minimal Churn via Cache-Local Candidates
by: Guan, Yongjie
Published: (2025)
by: Guan, Yongjie
Published: (2025)
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services
by: Tang, Lingfeng, et al.
Published: (2025)
by: Tang, Lingfeng, et al.
Published: (2025)
BlockSDN: Towards a High-Performance Blockchain via Software-Defined Cross Networking optimization
by: Jia, Wenyang, et al.
Published: (2025)
by: Jia, Wenyang, et al.
Published: (2025)
Decentralized Network Topology Design for Task Offloading in Mobile Edge Computing
by: Ma, Ke, et al.
Published: (2024)
by: Ma, Ke, et al.
Published: (2024)
A Multi-Layered Distributed Computing Framework for Enhanced Edge Computing
by: Ma, Ke, et al.
Published: (2024)
by: Ma, Ke, et al.
Published: (2024)
KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving
by: Liu, Zedong, et al.
Published: (2026)
by: Liu, Zedong, et al.
Published: (2026)
Revisiting Cache Freshness for Emerging Real-Time Applications
by: Mao, Ziming, et al.
Published: (2024)
by: Mao, Ziming, et al.
Published: (2024)
Placing Timely Refreshing Services at the Network Edge
by: Li, Xishuo, et al.
Published: (2024)
by: Li, Xishuo, et al.
Published: (2024)
Towards Timely Video Analytics Services at the Network Edge
by: Li, Xishuo, et al.
Published: (2024)
by: Li, Xishuo, et al.
Published: (2024)
A Learning-Based Caching Mechanism for Edge Content Delivery
by: Torabi, Hoda, et al.
Published: (2024)
by: Torabi, Hoda, et al.
Published: (2024)
Sensing and Storing Less: A MARL-based Solution for Energy Saving in Edge Internet of Things
by: Yuan, Zongyang, et al.
Published: (2025)
by: Yuan, Zongyang, et al.
Published: (2025)
ONCache: A Cache-Based Low-Overhead Container Overlay Network
by: Lin, Shengkai, et al.
Published: (2023)
by: Lin, Shengkai, et al.
Published: (2023)
A Survey on Privacy-Preserving Caching at Network Edge: Classification, Solutions, and Challenges
by: Zhang, Xianzhi, et al.
Published: (2024)
by: Zhang, Xianzhi, et al.
Published: (2024)
ScaleAcross Explorer: Exploring Communication Optimization for Scale-Across AI Model Training
by: Li, Minghao, et al.
Published: (2026)
by: Li, Minghao, et al.
Published: (2026)
cMPI: Using CXL Memory Sharing for MPI One-Sided and Two-Sided Inter-Node Communications
by: Wang, Xi, et al.
Published: (2025)
by: Wang, Xi, et al.
Published: (2025)
RailX: A Flexible, Scalable, and Low-Cost Network Architecture for Hyper-Scale LLM Training Systems
by: Feng, Yinxiao, et al.
Published: (2025)
by: Feng, Yinxiao, et al.
Published: (2025)
JANUS: Resilient and Adaptive Data Transmission for Enabling Timely and Efficient Cross-Facility Scientific Workflows
by: Esaulov, Vladislav, et al.
Published: (2025)
by: Esaulov, Vladislav, et al.
Published: (2025)
ForestColl: Throughput-Optimal Collective Communications on Heterogeneous Network Fabrics
by: Zhao, Liangyu, et al.
Published: (2024)
by: Zhao, Liangyu, et al.
Published: (2024)
MULTI-SCOUT: Multistatic Integrated Sensing and Communications in 5G and Beyond for Moving Target Detection, Positioning, and Tracking
by: Sagduyu, Yalin E., et al.
Published: (2025)
by: Sagduyu, Yalin E., et al.
Published: (2025)
Trivance: Latency-Optimal AllReduce by Shortcutting Multiport Networks
by: Juerss, Anton, et al.
Published: (2026)
by: Juerss, Anton, et al.
Published: (2026)
A Uniqueness Theorem for Distributed Computation under Physical Constraint
by: Ren, Zhiyuan, et al.
Published: (2025)
by: Ren, Zhiyuan, et al.
Published: (2025)
Multi-stage Flow Scheduling for LLM Serving
by: Sun, Yijun, et al.
Published: (2026)
by: Sun, Yijun, et al.
Published: (2026)
Similar Items
-
Irminsul: MLA-Native Position-Independent Caching for Agentic LLM Serving
by: Ma, Bole, et al.
Published: (2026) -
Diagnosing Overhead in Dispatch Operations: Cross-architecture Observatory
by: Ma, Bole, et al.
Published: (2026) -
The Illusion of Power Capping in LLM Decode: A Phase-Aware Energy Characterisation Across Attention Architectures
by: Ma, Bole, et al.
Published: (2026) -
Leyline: KV Cache Directives for Agentic Inference
by: Ma, Bole, et al.
Published: (2026) -
GORGO: Maximizing KV-Cache Reuse While Minimizing Network Latency in Cross-Region LLM Load Balancing
by: Toniolo, Alessio Ricci, et al.
Published: (2026)