CUCo: An Agentic Framework for Compute and Communication Co-design
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Bodun, Varshan V, Yoga Sri, Agarwal, Saurabh, Akella, Aditya |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Patchwork: A Unified Framework for RAG Serving
by: Hu, Bodun, et al.
Published: (2025)
by: Hu, Bodun, et al.
Published: (2025)
Software-Defined Agentic Serving
by: Agarwal, Saurabh, et al.
Published: (2026)
by: Agarwal, Saurabh, et al.
Published: (2026)
Warp-Cortex: An Asynchronous, Memory-Efficient Architecture for Million-Agent Cognitive Scaling on Consumer Hardware
by: Williams, Jorge L. Ruiz
Published: (2026)
by: Williams, Jorge L. Ruiz
Published: (2026)
Next-generation Probabilistic Computing Hardware with 3D MOSAICs, Illusion Scale-up, and Co-design
by: Srimani, Tathagata, et al.
Published: (2024)
by: Srimani, Tathagata, et al.
Published: (2024)
Nalar: An agent serving framework
by: Laju, Marco, et al.
Published: (2026)
by: Laju, Marco, et al.
Published: (2026)
ASIC-Agent: An Autonomous Multi-Agent System for ASIC Design with Benchmark Evaluation
by: Allam, Ahmed, et al.
Published: (2025)
by: Allam, Ahmed, et al.
Published: (2025)
Transforming the Hybrid Cloud for Emerging AI Workloads
by: Chen, Deming, et al.
Published: (2024)
by: Chen, Deming, et al.
Published: (2024)
CELLO: Co-designing Schedule and Hybrid Implicit/Explicit Buffer for Complex Tensor Reuse
by: Garg, Raveesh, et al.
Published: (2023)
by: Garg, Raveesh, et al.
Published: (2023)
Optimizing Distributed ML Communication with Fused Computation-Collective Operations
by: Punniyamurthy, Kishore, et al.
Published: (2023)
by: Punniyamurthy, Kishore, et al.
Published: (2023)
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines
by: Agrawal, Anirudha, et al.
Published: (2024)
by: Agrawal, Anirudha, et al.
Published: (2024)
Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap
by: Pal, Shagnik, et al.
Published: (2025)
by: Pal, Shagnik, et al.
Published: (2025)
Observation, Not Prediction: Conversation-Level Disaggregated Scheduling for Agentic Serving
by: Ding, Jianru, et al.
Published: (2026)
by: Ding, Jianru, et al.
Published: (2026)
PID-Comm: A Fast and Flexible Collective Communication Framework for Commodity Processing-in-DIMM Devices
by: Noh, Si Ung, et al.
Published: (2024)
by: Noh, Si Ung, et al.
Published: (2024)
Efficient, VRAM-Constrained xLM Inference on Clients
by: Ukarande, Aditya, et al.
Published: (2026)
by: Ukarande, Aditya, et al.
Published: (2026)
Memory-Centric Computing: Solving Computing's Memory Problem
by: Mutlu, Onur, et al.
Published: (2025)
by: Mutlu, Onur, et al.
Published: (2025)
FastGL: A GPU-Efficient Framework for Accelerating Sampling-Based GNN Training at Large Scale
by: Zhu, Zeyu, et al.
Published: (2024)
by: Zhu, Zeyu, et al.
Published: (2024)
Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems
by: Zhang, Chen, et al.
Published: (2026)
by: Zhang, Chen, et al.
Published: (2026)
Context-aware Simopt-Power: Using structural data with simulation metadata to optimise FPGA designs
by: Wadhwa, Eashan, et al.
Published: (2026)
by: Wadhwa, Eashan, et al.
Published: (2026)
WWW: What, When, Where to Compute-in-Memory
by: Sharma, Tanvi, et al.
Published: (2023)
by: Sharma, Tanvi, et al.
Published: (2023)
Revisiting Computational Storage for Data Integrity and Security
by: Shi, Chao, et al.
Published: (2025)
by: Shi, Chao, et al.
Published: (2025)
FedUHD: Unsupervised Federated Learning using Hyperdimensional Computing
by: Lee, You Hak, et al.
Published: (2025)
by: Lee, You Hak, et al.
Published: (2025)
Optimizing Task Scheduling in Fog Computing with Deadline Awareness
by: Sirjani, Mohammad Sadegh, et al.
Published: (2025)
by: Sirjani, Mohammad Sadegh, et al.
Published: (2025)
Memory-Centric Computing: Recent Advances in Processing-in-DRAM
by: Mutlu, Onur, et al.
Published: (2024)
by: Mutlu, Onur, et al.
Published: (2024)
Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUs
by: Zhang, Qijun, et al.
Published: (2026)
by: Zhang, Qijun, et al.
Published: (2026)
Knowledge-Guided Attention-Inspired Learning for Task Offloading in Vehicle Edge Computing
by: Ma, Ke, et al.
Published: (2025)
by: Ma, Ke, et al.
Published: (2025)
BlockAMC: Scalable In-Memory Analog Matrix Computing for Solving Linear Systems
by: Pan, Lunshuai, et al.
Published: (2024)
by: Pan, Lunshuai, et al.
Published: (2024)
Deep Learning and Machine Learning with GPGPU and CUDA: Unlocking the Power of Parallel Computing
by: Li, Ming, et al.
Published: (2024)
by: Li, Ming, et al.
Published: (2024)
How Fast Can Graph Computations Go on Fine-grained Parallel Architectures
by: Wang, Yuqing, et al.
Published: (2025)
by: Wang, Yuqing, et al.
Published: (2025)
UniFormer: Unified and Efficient Transformer for Reasoning Across General and Custom Computing
by: Ran, Zhuoheng, et al.
Published: (2025)
by: Ran, Zhuoheng, et al.
Published: (2025)
DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication
by: Pati, Suchita, et al.
Published: (2025)
by: Pati, Suchita, et al.
Published: (2025)
RapidOMS: FPGA-based Open Modification Spectral Library Searching with HD Computing
by: Pinge, Sumukh, et al.
Published: (2024)
by: Pinge, Sumukh, et al.
Published: (2024)
FpgaHub: Fpga-centric Hyper-heterogeneous Computing Platform for Big Data Analytics
by: Wang, Zeke, et al.
Published: (2025)
by: Wang, Zeke, et al.
Published: (2025)
TT-Edge: A Hardware-Software Co-Design for Energy-Efficient Tensor-Train Decomposition on Edge AI
by: Kwak, Hyunseok, et al.
Published: (2025)
by: Kwak, Hyunseok, et al.
Published: (2025)
Optimizing Communication for Latency Sensitive HPC Applications on up to 48 FPGAs Using ACCL
by: Meyer, Marius, et al.
Published: (2024)
by: Meyer, Marius, et al.
Published: (2024)
CCSS: Hardware-Accelerated RTL Simulation with Fast Combinational Logic Computing and Sequential Logic Synchronization
by: Feng, Weigang, et al.
Published: (2025)
by: Feng, Weigang, et al.
Published: (2025)
INR-Arch: A Dataflow Architecture and Compiler for Arbitrary-Order Gradient Computations in Implicit Neural Representation Processing
by: Abi-Karam, Stefan, et al.
Published: (2023)
by: Abi-Karam, Stefan, et al.
Published: (2023)
NetSmith: An Optimization Framework for Machine-Discovered Network Topologies
by: Green, Conor, et al.
Published: (2024)
by: Green, Conor, et al.
Published: (2024)
PULSAR: Simultaneous Many-Row Activation for Reliable and High-Performance Computing in Off-the-Shelf DRAM Chips
by: Yuksel, Ismail Emir, et al.
Published: (2023)
by: Yuksel, Ismail Emir, et al.
Published: (2023)
Conduit: Programmer-Transparent Near-Data Processing Using Multiple Compute-Capable Resources in Solid State Drives
by: Nadig, Rakesh, et al.
Published: (2026)
by: Nadig, Rakesh, et al.
Published: (2026)
COMET: A Framework for Modeling Compound Operation Dataflows with Explicit Collectives
by: Negi, Shubham, et al.
Published: (2025)
by: Negi, Shubham, et al.
Published: (2025)
Similar Items
-
Patchwork: A Unified Framework for RAG Serving
by: Hu, Bodun, et al.
Published: (2025) -
Software-Defined Agentic Serving
by: Agarwal, Saurabh, et al.
Published: (2026) -
Warp-Cortex: An Asynchronous, Memory-Efficient Architecture for Million-Agent Cognitive Scaling on Consumer Hardware
by: Williams, Jorge L. Ruiz
Published: (2026) -
Next-generation Probabilistic Computing Hardware with 3D MOSAICs, Illusion Scale-up, and Co-design
by: Srimani, Tathagata, et al.
Published: (2024) -
Nalar: An agent serving framework
by: Laju, Marco, et al.
Published: (2026)