MAP-UOT: A Memory-Efficient Approach to Unbalanced Optimal Transport Implementation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Sun, Chengyu, Hu, Jinyu, Jiang, Hong |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
What Cannot Be Implemented on Weak Memory?
par: Castañeda, Armando, et autres
Publié: (2024)
par: Castañeda, Armando, et autres
Publié: (2024)
Some New Approaches to MPI Implementations
par: Xiong, Yuqing
Publié: (2024)
par: Xiong, Yuqing
Publié: (2024)
Amortized Asynchronous Byzantine Reliable Broadcast with Optimal Resilience
par: Hu, Michael Yiqing, et autres
Publié: (2026)
par: Hu, Michael Yiqing, et autres
Publié: (2026)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
par: Hu, Cunchen, et autres
Publié: (2024)
par: Hu, Cunchen, et autres
Publié: (2024)
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management
par: Qianli, Liu, et autres
Publié: (2025)
par: Qianli, Liu, et autres
Publié: (2025)
An Asynchronous Many-Task Algorithm for Unstructured $S_{N}$ Transport on Shared Memory Systems
par: Elwood, Alex, et autres
Publié: (2025)
par: Elwood, Alex, et autres
Publié: (2025)
Cicada: A Pipeline-Efficient Approach to Serverless Inference with Decoupled Management
par: Wu, Z., et autres
Publié: (2025)
par: Wu, Z., et autres
Publié: (2025)
eLLM: Elastic Memory Management Framework for Efficient LLM Serving
par: Xu, Jiale, et autres
Publié: (2025)
par: Xu, Jiale, et autres
Publié: (2025)
Do MPI Derived Datatypes Actually Help? A Single-Node Cross-Implementation Study on Shared-Memory Communication
par: Adefemi, Temitayo
Publié: (2025)
par: Adefemi, Temitayo
Publié: (2025)
Sizey: Memory-Efficient Execution of Scientific Workflow Tasks
par: Bader, Jonathan, et autres
Publié: (2024)
par: Bader, Jonathan, et autres
Publié: (2024)
AME: An Efficient Heterogeneous Agentic Memory Engine for Smartphones
par: Zhao, Xinkui, et autres
Publié: (2025)
par: Zhao, Xinkui, et autres
Publié: (2025)
GMLake: Efficient and Transparent GPU Memory Defragmentation for Large-scale DNN Training with Virtual Memory Stitching
par: Guo, Cong, et autres
Publié: (2024)
par: Guo, Cong, et autres
Publié: (2024)
Efficient Parallel Implementation of the Pilot Assignment Problem in Massive MIMO Systems
par: Alqudah, Eman, et autres
Publié: (2025)
par: Alqudah, Eman, et autres
Publié: (2025)
DAK: Direct-Access-Enabled GPU Memory Offloading with Optimal Efficiency for LLM Inference
par: Lin, Shouxu, et autres
Publié: (2026)
par: Lin, Shouxu, et autres
Publié: (2026)
Efficient and Portable Support for Overdecomposition on Distributed Memory GPGPU Platforms
par: Bhosale, Aditya, et autres
Publié: (2026)
par: Bhosale, Aditya, et autres
Publié: (2026)
Flash-KMeans: Fast and Memory-Efficient Exact K-Means
par: Yang, Shuo, et autres
Publié: (2026)
par: Yang, Shuo, et autres
Publié: (2026)
OTAS: An Elastic Transformer Serving System via Token Adaptation
par: Chen, Jinyu, et autres
Publié: (2024)
par: Chen, Jinyu, et autres
Publié: (2024)
SiDP: Memory-Efficient Data Parallelism for Offline LLM Inference
par: Zhao, Alan, et autres
Publié: (2026)
par: Zhao, Alan, et autres
Publié: (2026)
Efficient GPU Implementation of Particle Interactions with Cutoff Radius and Few Particles per Cell
par: Algis, David, et autres
Publié: (2024)
par: Algis, David, et autres
Publié: (2024)
Prioritized-MVBA: A New Approach to Design an Optimal Asynchronous Byzantine Agreement Protocol
par: Sony, Nasit S, et autres
Publié: (2024)
par: Sony, Nasit S, et autres
Publié: (2024)
Silent Self-Stabilizing Ranking: Time Optimal and Space Efficient
par: Berenbrink, Petra, et autres
Publié: (2025)
par: Berenbrink, Petra, et autres
Publié: (2025)
Heterogeneity-Aware Memory Efficient Federated Learning via Progressive Layer Freezing
par: Yebo, Wu, et autres
Publié: (2024)
par: Yebo, Wu, et autres
Publié: (2024)
Picasso: Memory-Efficient Graph Coloring Using Palettes With Applications in Quantum Computing
par: Ferdous, S M, et autres
Publié: (2024)
par: Ferdous, S M, et autres
Publié: (2024)
DiFache: Efficient and Scalable Caching on Disaggregated Memory using Decentralized Coherence
par: Zhang, Hanze, et autres
Publié: (2025)
par: Zhang, Hanze, et autres
Publié: (2025)
Efficient Wait-Free Linearizable Implementations of Approximate Bounded Counters Using Read-Write Registers
par: Johnen, Colette, et autres
Publié: (2024)
par: Johnen, Colette, et autres
Publié: (2024)
UELLM: A Unified and Efficient Approach for LLM Inference Serving
par: He, Yiyuan, et autres
Publié: (2024)
par: He, Yiyuan, et autres
Publié: (2024)
Self-Evolving Distributed Memory Architecture for Scalable AI Systems
par: Li, Zixuan, et autres
Publié: (2026)
par: Li, Zixuan, et autres
Publié: (2026)
A Tale of Two Paths: Toward a Hybrid Data Plane for Efficient Far-Memory Applications
par: Chen, Lei, et autres
Publié: (2024)
par: Chen, Lei, et autres
Publié: (2024)
Serving Chain-structured Jobs with Large Memory Footprints with Application to Large Foundation Model Serving
par: Sun, Tingyang, et autres
Publié: (2026)
par: Sun, Tingyang, et autres
Publié: (2026)
Memory-Efficient Split Federated Learning for LLM Fine-Tuning on Heterogeneous Mobile Devices
par: Chen, Xiaopei, et autres
Publié: (2025)
par: Chen, Xiaopei, et autres
Publié: (2025)
Memory-Efficient Federated Fine-Tuning of Large Language Models via Layer Pruning
par: Wu, Yebo, et autres
Publié: (2025)
par: Wu, Yebo, et autres
Publié: (2025)
MemFine: Memory-Aware Fine-Grained Scheduling for MoE Training
par: Zhao, Lu, et autres
Publié: (2025)
par: Zhao, Lu, et autres
Publié: (2025)
FlexKV: Flexible Index Offloading for Memory-Disaggregated Key-Value Store
par: Hu, Zhisheng, et autres
Publié: (2025)
par: Hu, Zhisheng, et autres
Publié: (2025)
Efficient Training of Large Language Models on Distributed Infrastructures: A Survey
par: Duan, Jiangfei, et autres
Publié: (2024)
par: Duan, Jiangfei, et autres
Publié: (2024)
SW-TNC : Reaching the Most Complex Random Quantum Circuit via Tensor Network Contraction
par: Chen, Yaojian, et autres
Publié: (2025)
par: Chen, Yaojian, et autres
Publié: (2025)
Efficient CPU-GPU Collaborative Inference for MoE-based LLMs on Memory-Limited Systems
par: Huang, En-Ming, et autres
Publié: (2025)
par: Huang, En-Ming, et autres
Publié: (2025)
Memory Efficient and Staleness Free Pipeline Parallel DNN Training Framework with Improved Convergence Speed
par: Dutta, Ankita, et autres
Publié: (2025)
par: Dutta, Ankita, et autres
Publié: (2025)
NestedFP: High-Performance, Memory-Efficient Dual-Precision Floating Point Support for LLMs
par: Lee, Haeun, et autres
Publié: (2025)
par: Lee, Haeun, et autres
Publié: (2025)
FleetOpt: Analytical Fleet Provisioning for LLM Inference with Compress-and-Route as Implementation Mechanism
par: Chen, Huamin, et autres
Publié: (2026)
par: Chen, Huamin, et autres
Publié: (2026)
An Efficient Approach for Energy Conservation in Cloud Computing Environment
par: Pande, Sohan Kumar, et autres
Publié: (2025)
par: Pande, Sohan Kumar, et autres
Publié: (2025)
Documents similaires
-
What Cannot Be Implemented on Weak Memory?
par: Castañeda, Armando, et autres
Publié: (2024) -
Some New Approaches to MPI Implementations
par: Xiong, Yuqing
Publié: (2024) -
Amortized Asynchronous Byzantine Reliable Broadcast with Optimal Resilience
par: Hu, Michael Yiqing, et autres
Publié: (2026) -
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
par: Hu, Cunchen, et autres
Publié: (2024) -
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management
par: Qianli, Liu, et autres
Publié: (2025)