TStore: Rethinking AI Model Hub with Tensor-Centric Compression
Fuente:
arXiv
Saved in:
| Main Authors: | Lan, Tingfeng, Wang, Zirui, Zheng, Yunjia, Su, Zhaoyuan, Yang, Juncheng, Cheng, Yue |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LatentBox: Storing AI-Generated Images at Scale via a Latent-First Design
by: Wang, Zirui, et al.
Published: (2026)
by: Wang, Zirui, et al.
Published: (2026)
ZipLLM: Efficient LLM Storage via Model-Aware Synergistic Data Deduplication and Compression
by: Wang, Zirui, et al.
Published: (2025)
by: Wang, Zirui, et al.
Published: (2025)
MorphServe: Efficient and Workload-Aware LLM Serving via Runtime Quantized Layer Swapping and KV Cache Resizing
by: Su, Zhaoyuan, et al.
Published: (2025)
by: Su, Zhaoyuan, et al.
Published: (2025)
A Full Compression Pipeline for Green Federated Learning in Communication-Constrained Environments
by: Colybes, Elouan, et al.
Published: (2026)
by: Colybes, Elouan, et al.
Published: (2026)
ADF-LoRA: Alternating Low-Rank Aggregation for Decentralized Federated Fine-Tuning
by: Wang, Xiaoyu, et al.
Published: (2025)
by: Wang, Xiaoyu, et al.
Published: (2025)
AMP4EC: Adaptive Model Partitioning Framework for Efficient Deep Learning Inference in Edge Computing Environments
by: Zhang, Guilin, et al.
Published: (2025)
by: Zhang, Guilin, et al.
Published: (2025)
Experimentally Evaluating the Resource Efficiency of Big Data Autoscaling
by: Will, Jonathan, et al.
Published: (2025)
by: Will, Jonathan, et al.
Published: (2025)
TokenCake: A KV-Cache-centric Serving Framework for LLM-based Multi-Agent Applications
by: Bian, Zhuohang, et al.
Published: (2025)
by: Bian, Zhuohang, et al.
Published: (2025)
Asynchronous Federated Reinforcement Learning with Policy Gradient Updates: Algorithm Design and Convergence Analysis
by: Lan, Guangchen, et al.
Published: (2024)
by: Lan, Guangchen, et al.
Published: (2024)
EWSJF: An Adaptive Scheduler with Hybrid Partitioning for Mixed-Workload LLM Inference
by: Sidik, Bronislav, et al.
Published: (2026)
by: Sidik, Bronislav, et al.
Published: (2026)
ACME: Adaptive Customization of Large Models via Distributed Systems
by: Dai, Ziming, et al.
Published: (2025)
by: Dai, Ziming, et al.
Published: (2025)
An Empirical Study of the Impact of Federated Learning on Machine Learning Model Accuracy
by: Yang, Haotian, et al.
Published: (2025)
by: Yang, Haotian, et al.
Published: (2025)
SPARK: Igniting Communication-Efficient Decentralized Learning via Stage-wise Projected NTK and Accelerated Regularization
by: Xia, Li
Published: (2025)
by: Xia, Li
Published: (2025)
Accelerating Causal Algorithms for Industrial-scale Data: A Distributed Computing Approach with Ray Framework
by: Verma, Vishal, et al.
Published: (2024)
by: Verma, Vishal, et al.
Published: (2024)
Parameter-Efficient and Personalized Federated Training of Generative Models at the Edge
by: Khan, Kabir, et al.
Published: (2025)
by: Khan, Kabir, et al.
Published: (2025)
TAGC: Optimizing Gradient Communication in Distributed Transformer Training
by: Polyakov, Igor, et al.
Published: (2025)
by: Polyakov, Igor, et al.
Published: (2025)
Closing Africa's Early Warning Gap: AI Weather Forecasting for Disaster Prevention
by: Ndlovu, Qness
Published: (2026)
by: Ndlovu, Qness
Published: (2026)
Static Batching of Irregular Workloads on GPUs: Framework and Application to Efficient MoE Model Inference
by: Li, Yinghan, et al.
Published: (2025)
by: Li, Yinghan, et al.
Published: (2025)
DDS: DPU-optimized Disaggregated Storage [Extended Report]
by: Zhang, Qizhen, et al.
Published: (2024)
by: Zhang, Qizhen, et al.
Published: (2024)
Vectorized Adaptive Histograms for Sparse Oblique Forests
by: Lubonja, Ariel, et al.
Published: (2026)
by: Lubonja, Ariel, et al.
Published: (2026)
Challenges of Heterogeneity in Big Data: A Comparative Study of Classification in Large-Scale Structured and Unstructured Domains
by: Eduardo, González Trigueros Jesús, et al.
Published: (2025)
by: Eduardo, González Trigueros Jesús, et al.
Published: (2025)
A Taxonomy and Resolution Strategy for Client-Level Disagreements in Federated Learning
by: Rosendal, Daan, et al.
Published: (2026)
by: Rosendal, Daan, et al.
Published: (2026)
XFED: Non-Collusive Model Poisoning Attack Against Byzantine-Robust Federated Classifiers
by: Mouri, Israt Jahan, et al.
Published: (2026)
by: Mouri, Israt Jahan, et al.
Published: (2026)
CarbonEdge: Carbon-Aware Deep Learning Inference Framework for Sustainable Edge Computing
by: Zhang, Guilin, et al.
Published: (2026)
by: Zhang, Guilin, et al.
Published: (2026)
MILLION: Mastering Long-Context LLM Inference Via Outlier-Immunized KV Product Quantization
by: Wang, Zongwu, et al.
Published: (2025)
by: Wang, Zongwu, et al.
Published: (2025)
SepsisAI Orchestrator: A Containerized and Scalable Platform for Deploying AI Models and Real-Time Monitoring in Early Sepsis Detection
by: Ospitia, Santiago, et al.
Published: (2026)
by: Ospitia, Santiago, et al.
Published: (2026)
Flash-SD-KDE: Accelerating SD-KDE with Tensor Cores
by: Epstein, Elliot L., et al.
Published: (2026)
by: Epstein, Elliot L., et al.
Published: (2026)
Stream-K++: Adaptive GPU GEMM Kernel Scheduling and Selection using Bloom Filters
by: Sadasivan, Harisankar, et al.
Published: (2024)
by: Sadasivan, Harisankar, et al.
Published: (2024)
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended
by: Kamath, Aditya K, et al.
Published: (2026)
by: Kamath, Aditya K, et al.
Published: (2026)
λScale: Enabling Fast Scaling for Serverless Large Language Model Inference
by: Yu, Minchen, et al.
Published: (2025)
by: Yu, Minchen, et al.
Published: (2025)
Optimizing edge AI models on HPC systems with the edge in the loop
by: Aach, Marcel, et al.
Published: (2025)
by: Aach, Marcel, et al.
Published: (2025)
Federated Few-Shot Learning on Neuromorphic Hardware: An Empirical Study Across Physical Edge Nodes
by: Motta, Steven, et al.
Published: (2026)
by: Motta, Steven, et al.
Published: (2026)
SLO-Guard: Crash-Aware, Budget-Consistent Autotuning for SLO-Constrained LLM Serving
by: Lysenstøen, Christian
Published: (2026)
by: Lysenstøen, Christian
Published: (2026)
A Comparative Analysis of Distributed Linear Solvers under Data Heterogeneity
by: Velasevic, Boris, et al.
Published: (2023)
by: Velasevic, Boris, et al.
Published: (2023)
Quantize Once, Train Fast: Allreduce-Compatible Compression with Provable Guarantees
by: Xin, Jihao, et al.
Published: (2023)
by: Xin, Jihao, et al.
Published: (2023)
SRFed: Mitigating Poisoning Attacks in Privacy-Preserving Federated Learning with Heterogeneous Data
by: Lu, Yiwen
Published: (2026)
by: Lu, Yiwen
Published: (2026)
Cross-Platform Fused MoE Dispatch in Triton: Portable Expert Routing Without CUDA
by: Mitra, Subhadip
Published: (2026)
by: Mitra, Subhadip
Published: (2026)
Data Readiness for Scientific AI at Scale
by: Brewer, Wesley, et al.
Published: (2025)
by: Brewer, Wesley, et al.
Published: (2025)
Flexible Swapping for the Cloud
by: Pandurov, Milan, et al.
Published: (2024)
by: Pandurov, Milan, et al.
Published: (2024)
Cognitive Infrastructure: A Unified DCIM Framework for AI Data Centers
by: Sunkara, Krishna Chaitanya
Published: (2026)
by: Sunkara, Krishna Chaitanya
Published: (2026)
Similar Items
-
LatentBox: Storing AI-Generated Images at Scale via a Latent-First Design
by: Wang, Zirui, et al.
Published: (2026) -
ZipLLM: Efficient LLM Storage via Model-Aware Synergistic Data Deduplication and Compression
by: Wang, Zirui, et al.
Published: (2025) -
MorphServe: Efficient and Workload-Aware LLM Serving via Runtime Quantized Layer Swapping and KV Cache Resizing
by: Su, Zhaoyuan, et al.
Published: (2025) -
A Full Compression Pipeline for Green Federated Learning in Communication-Constrained Environments
by: Colybes, Elouan, et al.
Published: (2026) -
ADF-LoRA: Alternating Low-Rank Aggregation for Decentralized Federated Fine-Tuning
by: Wang, Xiaoyu, et al.
Published: (2025)