Asymmetric Virtual Memory Paging for Hybrid Mamba-Transformer Inference
Fuente:
arXiv
Saved in:
| Main Author: | Nguyen, An Xuan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference
by: Ganjihal, Sanjeev Rao
Published: (2026)
by: Ganjihal, Sanjeev Rao
Published: (2026)
DDS: DPU-optimized Disaggregated Storage [Extended Report]
by: Zhang, Qizhen, et al.
Published: (2024)
by: Zhang, Qizhen, et al.
Published: (2024)
Sky$^ε$-Tree: Embracing the Batch Updates of B$^ε$-trees through Access Port Parallelism on Skyrmion Racetrack Memory
by: Tsai, Yu-Shiang, et al.
Published: (2024)
by: Tsai, Yu-Shiang, et al.
Published: (2024)
Reexamining Paradigms of End-to-End Data Movement
by: Fang, Chin, et al.
Published: (2025)
by: Fang, Chin, et al.
Published: (2025)
SLO-Guard: Crash-Aware, Budget-Consistent Autotuning for SLO-Constrained LLM Serving
by: Lysenstøen, Christian
Published: (2026)
by: Lysenstøen, Christian
Published: (2026)
Flexible Swapping for the Cloud
by: Pandurov, Milan, et al.
Published: (2024)
by: Pandurov, Milan, et al.
Published: (2024)
Vectorized Adaptive Histograms for Sparse Oblique Forests
by: Lubonja, Ariel, et al.
Published: (2026)
by: Lubonja, Ariel, et al.
Published: (2026)
Reproduction Research of FSA-Benchmark
by: Ludolf, Joshua, et al.
Published: (2024)
by: Ludolf, Joshua, et al.
Published: (2024)
Intelligent Cloud Orchestration: A Hybrid Predictive and Heuristic Framework for Cost Optimization
by: Nagoriya, Heet, et al.
Published: (2026)
by: Nagoriya, Heet, et al.
Published: (2026)
Cost-Aware Logging: Measuring the Financial Impact of Excessive Log Retention in Small-Scale Cloud Deployments
by: Putra, Jody Almaida
Published: (2026)
by: Putra, Jody Almaida
Published: (2026)
TokenCake: A KV-Cache-centric Serving Framework for LLM-based Multi-Agent Applications
by: Bian, Zhuohang, et al.
Published: (2025)
by: Bian, Zhuohang, et al.
Published: (2025)
Comparative Analysis of Large Language Model Inference Serving Systems: A Performance Study of vLLM and HuggingFace TGI
by: Kolluru, Saicharan
Published: (2025)
by: Kolluru, Saicharan
Published: (2025)
DF* PageRank: Improved Incrementally Expanding Approaches for Updating PageRank on Dynamic Graphs
by: Sahu, Subhajit
Published: (2024)
by: Sahu, Subhajit
Published: (2024)
Polynomial Histograms for Memory-Efficient Representation of Long-tailed System Distributions
by: Stokely, Murray, et al.
Published: (2026)
by: Stokely, Murray, et al.
Published: (2026)
Federated Fine-Tuning of LLMs on the Very Edge: The Good, the Bad, the Ugly
by: Woisetschläger, Herbert, et al.
Published: (2023)
by: Woisetschläger, Herbert, et al.
Published: (2023)
DNA sequence alignment: An assignment for OpenMP, MPI, and CUDA/OpenCL
by: Gonzalez-Escribano, Arturo, et al.
Published: (2024)
by: Gonzalez-Escribano, Arturo, et al.
Published: (2024)
Send: Objects, History, and Transactions in a Single-Verb Kernel
by: Goes, Christopher
Published: (2026)
by: Goes, Christopher
Published: (2026)
cfdSCOPE: A Fluid-Dynamics Proxy App for Teaching Performance Engineering
by: Arzt, Peter, et al.
Published: (2025)
by: Arzt, Peter, et al.
Published: (2025)
Lock-Free Computation of PageRank in Dynamic Graphs
by: Sahu, Subhajit
Published: (2024)
by: Sahu, Subhajit
Published: (2024)
Unlocking Python's Cores: Hardware Usage and Energy Implications of Removing the GIL
by: Salazar, José Daniel Montoya
Published: (2026)
by: Salazar, José Daniel Montoya
Published: (2026)
A Methodology to Assess Power Modeling in Energy-Aware Federated Learning on Heterogeneous Mobile Devices
by: Jallouli, Chaimae, et al.
Published: (2026)
by: Jallouli, Chaimae, et al.
Published: (2026)
Scheduling the Unschedulable: Taming Black-Box LLM Inference at Scale
by: Yuan, Renzhong, et al.
Published: (2026)
by: Yuan, Renzhong, et al.
Published: (2026)
An Incrementally Expanding Approach for Updating PageRank on Dynamic Graphs
by: Sahu, Subhajit
Published: (2024)
by: Sahu, Subhajit
Published: (2024)
TStore: Rethinking AI Model Hub with Tensor-Centric Compression
by: Lan, Tingfeng, et al.
Published: (2026)
by: Lan, Tingfeng, et al.
Published: (2026)
AutoChunk: Automated Activation Chunk for Memory-Efficient Long Sequence Inference
by: Zhao, Xuanlei, et al.
Published: (2024)
by: Zhao, Xuanlei, et al.
Published: (2024)
Token Arena: A Continuous Benchmark Unifying Energy and Cognition in AI Inference
by: Gao, Yuxuan, et al.
Published: (2026)
by: Gao, Yuxuan, et al.
Published: (2026)
CoFormer: Collaborating with Heterogeneous Edge Devices for Scalable Transformer Inference
by: Xu, Guanyu, et al.
Published: (2025)
by: Xu, Guanyu, et al.
Published: (2025)
PIM-STM: Software Transactional Memory for Processing-In-Memory Systems
by: Lopes, André, et al.
Published: (2024)
by: Lopes, André, et al.
Published: (2024)
SGDRC: Software-Defined Dynamic Resource Control for Concurrent DNN Inference on NVIDIA GPUs
by: Zhang, Yongkang, et al.
Published: (2024)
by: Zhang, Yongkang, et al.
Published: (2024)
Stream parallel skeleton optimization
by: Aldinucci, Marco, et al.
Published: (2024)
by: Aldinucci, Marco, et al.
Published: (2024)
StreamFlow: cross-breeding cloud with HPC
by: Colonnelli, Iacopo, et al.
Published: (2020)
by: Colonnelli, Iacopo, et al.
Published: (2020)
I/O in Machine Learning Applications on HPC Systems: A 360-degree Survey
by: Lewis, Noah, et al.
Published: (2024)
by: Lewis, Noah, et al.
Published: (2024)
KVDirect: Distributed Disaggregated LLM Inference
by: Chen, Shiyang, et al.
Published: (2024)
by: Chen, Shiyang, et al.
Published: (2024)
Multi-DNN Inference of Sparse Models on Edge SoCs
by: Luo, Jiawei, et al.
Published: (2026)
by: Luo, Jiawei, et al.
Published: (2026)
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference
by: Hamadanian, Pouya, et al.
Published: (2025)
by: Hamadanian, Pouya, et al.
Published: (2025)
GVE-Louvain: Fast Louvain Algorithm for Community Detection in Shared Memory Setting
by: Sahu, Subhajit
Published: (2023)
by: Sahu, Subhajit
Published: (2023)
GVE-Leiden: Fast Leiden Algorithm for Community Detection in Shared Memory Setting
by: Sahu, Subhajit
Published: (2023)
by: Sahu, Subhajit
Published: (2023)
Libra: Unleashing GPU Heterogeneity for High-Performance Sparse Matrix Multiplication
by: Shi, Jinliang, et al.
Published: (2025)
by: Shi, Jinliang, et al.
Published: (2025)
IPA: Inference Pipeline Adaptation to Achieve High Accuracy and Cost-Efficiency
by: Ghafouri, Saeid, et al.
Published: (2023)
by: Ghafouri, Saeid, et al.
Published: (2023)
The Energy-Throughput Trade-off in Lossless-Compressed Source Code Storage
by: Ferragina, Paolo, et al.
Published: (2026)
by: Ferragina, Paolo, et al.
Published: (2026)
Similar Items
-
Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference
by: Ganjihal, Sanjeev Rao
Published: (2026) -
DDS: DPU-optimized Disaggregated Storage [Extended Report]
by: Zhang, Qizhen, et al.
Published: (2024) -
Sky$^ε$-Tree: Embracing the Batch Updates of B$^ε$-trees through Access Port Parallelism on Skyrmion Racetrack Memory
by: Tsai, Yu-Shiang, et al.
Published: (2024) -
Reexamining Paradigms of End-to-End Data Movement
by: Fang, Chin, et al.
Published: (2025) -
SLO-Guard: Crash-Aware, Budget-Consistent Autotuning for SLO-Constrained LLM Serving
by: Lysenstøen, Christian
Published: (2026)