NestedFP: High-Performance, Memory-Efficient Dual-Precision Floating Point Support for LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Haeun, Kwon, Omin, Park, Yeonhong, Lee, Jae W. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SGEMM-cube: Precision-Recovery FP32 GEMM Approximation on Ascend NPUs with FP16 Matrix Engines
by: Xue, Weicheng, et al.
Published: (2025)
by: Xue, Weicheng, et al.
Published: (2025)
Efficient CPU-GPU Collaborative Inference for MoE-based LLMs on Memory-Limited Systems
by: Huang, En-Ming, et al.
Published: (2025)
by: Huang, En-Ming, et al.
Published: (2025)
Double-Precision Matrix Multiplication Emulation via Ozaki-II Scheme with FP8 Quantization
by: Uchino, Yuki, et al.
Published: (2026)
by: Uchino, Yuki, et al.
Published: (2026)
Floating-Point Data Transformation for Lossless Compression
by: Jamalidinan, Samirasadat, et al.
Published: (2025)
by: Jamalidinan, Samirasadat, et al.
Published: (2025)
LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models
by: Park, Gunho, et al.
Published: (2022)
by: Park, Gunho, et al.
Published: (2022)
Efficient and Portable Support for Overdecomposition on Distributed Memory GPGPU Platforms
by: Bhosale, Aditya, et al.
Published: (2026)
by: Bhosale, Aditya, et al.
Published: (2026)
KubePACS: Kubernetes Cluster Using Performant, Highly Available, and Cost Efficient Spot Instances
by: Kim, Taeyoon, et al.
Published: (2026)
by: Kim, Taeyoon, et al.
Published: (2026)
PathWeaver: A High-Throughput Multi-GPU System for Graph-Based Approximate Nearest Neighbor Search
by: Kim, Sukjin, et al.
Published: (2025)
by: Kim, Sukjin, et al.
Published: (2025)
PIM-SHERPA: Software Method for On-device LLM Inference by Resolving PIM Memory Attribute and Layout Inconsistencies
by: Lee, Sunjung, et al.
Published: (2026)
by: Lee, Sunjung, et al.
Published: (2026)
HP2C-DT: High-Precision High-Performance Computer-enabled Digital Twin
by: Iraola, E., et al.
Published: (2025)
by: Iraola, E., et al.
Published: (2025)
PVU: Design and Implementation of a Posit Vector Arithmetic Unit (PVU) for Enhanced Floating-Point Computing in Edge and AI Applications
by: Wu, Xinyu, et al.
Published: (2025)
by: Wu, Xinyu, et al.
Published: (2025)
NAVIS: Concurrent Search and Update with Low Position-Seeking Overhead in On-SSD Graph-Based Vector Search
by: Song, Jaeyong, et al.
Published: (2026)
by: Song, Jaeyong, et al.
Published: (2026)
GriNNder: Breaking the Memory Capacity Wall in Full-Graph GNN Training with Storage Offloading
by: Song, Jaeyong, et al.
Published: (2026)
by: Song, Jaeyong, et al.
Published: (2026)
Performance Trade-offs of High Order Meshless Approximation on Distributed Memory Systems
by: Vehovar, Jon, et al.
Published: (2025)
by: Vehovar, Jon, et al.
Published: (2025)
Shared Virtual Memory: Its Design and Performance Implications for Diverse Applications
by: Cooper, Bennett, et al.
Published: (2024)
by: Cooper, Bennett, et al.
Published: (2024)
Characterizing Compute-Communication Overlap in GPU-Accelerated Distributed Deep Learning: Performance and Power Implications
by: Lee, Seonho, et al.
Published: (2025)
by: Lee, Seonho, et al.
Published: (2025)
FlexiWalker: Extensible GPU Framework for Efficient Dynamic Random Walks with Runtime Adaptation
by: Park, Seongyeon, et al.
Published: (2025)
by: Park, Seongyeon, et al.
Published: (2025)
The Design and Implementation of a High-Performance Log-Structured RAID System for ZNS SSDs
by: Li, Jinhong, et al.
Published: (2024)
by: Li, Jinhong, et al.
Published: (2024)
AGAThA: Fast and Efficient GPU Acceleration of Guided Sequence Alignment for Long Read Mapping
by: Park, Seongyeon, et al.
Published: (2024)
by: Park, Seongyeon, et al.
Published: (2024)
PULSE: Accelerating Distributed Pointer-Traversals on Disaggregated Memory (Extended Version)
by: Tang, Yupeng, et al.
Published: (2023)
by: Tang, Yupeng, et al.
Published: (2023)
F-BFQ: Flexible Block Floating-Point Quantization Accelerator for LLMs
by: Haris, Jude, et al.
Published: (2025)
by: Haris, Jude, et al.
Published: (2025)
Parallel Writing of Nested Data in Columnar Formats
by: Hahnfeld, Jonas, et al.
Published: (2024)
by: Hahnfeld, Jonas, et al.
Published: (2024)
EPIC: An Energy-Efficient, High-Performance GPGPU Computing Research Infrastructure
by: Själander, Magnus, et al.
Published: (2019)
by: Själander, Magnus, et al.
Published: (2019)
System-Level Performance Modeling of Photonic In-Memory Computing
by: Arockiaraj, Jebacyril, et al.
Published: (2026)
by: Arockiaraj, Jebacyril, et al.
Published: (2026)
Transactional Dynamics in Hyperledger Fabric: A Stochastic Modeling and Performance Evaluation of Permissioned Blockchains
by: Melo, Carlos, et al.
Published: (2025)
by: Melo, Carlos, et al.
Published: (2025)
Towards Federated Learning with On-device Training and Communication in 8-bit Floating Point
by: Wang, Bokun, et al.
Published: (2024)
by: Wang, Bokun, et al.
Published: (2024)
Distributed Inference Performance Optimization for LLMs on CPUs
by: He, Pujiang, et al.
Published: (2024)
by: He, Pujiang, et al.
Published: (2024)
AQUA: Network-Accelerated Memory Offloading for LLMs in Scale-Up GPU Domains
by: Kumar, Abhishek Vijaya, et al.
Published: (2024)
by: Kumar, Abhishek Vijaya, et al.
Published: (2024)
Sizey: Memory-Efficient Execution of Scientific Workflow Tasks
by: Bader, Jonathan, et al.
Published: (2024)
by: Bader, Jonathan, et al.
Published: (2024)
AME: An Efficient Heterogeneous Agentic Memory Engine for Smartphones
by: Zhao, Xinkui, et al.
Published: (2025)
by: Zhao, Xinkui, et al.
Published: (2025)
On the Performance and Memory Footprint of Distributed Training: An Empirical Study on Transformers
by: Lu, Zhengxian, et al.
Published: (2024)
by: Lu, Zhengxian, et al.
Published: (2024)
GMLake: Efficient and Transparent GPU Memory Defragmentation for Large-scale DNN Training with Virtual Memory Stitching
by: Guo, Cong, et al.
Published: (2024)
by: Guo, Cong, et al.
Published: (2024)
High-Performance and Power-Efficient Emulation of Matrix Multiplication using INT8 Matrix Engines
by: Uchino, Yuki, et al.
Published: (2025)
by: Uchino, Yuki, et al.
Published: (2025)
IntentContinuum: Using LLMs to Support Intent-Based Computing Across the Compute Continuum
by: Akbari, Negin, et al.
Published: (2025)
by: Akbari, Negin, et al.
Published: (2025)
Flash-KMeans: Fast and Memory-Efficient Exact K-Means
by: Yang, Shuo, et al.
Published: (2026)
by: Yang, Shuo, et al.
Published: (2026)
AGoQ: Activation and Gradient Quantization for Memory-Efficient Distributed Training of LLMs
by: Lin, Wenxiang, et al.
Published: (2026)
by: Lin, Wenxiang, et al.
Published: (2026)
Predictive Performance of Photonic SRAM-based In-Memory Computing for Tensor Decomposition
by: Wijeratne, Sasindu, et al.
Published: (2025)
by: Wijeratne, Sasindu, et al.
Published: (2025)
A Hierarchical Sharded Blockchain Balancing Performance and Availability
by: Jo, Yongrae, et al.
Published: (2025)
by: Jo, Yongrae, et al.
Published: (2025)
Performance Comparison of Graph Representations Which Support Dynamic Graph Updates
by: Sahu, Subhajit
Published: (2025)
by: Sahu, Subhajit
Published: (2025)
SiDP: Memory-Efficient Data Parallelism for Offline LLM Inference
by: Zhao, Alan, et al.
Published: (2026)
by: Zhao, Alan, et al.
Published: (2026)
Similar Items
-
SGEMM-cube: Precision-Recovery FP32 GEMM Approximation on Ascend NPUs with FP16 Matrix Engines
by: Xue, Weicheng, et al.
Published: (2025) -
Efficient CPU-GPU Collaborative Inference for MoE-based LLMs on Memory-Limited Systems
by: Huang, En-Ming, et al.
Published: (2025) -
Double-Precision Matrix Multiplication Emulation via Ozaki-II Scheme with FP8 Quantization
by: Uchino, Yuki, et al.
Published: (2026) -
Floating-Point Data Transformation for Lossless Compression
by: Jamalidinan, Samirasadat, et al.
Published: (2025) -
LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models
by: Park, Gunho, et al.
Published: (2022)