FPTC: A Fast Parallel Transform-based Codec for Efficient Asymmetric Signal Compression
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mechels, Ben, Billmeyer, Ryan, Chen, Alexander, Li, Shiyang, Ding, Caiwen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RTop-K: Ultra-Fast Row-Wise Top-K Selection for Neural Network Acceleration on GPUs
von: Xie, Xi, et al.
Veröffentlicht: (2024)
von: Xie, Xi, et al.
Veröffentlicht: (2024)
FastSet: Parallel Claim Settlement
von: Chen, Xiaohong, et al.
Veröffentlicht: (2025)
von: Chen, Xiaohong, et al.
Veröffentlicht: (2025)
Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services
von: Chen, Haoyu, et al.
Veröffentlicht: (2025)
von: Chen, Haoyu, et al.
Veröffentlicht: (2025)
CoCoDiff: Optimizing Collective Communications for Distributed Diffusion Transformer Inference Under Ulysses Sequence Parallelism
von: Ma, Bin, et al.
Veröffentlicht: (2026)
von: Ma, Bin, et al.
Veröffentlicht: (2026)
cuFastTuckerPlus: A Stochastic Parallel Sparse FastTucker Decomposition Using GPU Tensor Cores
von: Li, Zixuan, et al.
Veröffentlicht: (2024)
von: Li, Zixuan, et al.
Veröffentlicht: (2024)
Communication-Efficient Serving for Video Diffusion Models with Latent Parallelism
von: Wu, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Wu, Zhiyuan, et al.
Veröffentlicht: (2025)
Fast and Space-Efficient Parallel Algorithms for Influence Maximization
von: Wang, Letong, et al.
Veröffentlicht: (2023)
von: Wang, Letong, et al.
Veröffentlicht: (2023)
CodecSight: Leveraging Video Codec Signals for Efficient Streaming VLM Inference
von: Zou, Yulin, et al.
Veröffentlicht: (2026)
von: Zou, Yulin, et al.
Veröffentlicht: (2026)
StarTrail: Concentric Ring Sequence Parallelism for Efficient Near-Infinite-Context Transformer Model Training
von: Liu, Ziming, et al.
Veröffentlicht: (2024)
von: Liu, Ziming, et al.
Veröffentlicht: (2024)
WRATH: Workload Resilience Across Task Hierarchies in Task-based Parallel Programming Frameworks
von: Zhou, Sicheng, et al.
Veröffentlicht: (2025)
von: Zhou, Sicheng, et al.
Veröffentlicht: (2025)
EXaCTz: Guaranteed Extremum Graph and Contour Tree Preservation for Distributed- and GPU-Parallel Lossy Compression
von: Li, Yuxiao, et al.
Veröffentlicht: (2026)
von: Li, Yuxiao, et al.
Veröffentlicht: (2026)
ZeroPP: Unleashing Exceptional Parallelism Efficiency through Tensor-Parallelism-Free Methodology
von: Tang, Ding, et al.
Veröffentlicht: (2024)
von: Tang, Ding, et al.
Veröffentlicht: (2024)
NetSenseML: Network-Adaptive Compression for Efficient Distributed Machine Learning
von: Wang, Yisu, et al.
Veröffentlicht: (2025)
von: Wang, Yisu, et al.
Veröffentlicht: (2025)
pMSz: A Distributed Parallel Algorithm for Correcting Extrema and Morse Smale Segmentations in Lossy Compression
von: Li, Yuxiao, et al.
Veröffentlicht: (2026)
von: Li, Yuxiao, et al.
Veröffentlicht: (2026)
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism
von: Zhang, Zheng, et al.
Veröffentlicht: (2025)
von: Zhang, Zheng, et al.
Veröffentlicht: (2025)
SPPO:Efficient Long-sequence LLM Training via Adaptive Sequence Pipeline Parallel Offloading
von: Chen, Qiaoling, et al.
Veröffentlicht: (2025)
von: Chen, Qiaoling, et al.
Veröffentlicht: (2025)
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism
von: Wei, Jinhui, et al.
Veröffentlicht: (2025)
von: Wei, Jinhui, et al.
Veröffentlicht: (2025)
Heimdall++: Optimizing GPU Utilization and Pipeline Parallelism for Efficient Single-Pulse Detection
von: Xia, Bingzheng, et al.
Veröffentlicht: (2025)
von: Xia, Bingzheng, et al.
Veröffentlicht: (2025)
HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference
von: Lin, Haoran, et al.
Veröffentlicht: (2025)
von: Lin, Haoran, et al.
Veröffentlicht: (2025)
S-HPLB: Efficient LLM Attention Serving via Sparsity-Aware Head Parallelism Load Balance
von: Liu, Di, et al.
Veröffentlicht: (2026)
von: Liu, Di, et al.
Veröffentlicht: (2026)
A Flexible Programmable Pipeline Parallelism Framework for Efficient DNN Training
von: Jiang, Lijuan, et al.
Veröffentlicht: (2025)
von: Jiang, Lijuan, et al.
Veröffentlicht: (2025)
pdGRASS: A Fast Parallel Density-Aware Algorithm for Graph Spectral Sparsification
von: Zhao, Tiancheng, et al.
Veröffentlicht: (2025)
von: Zhao, Tiancheng, et al.
Veröffentlicht: (2025)
LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelism
von: Gu, Diandian, et al.
Veröffentlicht: (2024)
von: Gu, Diandian, et al.
Veröffentlicht: (2024)
Parallel Track Transformers: Enabling Fast GPU Inference with Reduced Synchronization
von: Wang, Chong, et al.
Veröffentlicht: (2026)
von: Wang, Chong, et al.
Veröffentlicht: (2026)
JanusPipe: Efficient Pipeline Parallel Training for Machine Learning Interatomic Potentials
von: Wang, Hongyu, et al.
Veröffentlicht: (2026)
von: Wang, Hongyu, et al.
Veröffentlicht: (2026)
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism
von: Qing, Yuhao, et al.
Veröffentlicht: (2025)
von: Qing, Yuhao, et al.
Veröffentlicht: (2025)
Exploring Fast Fourier Transforms on the Tenstorrent Wormhole
von: Brown, Nick, et al.
Veröffentlicht: (2025)
von: Brown, Nick, et al.
Veröffentlicht: (2025)
Efficient Remote KV Cache Reuse with GPU-native Video Codec
von: Mi, Liang, et al.
Veröffentlicht: (2026)
von: Mi, Liang, et al.
Veröffentlicht: (2026)
Dynamic Contract Analysis for Parallel Programming Models
von: Oraji, Yussur Mustafa, et al.
Veröffentlicht: (2026)
von: Oraji, Yussur Mustafa, et al.
Veröffentlicht: (2026)
Efficient Parallel Compilation and Profiling of Quantum Circuits at Large Scales
von: Moore, Jane, et al.
Veröffentlicht: (2026)
von: Moore, Jane, et al.
Veröffentlicht: (2026)
Efficient Task Graph Scheduling for Parallel QR Factorization in SLSQP
von: Chatterjee, Soumyajit, et al.
Veröffentlicht: (2025)
von: Chatterjee, Soumyajit, et al.
Veröffentlicht: (2025)
Efficient Parallel Execution of Blockchain Transactions Leveraging Conflict Specifications
von: Anjana, Parwat Singh, et al.
Veröffentlicht: (2025)
von: Anjana, Parwat Singh, et al.
Veröffentlicht: (2025)
Towards Efficient Verification of Parallel Applications with Mc SimGrid
von: Laurent, Matthieu, et al.
Veröffentlicht: (2025)
von: Laurent, Matthieu, et al.
Veröffentlicht: (2025)
DAG-based Consensus with Asymmetric Trust [Extended Version]
von: Amores-Sesar, Ignacio, et al.
Veröffentlicht: (2025)
von: Amores-Sesar, Ignacio, et al.
Veröffentlicht: (2025)
CudaForge: An Agent Framework with Hardware Feedback for CUDA Kernel Optimization
von: Zhang, Zijian, et al.
Veröffentlicht: (2025)
von: Zhang, Zijian, et al.
Veröffentlicht: (2025)
Folding Tensor and Sequence Parallelism for Memory-Efficient Transformer Training & Inference
von: Shyam, Vasu, et al.
Veröffentlicht: (2026)
von: Shyam, Vasu, et al.
Veröffentlicht: (2026)
GRNND: A GPU-Parallel Relative NN-Descent Algorithm for Efficient Approximate Nearest Neighbor Graph Construction
von: Li, Xiang, et al.
Veröffentlicht: (2025)
von: Li, Xiang, et al.
Veröffentlicht: (2025)
FourierCompress: Layer-Aware Spectral Activation Compression for Efficient and Accurate Collaborative LLM Inference
von: Ma, Jian, et al.
Veröffentlicht: (2025)
von: Ma, Jian, et al.
Veröffentlicht: (2025)
FlashMP: Fast Discrete Transform-Based Solver for Preconditioning Maxwell's Equations on GPUs
von: Zhang, Haoyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Haoyuan, et al.
Veröffentlicht: (2025)
Flash-KMeans: Fast and Memory-Efficient Exact K-Means
von: Yang, Shuo, et al.
Veröffentlicht: (2026)
von: Yang, Shuo, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
RTop-K: Ultra-Fast Row-Wise Top-K Selection for Neural Network Acceleration on GPUs
von: Xie, Xi, et al.
Veröffentlicht: (2024) -
FastSet: Parallel Claim Settlement
von: Chen, Xiaohong, et al.
Veröffentlicht: (2025) -
Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services
von: Chen, Haoyu, et al.
Veröffentlicht: (2025) -
CoCoDiff: Optimizing Collective Communications for Distributed Diffusion Transformer Inference Under Ulysses Sequence Parallelism
von: Ma, Bin, et al.
Veröffentlicht: (2026) -
cuFastTuckerPlus: A Stochastic Parallel Sparse FastTucker Decomposition Using GPU Tensor Cores
von: Li, Zixuan, et al.
Veröffentlicht: (2024)