Saved in:
| Main Authors: | Xiang, Maoyang, Wang, Bo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2603.19338 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On-Device Qwen2.5: Efficient LLM Inference with Model Compression and Hardware Acceleration
by: Xiang, Maoyang, et al.
Published: (2025)
by: Xiang, Maoyang, et al.
Published: (2025)
GoQuant: Geometric Orthogonal Residual Projection for Multiplier-Free Power-of-Two Transformer Quantization
by: Xiang, Maoyang, et al.
Published: (2026)
by: Xiang, Maoyang, et al.
Published: (2026)
Developing Training Procedures for Piecewise-linear Spline Activation Functions in Neural Networks
by: Patty, William H
Published: (2025)
by: Patty, William H
Published: (2025)
End-to-End On-Device Quantization-Aware Training for LLMs at Inference Cost
by: Tan, Qitao, et al.
Published: (2025)
by: Tan, Qitao, et al.
Published: (2025)
On the Expressive Power of Transformers for Maxout Networks and Continuous Piecewise Linear Functions
by: Gu, Linyan, et al.
Published: (2026)
by: Gu, Linyan, et al.
Published: (2026)
Accelerating Transformer Inference and Training with 2:4 Activation Sparsity
by: Haziza, Daniel, et al.
Published: (2025)
by: Haziza, Daniel, et al.
Published: (2025)
EdgeFlex-Transformer: Transformer Inference for Edge Devices
by: Mohammad, Shoaib, et al.
Published: (2025)
by: Mohammad, Shoaib, et al.
Published: (2025)
Going Beyond the Edge: Distributed Inference of Transformer Models on Ultra-Low-Power Wireless Devices
by: Gräfe, Alexander, et al.
Published: (2026)
by: Gräfe, Alexander, et al.
Published: (2026)
Privacy-Aware Multi-Device Cooperative Edge Inference with Distributed Resource Bidding
by: Zhuang, Wenhao, et al.
Published: (2024)
by: Zhuang, Wenhao, et al.
Published: (2024)
R-Sparse: Rank-Aware Activation Sparsity for Efficient LLM Inference
by: Zhang, Zhenyu, et al.
Published: (2025)
by: Zhang, Zhenyu, et al.
Published: (2025)
Representing Piecewise-Linear Functions by Functions with Minimal Arity
by: Koutschan, Christoph, et al.
Published: (2024)
by: Koutschan, Christoph, et al.
Published: (2024)
Non-Singularity of the Gradient Descent map for Neural Networks with Piecewise Analytic Activations
by: Crăciun, Alexandru, et al.
Published: (2025)
by: Crăciun, Alexandru, et al.
Published: (2025)
Don't Forget the Nonlinearity: Unlocking Activation Functions in Efficient Fine-Tuning
by: Yin, Bo, et al.
Published: (2025)
by: Yin, Bo, et al.
Published: (2025)
Piecewise Deterministic Markov Processes for Bayesian Inference of PDE Coefficients
by: Riccius, Leon, et al.
Published: (2026)
by: Riccius, Leon, et al.
Published: (2026)
AffineLens: Capturing the Continuous Piecewise Affine Functions of Neural Networks
by: Wei, Yi, et al.
Published: (2026)
by: Wei, Yi, et al.
Published: (2026)
Hysteresis Activation Function for Efficient Inference
by: Kimhi, Moshe, et al.
Published: (2024)
by: Kimhi, Moshe, et al.
Published: (2024)
Piecewise Normalizing Flows
by: Bevins, Harry, et al.
Published: (2023)
by: Bevins, Harry, et al.
Published: (2023)
ASTRA: Communication-Efficient Acceleration for Multi-Device Transformer Inference
by: Liu, Xiao, et al.
Published: (2025)
by: Liu, Xiao, et al.
Published: (2025)
Secure Transformer Inference Protocol
by: Yuan, Mu, et al.
Published: (2023)
by: Yuan, Mu, et al.
Published: (2023)
Region Seeding via Pre-Activation Regularization: A Geometric View of Piecewise Affine Neural Networks
by: Wei, Yi, et al.
Published: (2026)
by: Wei, Yi, et al.
Published: (2026)
Training-Time Batch Normalization Reshapes Local Partition Geometry in Piecewise-Affine Networks
by: Qi, Xuan, et al.
Published: (2026)
by: Qi, Xuan, et al.
Published: (2026)
Decomposition Polyhedra of Piecewise Linear Functions
by: Brandenburg, Marie-Charlotte, et al.
Published: (2024)
by: Brandenburg, Marie-Charlotte, et al.
Published: (2024)
Topology-Aware Activation Functions in Neural Networks
by: Snopov, Pavel, et al.
Published: (2025)
by: Snopov, Pavel, et al.
Published: (2025)
WiSparse: Boosting LLM Inference Efficiency with Weight-Aware Mixed Activation Sparsity
by: Chen, Lei, et al.
Published: (2026)
by: Chen, Lei, et al.
Published: (2026)
Piecewise deterministic generative models
by: Bertazzi, Andrea, et al.
Published: (2024)
by: Bertazzi, Andrea, et al.
Published: (2024)
TTQ: Activation-Aware Test-Time Quantization to Accelerate LLM Inference On The Fly
by: Koike-Akino, Toshiaki, et al.
Published: (2026)
by: Koike-Akino, Toshiaki, et al.
Published: (2026)
Transformer-Based Approach for Automated Functional Group Replacement in Chemical Compounds
by: Pan, Bo, et al.
Published: (2026)
by: Pan, Bo, et al.
Published: (2026)
DAF: An Efficient End-to-End Dynamic Activation Framework for on-Device DNN Training
by: Liu, Renyuan, et al.
Published: (2025)
by: Liu, Renyuan, et al.
Published: (2025)
Beat the long tail: Distribution-Aware Speculative Decoding for RL Training
by: Shao, Zelei, et al.
Published: (2025)
by: Shao, Zelei, et al.
Published: (2025)
Continuity, Piecewise Corrections, and Functor Models in Function-Based Learning
by: Harby, John
Published: (2026)
by: Harby, John
Published: (2026)
Kraken: Inherently Parallel Transformers For Efficient Multi-Device Inference
by: Prabhakar, Rohan Baskar, et al.
Published: (2024)
by: Prabhakar, Rohan Baskar, et al.
Published: (2024)
ChunkFlow: Communication-Aware Chunked Prefetching for Layerwise Offloading in Distributed Diffusion Transformer Inference
by: Meng, Han, et al.
Published: (2026)
by: Meng, Han, et al.
Published: (2026)
Steklov Activations: Piecewise-Polynomial Gates with Compact Support and Tunable Sparsity
by: Masalskikh, Aleksandr
Published: (2026)
by: Masalskikh, Aleksandr
Published: (2026)
CoFormer: Collaborating with Heterogeneous Edge Devices for Scalable Transformer Inference
by: Xu, Guanyu, et al.
Published: (2025)
by: Xu, Guanyu, et al.
Published: (2025)
Distributed Sign Momentum with Local Steps for Training Transformers
by: Yu, Shuhua, et al.
Published: (2024)
by: Yu, Shuhua, et al.
Published: (2024)
Symmetry-Aware Transformer Training for Automated Planning
by: Fritzsche, Markus, et al.
Published: (2025)
by: Fritzsche, Markus, et al.
Published: (2025)
Flow-Transformed Implicit Processes for Function-Space Variational Inference
by: Ortega, Luis A., et al.
Published: (2026)
by: Ortega, Luis A., et al.
Published: (2026)
NEST: Network- and Memory-Aware Device Placement For Distributed Deep Learning
by: Wang, Irene, et al.
Published: (2026)
by: Wang, Irene, et al.
Published: (2026)
Digi-Q: Learning Q-Value Functions for Training Device-Control Agents
by: Bai, Hao, et al.
Published: (2025)
by: Bai, Hao, et al.
Published: (2025)
PolyLUT: Learning Piecewise Polynomials for Ultra-Low Latency FPGA LUT-based Inference
by: Andronic, Marta, et al.
Published: (2023)
by: Andronic, Marta, et al.
Published: (2023)
Similar Items
-
On-Device Qwen2.5: Efficient LLM Inference with Model Compression and Hardware Acceleration
by: Xiang, Maoyang, et al.
Published: (2025) -
GoQuant: Geometric Orthogonal Residual Projection for Multiplier-Free Power-of-Two Transformer Quantization
by: Xiang, Maoyang, et al.
Published: (2026) -
Developing Training Procedures for Piecewise-linear Spline Activation Functions in Neural Networks
by: Patty, William H
Published: (2025) -
End-to-End On-Device Quantization-Aware Training for LLMs at Inference Cost
by: Tan, Qitao, et al.
Published: (2025) -
On the Expressive Power of Transformers for Maxout Networks and Continuous Piecewise Linear Functions
by: Gu, Linyan, et al.
Published: (2026)