FFSplit: Split Feed-Forward Network For Optimizing Accuracy-Efficiency Trade-off in Language Model Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Zirui, Song, Qingquan, Xiao, Qiang Charles, Selvaraj, Sathiya Keerthi, Mazumder, Rahul, Gupta, Aman, Hu, Xia |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Precise Characterization of SGD Stability Using Loss Surface Geometry
by: Dexter, Gregory, et al.
Published: (2024)
by: Dexter, Gregory, et al.
Published: (2024)
Exploring Accuracy-Fairness Trade-off in Large Language Models
by: Zhang, Qingquan, et al.
Published: (2024)
by: Zhang, Qingquan, et al.
Published: (2024)
AlphaPO: Reward Shape Matters for LLM Alignment
by: Gupta, Aman, et al.
Published: (2025)
by: Gupta, Aman, et al.
Published: (2025)
You Only Debias Once: Towards Flexible Accuracy-Fairness Trade-offs at Inference Time
by: Han, Xiaotian, et al.
Published: (2025)
by: Han, Xiaotian, et al.
Published: (2025)
Effective Quantization of Muon Optimizer States
by: Gupta, Aman, et al.
Published: (2025)
by: Gupta, Aman, et al.
Published: (2025)
From Rays to Projections: Better Inputs for Feed-Forward View Synthesis
by: Wu, Zirui, et al.
Published: (2026)
by: Wu, Zirui, et al.
Published: (2026)
Enhancing Accuracy-Privacy Trade-off in Differentially Private Split Learning
by: Pham, Ngoc Duy, et al.
Published: (2023)
by: Pham, Ngoc Duy, et al.
Published: (2023)
Efficient user history modeling with amortized inference for deep learning recommendation models
by: Hertel, Lars, et al.
Published: (2024)
by: Hertel, Lars, et al.
Published: (2024)
Accuracy-Privacy Trade-off in the Mitigation of Membership Inference Attack in Federated Learning
by: Ahamed, Sayyed Farid, et al.
Published: (2024)
by: Ahamed, Sayyed Farid, et al.
Published: (2024)
Reasoning Models Can be Accurately Pruned Via Chain-of-Thought Reconstruction
by: Lucas, Ryan, et al.
Published: (2025)
by: Lucas, Ryan, et al.
Published: (2025)
MOSS: Multi-Objective Optimization for Stable Rule Sets
by: Liu, Brian, et al.
Published: (2025)
by: Liu, Brian, et al.
Published: (2025)
Characterizing the Accuracy -- Efficiency Trade-off of Low-rank Decomposition in Language Models
by: Moar, Chakshu, et al.
Published: (2024)
by: Moar, Chakshu, et al.
Published: (2024)
FAST: An Optimization Framework for Fast Additive Segmentation in Transparent ML
by: Liu, Brian, et al.
Published: (2024)
by: Liu, Brian, et al.
Published: (2024)
The Efficiency vs. Accuracy Trade-off: Optimizing RAG-Enhanced LLM Recommender Systems Using Multi-Head Early Exit
by: Zhou, Huixue, et al.
Published: (2025)
by: Zhou, Huixue, et al.
Published: (2025)
Preserving Deep Representations In One-Shot Pruning: A Hessian-Free Second-Order Optimization Framework
by: Lucas, Ryan, et al.
Published: (2024)
by: Lucas, Ryan, et al.
Published: (2024)
Meta-Black-Box Optimization with Ensemble Surrogate Modeling for Robustness-Accuracy Trade-off within SAEA
by: Jin, Xiao, et al.
Published: (2026)
by: Jin, Xiao, et al.
Published: (2026)
Variational Inference for Uncertainty Quantification: an Analysis of Trade-offs
by: Margossian, Charles C., et al.
Published: (2024)
by: Margossian, Charles C., et al.
Published: (2024)
Quantization Impact on the Accuracy and Communication Efficiency Trade-off in Federated Learning for Aerospace Predictive Maintenance
by: Loukili, Abdelkarim
Published: (2026)
by: Loukili, Abdelkarim
Published: (2026)
optimizn: a Python Library for Developing Customized Optimization Algorithms
by: Sathiya, Akshay, et al.
Published: (2025)
by: Sathiya, Akshay, et al.
Published: (2025)
Neural Optimization with Adaptive Heuristics for Intelligent Marketing System
by: Wei, Changshuai, et al.
Published: (2024)
by: Wei, Changshuai, et al.
Published: (2024)
Optimizing Dense Feed-Forward Neural Networks
by: Balderas, Luis, et al.
Published: (2023)
by: Balderas, Luis, et al.
Published: (2023)
Characterizing the Accuracy-Communication-Privacy Trade-off in Distributed Stochastic Convex Optimization
by: Salgia, Sudeep, et al.
Published: (2025)
by: Salgia, Sudeep, et al.
Published: (2025)
Stability and Accuracy Trade-offs in Statistical Estimation
by: Chakraborty, Abhinav, et al.
Published: (2026)
by: Chakraborty, Abhinav, et al.
Published: (2026)
Understanding the Trade-offs in Accuracy and Uncertainty Quantification: Architecture and Inference Choices in Bayesian Neural Networks
by: Sheinkman, Alisa, et al.
Published: (2025)
by: Sheinkman, Alisa, et al.
Published: (2025)
From Raw Data to Structural Semantics: Trade-offs among Distortion, Rate, and Inference Accuracy
by: Asirimath, Charmin, et al.
Published: (2024)
by: Asirimath, Charmin, et al.
Published: (2024)
On Efficiency-Effectiveness Trade-off of Diffusion-based Recommenders
by: Mao, Wenyu, et al.
Published: (2025)
by: Mao, Wenyu, et al.
Published: (2025)
Feed-Forward Latent Domain Adaptation
by: Bohdal, Ondrej, et al.
Published: (2022)
by: Bohdal, Ondrej, et al.
Published: (2022)
Enhancing Energy Efficiency and Battery Lifetime in Cardiac Implantable Devices using Optimized RNN‐LSTM
by: Subramanian Nagakumararaj, et al.
Published: (2025)
by: Subramanian Nagakumararaj, et al.
Published: (2025)
Sparse Gaussian Graphical Models with Discrete Optimization: Computational and Statistical Perspectives
by: Behdin, Kayhan, et al.
Published: (2023)
by: Behdin, Kayhan, et al.
Published: (2023)
Queueing-Aware Optimization of Reasoning Tokens for Accuracy-Latency Trade-offs in LLM Servers
by: Ozbas, Emre, et al.
Published: (2026)
by: Ozbas, Emre, et al.
Published: (2026)
Controllable Pareto Trade-off between Fairness and Accuracy
by: Du, Yongkang, et al.
Published: (2025)
by: Du, Yongkang, et al.
Published: (2025)
Trading-off Accuracy and Communication Cost in Federated Learning
by: Villani, Mattia Jacopo, et al.
Published: (2025)
by: Villani, Mattia Jacopo, et al.
Published: (2025)
Towards Understanding Systems Trade-offs in Retrieval-Augmented Generation Model Inference
by: Shen, Michael, et al.
Published: (2024)
by: Shen, Michael, et al.
Published: (2024)
Flash Multi-Head Feed-Forward Network
by: Zhang, Minshen, et al.
Published: (2025)
by: Zhang, Minshen, et al.
Published: (2025)
LLM Query Scheduling with Prefix Reuse and Latency Constraints
by: Dexter, Gregory, et al.
Published: (2025)
by: Dexter, Gregory, et al.
Published: (2025)
Feed-Forward Optimization With Delayed Feedback for Neural Network Training
by: Flügel, Katharina, et al.
Published: (2023)
by: Flügel, Katharina, et al.
Published: (2023)
Sparse NMF with Archetypal Regularization: Computational and Robustness Properties
by: Behdin, Kayhan, et al.
Published: (2021)
by: Behdin, Kayhan, et al.
Published: (2021)
Randomization Can Reduce Both Bias and Variance: A Case Study in Random Forests
by: Liu, Brian, et al.
Published: (2024)
by: Liu, Brian, et al.
Published: (2024)
Sparse PCA: A New Scalable Estimator Based On Integer Programming
by: Behdin, Kayhan, et al.
Published: (2021)
by: Behdin, Kayhan, et al.
Published: (2021)
Characterizing Accuracy Trade-offs of EEG Applications on Embedded HMPs
by: Taufique, Zain, et al.
Published: (2024)
by: Taufique, Zain, et al.
Published: (2024)
Similar Items
-
A Precise Characterization of SGD Stability Using Loss Surface Geometry
by: Dexter, Gregory, et al.
Published: (2024) -
Exploring Accuracy-Fairness Trade-off in Large Language Models
by: Zhang, Qingquan, et al.
Published: (2024) -
AlphaPO: Reward Shape Matters for LLM Alignment
by: Gupta, Aman, et al.
Published: (2025) -
You Only Debias Once: Towards Flexible Accuracy-Fairness Trade-offs at Inference Time
by: Han, Xiaotian, et al.
Published: (2025) -
Effective Quantization of Muon Optimizer States
by: Gupta, Aman, et al.
Published: (2025)