XFP: Quality-Targeted Adaptive Codebook Quantization with Sparse Outlier Separation for LLM Inference
Fuente:
arXiv
Saved in:
| Main Author: | Witt, Thomas |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
What Scalable Second-Order Information Knows for Pruning at Initialization
by: Navarrete, Ivo Gollini, et al.
Published: (2025)
by: Navarrete, Ivo Gollini, et al.
Published: (2025)
S$^3$: Structured Sparsity Specification
by: Ghriss, Ayoub
Published: (2026)
by: Ghriss, Ayoub
Published: (2026)
The E$Δ$-MHC-Geo Transformer: Adaptive Geodesic Operations with Guaranteed Orthogonality
by: Shahmansoori, Arash
Published: (2026)
by: Shahmansoori, Arash
Published: (2026)
Energy-Efficient Neuromorphic Computing for Edge AI: A Framework with Adaptive Spiking Neural Networks and Hardware-Aware Optimization
by: Imanov, Olaf Yunus Laitinen, et al.
Published: (2026)
by: Imanov, Olaf Yunus Laitinen, et al.
Published: (2026)
GRALIS: A Unified Canonical Framework for Linear Attribution Methods via Riesz Representation
by: Fanale, Raimondo
Published: (2026)
by: Fanale, Raimondo
Published: (2026)
From Features to Graphs: Exploring Graph Structures and Pairwise Interactions via GNNs
by: Yamchote, Phaphontee, et al.
Published: (2025)
by: Yamchote, Phaphontee, et al.
Published: (2025)
HyperMask: Adaptive Hypernetwork-based Masks for Continual Learning
by: Książek, Kamil, et al.
Published: (2023)
by: Książek, Kamil, et al.
Published: (2023)
Semantic Reward Collapse and the Preservation of Epistemic Integrity in Adaptive AI Systems
by: Parris, William
Published: (2026)
by: Parris, William
Published: (2026)
SafeAgent: Safeguarding LLM Agents via an Automated Risk Simulator
by: Zhou, Xueyang, et al.
Published: (2025)
by: Zhou, Xueyang, et al.
Published: (2025)
GraphNNK -- Graph Classification and Interpretability
by: Bolevic, Zeljko, et al.
Published: (2026)
by: Bolevic, Zeljko, et al.
Published: (2026)
multivariateGPT: a decoder-only transformer for multivariate categorical and numeric data
by: Loza, Andrew J., et al.
Published: (2025)
by: Loza, Andrew J., et al.
Published: (2025)
DecompKAN: Decomposed Patch-KAN for Long-Term Time Series Forecasting
by: Mysore, Naveen
Published: (2026)
by: Mysore, Naveen
Published: (2026)
Scaling Laws in the Tiny Regime: How Small Models Change Their Mistakes
by: Alnemari, Mohammed, et al.
Published: (2026)
by: Alnemari, Mohammed, et al.
Published: (2026)
Deep Learning-Based Forecasting of Boarding Patient Counts to Address ED Overcrowding
by: Vural, Orhun, et al.
Published: (2025)
by: Vural, Orhun, et al.
Published: (2025)
Binarized Neural Networks Converge Toward Algorithmic Simplicity: Empirical Support for the Learning-as-Compression Hypothesis
by: Sakabe, Eduardo Y., et al.
Published: (2025)
by: Sakabe, Eduardo Y., et al.
Published: (2025)
A general language model for peptide function identification
by: Zhai, Jixiu, et al.
Published: (2025)
by: Zhai, Jixiu, et al.
Published: (2025)
Momentum Attention: The Physics of In-Context Learning and Spectral Forensics for Mechanistic Interpretability
by: Maitra, Kingsuk
Published: (2026)
by: Maitra, Kingsuk
Published: (2026)
How Pruning Reshapes Features: Sparse Autoencoder Analysis of Weight-Pruned Language Models
by: Borobia, Hector, et al.
Published: (2026)
by: Borobia, Hector, et al.
Published: (2026)
Plug-and-Play Spiking Operators: Breaking the Nonlinearity Bottleneck in Spiking Transformers
by: Yuan, Xinzhe, et al.
Published: (2026)
by: Yuan, Xinzhe, et al.
Published: (2026)
Massive Redundancy in Gradient Transport Enables Sparse Online Learning
by: Merin, Aur Shalev
Published: (2026)
by: Merin, Aur Shalev
Published: (2026)
Ghosts of Softmax: Complex Singularities That Limit Safe Step Sizes in Cross-Entropy
by: Sao, Piyush
Published: (2026)
by: Sao, Piyush
Published: (2026)
Enhancing mortality prediction in cardiac arrest ICU patients through meta-modeling of structured clinical data from MIMIC-IV
by: Mamatov, Nursultan, et al.
Published: (2025)
by: Mamatov, Nursultan, et al.
Published: (2025)
Reliable Unlearning Harmful Information in LLMs with Metamorphosis Representation Projection
by: Wu, Chengcan, et al.
Published: (2025)
by: Wu, Chengcan, et al.
Published: (2025)
A Machine Learning Framework for Turbofan Health Estimation via Inverse Problem Formulation
by: Leyli-Abadi, Milad, et al.
Published: (2026)
by: Leyli-Abadi, Milad, et al.
Published: (2026)
HGTUL: A Hypergraph-based Model For Trajectory User Linking
by: Chang, Fengjie, et al.
Published: (2025)
by: Chang, Fengjie, et al.
Published: (2025)
Is ReLU Adversarially Robust?
by: Sooksatra, Korn, et al.
Published: (2024)
by: Sooksatra, Korn, et al.
Published: (2024)
Upside Down Reinforcement Learning with Policy Generators
by: Di Ventura, Jacopo, et al.
Published: (2025)
by: Di Ventura, Jacopo, et al.
Published: (2025)
Multi-Level Fusion Graph Neural Network for Molecule Property Prediction
by: Liu, XiaYu, et al.
Published: (2025)
by: Liu, XiaYu, et al.
Published: (2025)
CGLearn: Consistent Gradient-Based Learning for Out-of-Distribution Generalization
by: Chowdhury, Jawad, et al.
Published: (2024)
by: Chowdhury, Jawad, et al.
Published: (2024)
Concept Prerequisite Relation Prediction by Using Permutation-Equivariant Directed Graph Neural Networks
by: Qu, Xiran, et al.
Published: (2023)
by: Qu, Xiran, et al.
Published: (2023)
Simple Network Graph Comparative Learning
by: Yu, Qiang, et al.
Published: (2026)
by: Yu, Qiang, et al.
Published: (2026)
Mapping representations in Reinforcement Learning via Semantic Alignment for Zero-Shot Stitching
by: Ricciardi, Antonio Pio, et al.
Published: (2025)
by: Ricciardi, Antonio Pio, et al.
Published: (2025)
Efficient Morphology-Control Co-Design via Stackelberg Proximal Policy Optimization
by: Dai, Yanning, et al.
Published: (2026)
by: Dai, Yanning, et al.
Published: (2026)
Using latent representations to link disjoint longitudinal data for mixed-effects regression
by: Schächter, Clemens, et al.
Published: (2025)
by: Schächter, Clemens, et al.
Published: (2025)
Optimizing Inference in Transformer-Based Models: A Multi-Method Benchmark
by: Ho, Siu Hang, et al.
Published: (2025)
by: Ho, Siu Hang, et al.
Published: (2025)
Neural Concept Verifier: Scaling Prover-Verifier Games via Concept Encodings
by: Turan, Berkant, et al.
Published: (2025)
by: Turan, Berkant, et al.
Published: (2025)
CNN-LSTM Hybrid Model for AI-Driven Prediction of COVID-19 Severity from Spike Sequences and Clinical Data
by: Cheohen, Caio, et al.
Published: (2025)
by: Cheohen, Caio, et al.
Published: (2025)
Lost or Hidden? A Concept-Level Forgetting in Supervised Continual Learning
by: Filus, Katarzyna, et al.
Published: (2026)
by: Filus, Katarzyna, et al.
Published: (2026)
AMAR: Lightweight Attention-Based Multi-User Activity Recognition from Wi-Fi CSI
by: Mohammadi, Amirhossein, et al.
Published: (2026)
by: Mohammadi, Amirhossein, et al.
Published: (2026)
CVCM Track Circuits Pre-emptive Failure Diagnostics for Predictive Maintenance Using Deep Neural Networks
by: Mukherjee, Debdeep, et al.
Published: (2025)
by: Mukherjee, Debdeep, et al.
Published: (2025)
Similar Items
-
What Scalable Second-Order Information Knows for Pruning at Initialization
by: Navarrete, Ivo Gollini, et al.
Published: (2025) -
S$^3$: Structured Sparsity Specification
by: Ghriss, Ayoub
Published: (2026) -
The E$Δ$-MHC-Geo Transformer: Adaptive Geodesic Operations with Guaranteed Orthogonality
by: Shahmansoori, Arash
Published: (2026) -
Energy-Efficient Neuromorphic Computing for Edge AI: A Framework with Adaptive Spiking Neural Networks and Hardware-Aware Optimization
by: Imanov, Olaf Yunus Laitinen, et al.
Published: (2026) -
GRALIS: A Unified Canonical Framework for Linear Attribution Methods via Riesz Representation
by: Fanale, Raimondo
Published: (2026)