Influence-Inspired Spectral Rotations for Extreme Low-Bit LLM Quantization
Fuente:
arXiv
Saved in:
| Main Author: | Pavlov, Gorgi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Generalization Bound for a Family of Implicit Networks
by: Fung, Samy Wu, et al.
Published: (2024)
by: Fung, Samy Wu, et al.
Published: (2024)
From Features to Graphs: Exploring Graph Structures and Pairwise Interactions via GNNs
by: Yamchote, Phaphontee, et al.
Published: (2025)
by: Yamchote, Phaphontee, et al.
Published: (2025)
Optimal Abstractions for Verifying Properties of Kolmogorov-Arnold Networks (KANs)
by: Schwartz, Noah, et al.
Published: (2026)
by: Schwartz, Noah, et al.
Published: (2026)
Are Targeted Messages More Effective?
by: Grohe, Martin, et al.
Published: (2024)
by: Grohe, Martin, et al.
Published: (2024)
On Halting vs Converging in Recurrent Graph Neural Networks
by: Bollen, Jeroen, et al.
Published: (2026)
by: Bollen, Jeroen, et al.
Published: (2026)
Attention Maps in 3D Shape Classification for Dental Stage Estimation with Class Node Graph Attention Networks
by: Buyukcakir, Barkin, et al.
Published: (2025)
by: Buyukcakir, Barkin, et al.
Published: (2025)
Marrying Compressed Sensing and Deep Signal Separation
by: Hickok, Truman, et al.
Published: (2024)
by: Hickok, Truman, et al.
Published: (2024)
Optimizing MoE Routers: Design, Implementation, and Evaluation in Transformer Models
by: Harvey, Daniel Fidel, et al.
Published: (2025)
by: Harvey, Daniel Fidel, et al.
Published: (2025)
SplInterp: Improving our Understanding and Training of Sparse Autoencoders
by: Budd, Jeremy, et al.
Published: (2025)
by: Budd, Jeremy, et al.
Published: (2025)
The Inhibitor: ReLU and Addition-Based Attention for Efficient Transformers under Fully Homomorphic Encryption on the Torus
by: Brännvall, Rickard, et al.
Published: (2023)
by: Brännvall, Rickard, et al.
Published: (2023)
OptPO: Optimal Rollout Allocation for Test-time Policy Optimization
by: Wang, Youkang, et al.
Published: (2025)
by: Wang, Youkang, et al.
Published: (2025)
A Language Model-Driven Semi-Supervised Ensemble Framework for Illicit Market Detection Across Deep/Dark Web and Social Platforms
by: Yazdanjue, Navid, et al.
Published: (2025)
by: Yazdanjue, Navid, et al.
Published: (2025)
Flow Matching on Symmetric Spaces
by: Ruscelli, Francesco, et al.
Published: (2026)
by: Ruscelli, Francesco, et al.
Published: (2026)
BrainDistill: Implantable Motor Decoding with Task-Specific Knowledge Distillation
by: Xie, Yuhan, et al.
Published: (2026)
by: Xie, Yuhan, et al.
Published: (2026)
Unveiling Memorization-Generalization Coexistence: A Case Study on Arithmetic Tasks with Label Noise
by: Liu, Linyu, et al.
Published: (2026)
by: Liu, Linyu, et al.
Published: (2026)
Assessing local deformation and computing scalar curvature with nonlinear conformal regularization of decoders
by: Couéraud, Benjamin, et al.
Published: (2025)
by: Couéraud, Benjamin, et al.
Published: (2025)
Deep learning the Hurst parameter of linear fractional processes and assessing its reliability
by: Boros, Dániel, et al.
Published: (2024)
by: Boros, Dániel, et al.
Published: (2024)
Embracing the black box: Heading towards foundation models for causal discovery from time series data
by: Stein, Gideon, et al.
Published: (2024)
by: Stein, Gideon, et al.
Published: (2024)
Beyond Backpropagation: Exploring Innovative Algorithms for Energy-Efficient Deep Neural Network Training
by: Spyra, Przemysław
Published: (2025)
by: Spyra, Przemysław
Published: (2025)
Front-propagation Algorithm: Explainable AI Technique for Extracting Linear Function Approximations from Neural Networks
by: Viaña, Javier
Published: (2024)
by: Viaña, Javier
Published: (2024)
rETF-semiSL: Semi-Supervised Learning for Neural Collapse in Temporal Data
by: Xie, Yuhan, et al.
Published: (2025)
by: Xie, Yuhan, et al.
Published: (2025)
TrackGPT -- A generative pre-trained transformer for cross-domain entity trajectory forecasting
by: Stroh, Nicholas
Published: (2024)
by: Stroh, Nicholas
Published: (2024)
OPENXRD: A Comprehensive Benchmark Framework for LLM/MLLM XRD Question Answering
by: Vosoughi, Ali, et al.
Published: (2025)
by: Vosoughi, Ali, et al.
Published: (2025)
Teaching and Learning under Deductive Errors
by: Telle, Jan Arne, et al.
Published: (2026)
by: Telle, Jan Arne, et al.
Published: (2026)
The Boundaries of Verifiable Accuracy, Robustness, and Generalisation in Deep Learning
by: Bastounis, Alexander, et al.
Published: (2023)
by: Bastounis, Alexander, et al.
Published: (2023)
Graph Neural Networks for Brain Graph Learning: A Survey
by: Luo, Xuexiong, et al.
Published: (2024)
by: Luo, Xuexiong, et al.
Published: (2024)
Enhanced QKNorm normalization for neural transformers with the Lp norm
by: Lopez-Rubio, Ezequiel, et al.
Published: (2026)
by: Lopez-Rubio, Ezequiel, et al.
Published: (2026)
Pay Attention to What You Need
by: Gao, Yifei, et al.
Published: (2023)
by: Gao, Yifei, et al.
Published: (2023)
Deriving Decoder-Free Sparse Autoencoders from First Principles
by: Oursland, Alan
Published: (2026)
by: Oursland, Alan
Published: (2026)
ProactBench: Beyond What The User Asked For
by: Harfi, Sepehr, et al.
Published: (2026)
by: Harfi, Sepehr, et al.
Published: (2026)
Generative Design of a Gas Turbine Combustor Using Invertible Neural Networks
by: Krüger, Patrick, et al.
Published: (2026)
by: Krüger, Patrick, et al.
Published: (2026)
Centralized vs. Decentralized Multi-Agent Reinforcement Learning for Enhanced Control of Electric Vehicle Charging Networks
by: Shojaeighadikolaei, Amin, et al.
Published: (2024)
by: Shojaeighadikolaei, Amin, et al.
Published: (2024)
ADAM: An AI Reasoning and Bioinformatics Model for Alzheimer's Disease Detection and Microbiome-Clinical Data Integration
by: Huang, Ziyuan, et al.
Published: (2025)
by: Huang, Ziyuan, et al.
Published: (2025)
Uncertainty Estimation and Quantification for LLMs: A Simple Supervised Approach
by: Liu, Linyu, et al.
Published: (2024)
by: Liu, Linyu, et al.
Published: (2024)
Thanos: A Block-wise Pruning Algorithm for Efficient Large Language Model Compression
by: Ilin, Ivan, et al.
Published: (2025)
by: Ilin, Ivan, et al.
Published: (2025)
3ET: Efficient Event-based Eye Tracking using a Change-Based ConvLSTM Network
by: Chen, Qinyu, et al.
Published: (2023)
by: Chen, Qinyu, et al.
Published: (2023)
Parameter-Efficient Transformer Embeddings
by: Ndubuaku, Henry, et al.
Published: (2025)
by: Ndubuaku, Henry, et al.
Published: (2025)
Explainable Image Similarity: Integrating Siamese Networks and Grad-CAM
by: Livieris, Ioannis E., et al.
Published: (2023)
by: Livieris, Ioannis E., et al.
Published: (2023)
DYNAMAX: Dynamic computing for Transformers and Mamba based architectures
by: Nogales, Miguel, et al.
Published: (2025)
by: Nogales, Miguel, et al.
Published: (2025)
Multimodal Multi-Agent Ransomware Analysis Using AutoGen
by: Khan, Asifullah, et al.
Published: (2026)
by: Khan, Asifullah, et al.
Published: (2026)
Similar Items
-
A Generalization Bound for a Family of Implicit Networks
by: Fung, Samy Wu, et al.
Published: (2024) -
From Features to Graphs: Exploring Graph Structures and Pairwise Interactions via GNNs
by: Yamchote, Phaphontee, et al.
Published: (2025) -
Optimal Abstractions for Verifying Properties of Kolmogorov-Arnold Networks (KANs)
by: Schwartz, Noah, et al.
Published: (2026) -
Are Targeted Messages More Effective?
by: Grohe, Martin, et al.
Published: (2024) -
On Halting vs Converging in Recurrent Graph Neural Networks
by: Bollen, Jeroen, et al.
Published: (2026)