InhibiDistilbert: Knowledge Distillation for a ReLU and Addition-based Transformer
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Tony, Brännvall, Rickard |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Inhibitor: ReLU and Addition-Based Attention for Efficient Transformers under Fully Homomorphic Encryption on the Torus
by: Brännvall, Rickard, et al.
Published: (2023)
by: Brännvall, Rickard, et al.
Published: (2023)
Graded Transformers
by: Shaska Sr, Tony
Published: (2025)
by: Shaska Sr, Tony
Published: (2025)
Think Thrice Before You Speak: Dual knowledge-enhanced Theory-of-Mind Reasoning for Persuasive Agents
by: Ma, Minghui, et al.
Published: (2026)
by: Ma, Minghui, et al.
Published: (2026)
Modularity in Transformers: Investigating Neuron Separability & Specialization
by: Pochinkov, Nicholas, et al.
Published: (2024)
by: Pochinkov, Nicholas, et al.
Published: (2024)
When Does Content-Based Routing Work? Representation Requirements for Selective Attention in Hybrid Sequence Models
by: Basu, Abhinaba
Published: (2026)
by: Basu, Abhinaba
Published: (2026)
NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution
by: Breneur, Oleksandr Marchenko, et al.
Published: (2026)
by: Breneur, Oleksandr Marchenko, et al.
Published: (2026)
Rethinking the Multilingual Reasoning Gap with Layer Swap
by: Lasbordes, Maxence, et al.
Published: (2026)
by: Lasbordes, Maxence, et al.
Published: (2026)
Energy-Efficient Information Representation in MNIST Classification Using Biologically Inspired Learning
by: Stricker, Patrick, et al.
Published: (2026)
by: Stricker, Patrick, et al.
Published: (2026)
Harnessing non-adversarial robustness in large language models
by: Zhou, Qinghua, et al.
Published: (2026)
by: Zhou, Qinghua, et al.
Published: (2026)
Semantic Retention and Extreme Compression in LLMs: Can We Have Both?
by: Laborde, Stanislas, et al.
Published: (2025)
by: Laborde, Stanislas, et al.
Published: (2025)
CLMN: Concept based Language Models via Neural Symbolic Reasoning
by: Yang, Yibo
Published: (2025)
by: Yang, Yibo
Published: (2025)
Do Reasoning Models Enhance Embedding Models?
by: Chan, Wun Yu, et al.
Published: (2026)
by: Chan, Wun Yu, et al.
Published: (2026)
Research on a hybrid LSTM-CNN-Attention model for text-based web content classification
by: Kuz, Mykola, et al.
Published: (2025)
by: Kuz, Mykola, et al.
Published: (2025)
Linguistic Collapse: Neural Collapse in (Large) Language Models
by: Wu, Robert, et al.
Published: (2024)
by: Wu, Robert, et al.
Published: (2024)
The Concept Allocation Zone: Tracking How Concepts Form Across Transformer Depth
by: Henry, James
Published: (2026)
by: Henry, James
Published: (2026)
Latent Object Permanence: Topological Phase Transitions, Free-Energy Principles, and Renormalization Group Flows in Deep Transformer Manifolds
by: Alpay, Faruk, et al.
Published: (2026)
by: Alpay, Faruk, et al.
Published: (2026)
Theoretical Analysis of Positional Encodings in Transformer Models: Impact on Expressiveness and Generalization
by: Li, Yin
Published: (2025)
by: Li, Yin
Published: (2025)
Understanding the Uncertainty of LLM Explanations: A Perspective Based on Reasoning Topology
by: Da, Longchao, et al.
Published: (2025)
by: Da, Longchao, et al.
Published: (2025)
Position: Uncertainty Quantification in LLMs is Just Unsupervised Clustering
by: Chen, Tiejin, et al.
Published: (2026)
by: Chen, Tiejin, et al.
Published: (2026)
Causal Dimensionality of Transformer Representations: Measurement, Scaling, and Layer Structure
by: Sarkar, Nilesh, et al.
Published: (2026)
by: Sarkar, Nilesh, et al.
Published: (2026)
CogniLoad: A Synthetic Natural Language Reasoning Benchmark With Tunable Length, Intrinsic Difficulty, and Distractor Density
by: Kaiser, Daniel, et al.
Published: (2025)
by: Kaiser, Daniel, et al.
Published: (2025)
How Pruning Reshapes Features: Sparse Autoencoder Analysis of Weight-Pruned Language Models
by: Borobia, Hector, et al.
Published: (2026)
by: Borobia, Hector, et al.
Published: (2026)
AIPsy-Affect: A Keyword-Free Clinical Stimulus Battery for Mechanistic Interpretability of Emotion in Language Models
by: Keeman, Michael
Published: (2026)
by: Keeman, Michael
Published: (2026)
Product-of-Experts Training Reduces Dataset Artifacts in Natural Language Inference
by: Mathew, Aby Mammen
Published: (2026)
by: Mathew, Aby Mammen
Published: (2026)
Mechanistic Analysis of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning
by: Imanov, Olaf Yunus Laitinen
Published: (2026)
by: Imanov, Olaf Yunus Laitinen
Published: (2026)
Inference acceleration for large language models using "stairs" assisted greedy generation
by: Grigaliūnas, Domas, et al.
Published: (2024)
by: Grigaliūnas, Domas, et al.
Published: (2024)
Strategic Doctrine Language Models (sdLM): A Learning-System Framework for Doctrinal Consistency and Geopolitical Forecasting
by: Imanov, Olaf Yunus Laitinen, et al.
Published: (2026)
by: Imanov, Olaf Yunus Laitinen, et al.
Published: (2026)
SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
Unpacking Hateful Memes: Presupposed Context and False Claims
by: Cai, Weibin, et al.
Published: (2025)
by: Cai, Weibin, et al.
Published: (2025)
Tricks and Plug-ins for Gradient Boosting with Transformers
by: Fang, Biyi, et al.
Published: (2025)
by: Fang, Biyi, et al.
Published: (2025)
Beyond Subtokens: A Rich Character Embedding for Low-resource and Morphologically Complex Languages
by: Schneider, Felix, et al.
Published: (2026)
by: Schneider, Felix, et al.
Published: (2026)
ProactBench: Beyond What The User Asked For
by: Harfi, Sepehr, et al.
Published: (2026)
by: Harfi, Sepehr, et al.
Published: (2026)
ReFactor GNNs: Revisiting Factorisation-based Models from a Message-Passing Perspective
by: Chen, Yihong, et al.
Published: (2022)
by: Chen, Yihong, et al.
Published: (2022)
Vis-CoT: A Human-in-the-Loop Framework for Interactive Visualization and Intervention in LLM Chain-of-Thought Reasoning
by: Pather, Kaviraj, et al.
Published: (2025)
by: Pather, Kaviraj, et al.
Published: (2025)
Approaches to Semantic Textual Similarity in Slovak Language: From Algorithms to Transformers
by: Radosky, Lukas, et al.
Published: (2026)
by: Radosky, Lukas, et al.
Published: (2026)
Extracting Sentence Embeddings from Pretrained Transformer Models
by: Stankevičius, Lukas, et al.
Published: (2024)
by: Stankevičius, Lukas, et al.
Published: (2024)
Diagnosing Multi-step Reasoning Failures in Black-box LLMs via Stepwise Confidence Attribution
by: Liu, Xiaoou, et al.
Published: (2026)
by: Liu, Xiaoou, et al.
Published: (2026)
$δ$-STEAL: LLM Stealing Attack with Local Differential Privacy
by: Dang, Kieu, et al.
Published: (2025)
by: Dang, Kieu, et al.
Published: (2025)
SAGE: Sign-Adaptive Gradient for Memory-Efficient LLM Optimization
by: Lee, Wooin, et al.
Published: (2026)
by: Lee, Wooin, et al.
Published: (2026)
Rule Extraction in Machine Learning: Chat Incremental Pattern Constructor
by: Nwokocha, Caleb Princewill
Published: (2022)
by: Nwokocha, Caleb Princewill
Published: (2022)
Similar Items
-
The Inhibitor: ReLU and Addition-Based Attention for Efficient Transformers under Fully Homomorphic Encryption on the Torus
by: Brännvall, Rickard, et al.
Published: (2023) -
Graded Transformers
by: Shaska Sr, Tony
Published: (2025) -
Think Thrice Before You Speak: Dual knowledge-enhanced Theory-of-Mind Reasoning for Persuasive Agents
by: Ma, Minghui, et al.
Published: (2026) -
Modularity in Transformers: Investigating Neuron Separability & Specialization
by: Pochinkov, Nicholas, et al.
Published: (2024) -
When Does Content-Based Routing Work? Representation Requirements for Selective Attention in Hybrid Sequence Models
by: Basu, Abhinaba
Published: (2026)