Attention and Compression is all you need for Controllably Efficient Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Prakash, Jatin, Puli, Aahlad, Ranganath, Rajesh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Explanations that reveal all through the definition of encoding
by: Puli, Aahlad, et al.
Published: (2024)
by: Puli, Aahlad, et al.
Published: (2024)
Contrasting with Symile: Simple Model-Agnostic Representation Learning for Unlimited Modalities
by: Saporta, Adriel, et al.
Published: (2024)
by: Saporta, Adriel, et al.
Published: (2024)
Nuisances via Negativa: Adjusting for Spurious Correlations via Data Augmentation
by: Puli, Aahlad, et al.
Published: (2022)
by: Puli, Aahlad, et al.
Published: (2022)
New-Onset Diabetes Assessment Using Artificial Intelligence-Enhanced Electrocardiography
by: Zhang, Hao, et al.
Published: (2022)
by: Zhang, Hao, et al.
Published: (2022)
Robust Anomaly Detection for Particle Physics Using Multi-Background Representation Learning
by: Gandrakota, Abhijith, et al.
Published: (2024)
by: Gandrakota, Abhijith, et al.
Published: (2024)
KL-Regularized Reinforcement Learning is Designed to Mode Collapse
by: GX-Chen, Anthony, et al.
Published: (2025)
by: GX-Chen, Anthony, et al.
Published: (2025)
Black Box Causal Inference: Effect Estimation via Meta Prediction
by: Bynum, Lucius E. J., et al.
Published: (2025)
by: Bynum, Lucius E. J., et al.
Published: (2025)
Let the Experts Speak: Improving Survival Prediction & Calibration via Mixture-of-Experts Heads
by: Morrill, Todd, et al.
Published: (2025)
by: Morrill, Todd, et al.
Published: (2025)
Large Language Models aren't all that you need
by: Holla, Kiran Voderhobli, et al.
Published: (2024)
by: Holla, Kiran Voderhobli, et al.
Published: (2024)
One protein is all you need
by: Bushuiev, Anton, et al.
Published: (2024)
by: Bushuiev, Anton, et al.
Published: (2024)
Experts are all you need: A Composable Framework for Large Language Model Inference
by: Sridharan, Shrihari, et al.
Published: (2025)
by: Sridharan, Shrihari, et al.
Published: (2025)
Addition is almost all you need: Compressing large language models with double binary factorization
by: Boža, Vladimír, et al.
Published: (2025)
by: Boža, Vladimír, et al.
Published: (2025)
Attention is all you need for boosting graph convolutional neural network
by: Wu, Yinwei
Published: (2024)
by: Wu, Yinwei
Published: (2024)
Kolmogorov GAM Networks are all you need!
by: Polson, Sarah, et al.
Published: (2025)
by: Polson, Sarah, et al.
Published: (2025)
KV-weights are all you need for skipless transformers
by: Graef, Nils
Published: (2024)
by: Graef, Nils
Published: (2024)
To Use or not to Use Muon: How Simplicity Bias in Optimizers Matters
by: Dragutinović, Sara, et al.
Published: (2026)
by: Dragutinović, Sara, et al.
Published: (2026)
Tabular Data: Is Deep Learning all you need?
by: Zabërgja, Guri, et al.
Published: (2024)
by: Zabërgja, Guri, et al.
Published: (2024)
Image compositing is all you need for data augmentation
by: Shermaine, Ang Jia Ning, et al.
Published: (2025)
by: Shermaine, Ang Jia Ning, et al.
Published: (2025)
Preference learning made easy: Everything should be understood through win rate
by: Zhang, Lily H., et al.
Published: (2025)
by: Zhang, Lily H., et al.
Published: (2025)
What's the score? Automated Denoising Score Matching for Nonlinear Diffusions
by: Singhal, Raghav, et al.
Published: (2024)
by: Singhal, Raghav, et al.
Published: (2024)
What Can You Do When You Have Zero Rewards During RL?
by: Prakash, Jatin, et al.
Published: (2025)
by: Prakash, Jatin, et al.
Published: (2025)
Attention is all you need for an improved CNN-based flash flood susceptibility modeling. The case of the ungauged Rheraya watershed, Morocco
by: Elghouat, Akram, et al.
Published: (2024)
by: Elghouat, Akram, et al.
Published: (2024)
Graph is all you need? Lightweight data-agnostic neural architecture search without training
by: Huang, Zhenhan, et al.
Published: (2024)
by: Huang, Zhenhan, et al.
Published: (2024)
Cross-Modal Safety Alignment: Is textual unlearning all you need?
by: Chakraborty, Trishna, et al.
Published: (2024)
by: Chakraborty, Trishna, et al.
Published: (2024)
Linear attention is (maybe) all you need (to understand transformer optimization)
by: Ahn, Kwangjun, et al.
Published: (2023)
by: Ahn, Kwangjun, et al.
Published: (2023)
Signformer is all you need: Towards Edge AI for Sign Language
by: Yang, Eta
Published: (2024)
by: Yang, Eta
Published: (2024)
Navigating LLM Valley: From AdamW to Memory-Efficient and Matrix-Based Optimizers
by: Ranganath, Aditya
Published: (2026)
by: Ranganath, Aditya
Published: (2026)
Simulation-based inference with scattering representations: scattering is all you need
by: Lin, Kiyam, et al.
Published: (2024)
by: Lin, Kiyam, et al.
Published: (2024)
DC is all you need: describing ReLU from a signal processing standpoint
by: Kechris, Christodoulos, et al.
Published: (2024)
by: Kechris, Christodoulos, et al.
Published: (2024)
Few Labels are all you need: A Weakly Supervised Framework for Appliance Localization in Smart-Meter Series
by: Petralia, Adrien, et al.
Published: (2025)
by: Petralia, Adrien, et al.
Published: (2025)
Slim attention: cut your context memory in half without loss -- K-cache is all you need for MHA
by: Graef, Nils, et al.
Published: (2025)
by: Graef, Nils, et al.
Published: (2025)
Three Forms of Stochastic Injection for Improved Distribution-to-Distribution Generative Modeling
by: Su, Shiye, et al.
Published: (2025)
by: Su, Shiye, et al.
Published: (2025)
Is attention all you need in medical image analysis? A review
by: Papanastasiou, Giorgos, et al.
Published: (2023)
by: Papanastasiou, Giorgos, et al.
Published: (2023)
Estimating Tail Risks in Language Model Output Distributions
by: Angell, Rico, et al.
Published: (2026)
by: Angell, Rico, et al.
Published: (2026)
Why you don't overfit, and don't need Bayes if you only train for one epoch
by: Aitchison, Laurence
Published: (2024)
by: Aitchison, Laurence
Published: (2024)
Neural Operator: Is data all you need to model the world? An insight into the paradigm of data-driven scientific ML
by: Viswanath, Hrishikesh, et al.
Published: (2023)
by: Viswanath, Hrishikesh, et al.
Published: (2023)
Compression is all you need: Modeling Mathematics
by: Aksenov, Vitaly, et al.
Published: (2026)
by: Aksenov, Vitaly, et al.
Published: (2026)
Enhancing Trust in Large Language Models via Uncertainty-Calibrated Fine-Tuning
by: Krishnan, Ranganath, et al.
Published: (2024)
by: Krishnan, Ranganath, et al.
Published: (2024)
Stochastic interpolants with data-dependent couplings
by: Albergo, Michael S., et al.
Published: (2023)
by: Albergo, Michael S., et al.
Published: (2023)
100 instances is all you need: predicting the success of a new LLM on unseen data by testing on a few instances
by: Pacchiardi, Lorenzo, et al.
Published: (2024)
by: Pacchiardi, Lorenzo, et al.
Published: (2024)
Similar Items
-
Explanations that reveal all through the definition of encoding
by: Puli, Aahlad, et al.
Published: (2024) -
Contrasting with Symile: Simple Model-Agnostic Representation Learning for Unlimited Modalities
by: Saporta, Adriel, et al.
Published: (2024) -
Nuisances via Negativa: Adjusting for Spurious Correlations via Data Augmentation
by: Puli, Aahlad, et al.
Published: (2022) -
New-Onset Diabetes Assessment Using Artificial Intelligence-Enhanced Electrocardiography
by: Zhang, Hao, et al.
Published: (2022) -
Robust Anomaly Detection for Particle Physics Using Multi-Background Representation Learning
by: Gandrakota, Abhijith, et al.
Published: (2024)