Learnable Multipliers: Freeing the Scale of Language Model Matrix Layers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Velikanov, Maksim, Chahed, Ilyas, Zuo, Jingwei, Rhaiem, Dhia Eddine, Belkada, Younes, Hacid, Hakim |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Falcon Mamba: The First Competitive Attention-free 7B Language Model
von: Zuo, Jingwei, et al.
Veröffentlicht: (2024)
von: Zuo, Jingwei, et al.
Veröffentlicht: (2024)
Re-thinking Human Activity Recognition with Hierarchy-aware Label Relationship Modeling
von: Zuo, Jingwei, et al.
Veröffentlicht: (2024)
von: Zuo, Jingwei, et al.
Veröffentlicht: (2024)
MAGNETO: Edge AI for Human Activity Recognition -- Privacy and Personalization
von: Zuo, Jingwei, et al.
Veröffentlicht: (2024)
von: Zuo, Jingwei, et al.
Veröffentlicht: (2024)
SGD with memory: fundamental properties and stochastic acceleration
von: Yarotsky, Dmitry, et al.
Veröffentlicht: (2024)
von: Yarotsky, Dmitry, et al.
Veröffentlicht: (2024)
Tight Convergence Rate Bounds for Optimization Under Power Law Spectral Conditions
von: Velikanov, Maksim, et al.
Veröffentlicht: (2022)
von: Velikanov, Maksim, et al.
Veröffentlicht: (2022)
Generalization error of spectral algorithms
von: Velikanov, Maksim, et al.
Veröffentlicht: (2024)
von: Velikanov, Maksim, et al.
Veröffentlicht: (2024)
Constrained Online Convex Optimization with Memory and Predictions
von: Abdullah, Mohammed, et al.
Veröffentlicht: (2026)
von: Abdullah, Mohammed, et al.
Veröffentlicht: (2026)
PORT: Preference Optimization on Reasoning Traces
von: Lahlou, Salem, et al.
Veröffentlicht: (2024)
von: Lahlou, Salem, et al.
Veröffentlicht: (2024)
Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance
von: Zuo, Jingwei, et al.
Veröffentlicht: (2025)
von: Zuo, Jingwei, et al.
Veröffentlicht: (2025)
BayesJudge: Bayesian Kernel Language Modelling with Confidence Uncertainty in Legal Judgment Prediction
von: Azam, Ubaid, et al.
Veröffentlicht: (2024)
von: Azam, Ubaid, et al.
Veröffentlicht: (2024)
LEAP: Learnable End-to-End Adaptive Pruning of Large Language Models
von: Mozaffari, Mohammad, et al.
Veröffentlicht: (2026)
von: Mozaffari, Mohammad, et al.
Veröffentlicht: (2026)
WavLink: Compact Audio-Text Embeddings with a Global Whisper Token
von: Kumar, Gokul Karthik, et al.
Veröffentlicht: (2026)
von: Kumar, Gokul Karthik, et al.
Veröffentlicht: (2026)
Maximizing the Potential of Synthetic Data: Insights from Random Matrix Theory
von: Firdoussi, Aymane El, et al.
Veröffentlicht: (2024)
von: Firdoussi, Aymane El, et al.
Veröffentlicht: (2024)
Competitive Audio-Language Models with Data-Efficient Single-Stage Training on Public Data
von: Kumar, Gokul Karthik, et al.
Veröffentlicht: (2025)
von: Kumar, Gokul Karthik, et al.
Veröffentlicht: (2025)
When Kernels Multiply, Clusters Unify: Fusing Embeddings with the Kronecker Product
von: Wu, Youqi, et al.
Veröffentlicht: (2025)
von: Wu, Youqi, et al.
Veröffentlicht: (2025)
Data Quality in Edge Machine Learning: A State-of-the-Art Survey
von: Belgoumri, Mohammed Djameleddine, et al.
Veröffentlicht: (2024)
von: Belgoumri, Mohammed Djameleddine, et al.
Veröffentlicht: (2024)
LESA: Learnable LLM Layer Scaling-Up
von: Yang, Yifei, et al.
Veröffentlicht: (2025)
von: Yang, Yifei, et al.
Veröffentlicht: (2025)
From Uncertainty to Trust: Kernel Dropout for AI-Powered Medical Predictions
von: Azam, Ubaid, et al.
Veröffentlicht: (2024)
von: Azam, Ubaid, et al.
Veröffentlicht: (2024)
Enforcing Consistency and Fairness in Multi-level Hierarchical Classification with a Mask-based Output Layer
von: Chen, Shijing, et al.
Veröffentlicht: (2025)
von: Chen, Shijing, et al.
Veröffentlicht: (2025)
Towards Fully FP8 GEMM LLM Training at Scale
von: Hernández-Cano, Alejandro, et al.
Veröffentlicht: (2025)
von: Hernández-Cano, Alejandro, et al.
Veröffentlicht: (2025)
PATCH: Learnable Tile-level Hybrid Sparsity for LLMs
von: Hourri, Younes, et al.
Veröffentlicht: (2025)
von: Hourri, Younes, et al.
Veröffentlicht: (2025)
Training Machine Learning models at the Edge: A Survey
von: Khouas, Aymen Rayane, et al.
Veröffentlicht: (2024)
von: Khouas, Aymen Rayane, et al.
Veröffentlicht: (2024)
Gaussian Approximation and Multiplier Bootstrap for Federated Linear Stochastic Approximation
von: Levin, Ilya, et al.
Veröffentlicht: (2026)
von: Levin, Ilya, et al.
Veröffentlicht: (2026)
Far From Sight, Far From Mind: Inverse Distance Weighting for Graph Federated Recommendation
von: Khouas, Aymen Rayane, et al.
Veröffentlicht: (2025)
von: Khouas, Aymen Rayane, et al.
Veröffentlicht: (2025)
Alignment with Preference Optimization Is All You Need for LLM Safety
von: Alami, Reda, et al.
Veröffentlicht: (2024)
von: Alami, Reda, et al.
Veröffentlicht: (2024)
Rolling Ball Optimizer: Learning by ironing out loss landscape wrinkles
von: Belgoumri, Mohammed Djameleddine, et al.
Veröffentlicht: (2025)
von: Belgoumri, Mohammed Djameleddine, et al.
Veröffentlicht: (2025)
A Linearized Alternating Direction Multiplier Method for Federated Matrix Completion Problems
von: Hytla, Patrick, et al.
Veröffentlicht: (2025)
von: Hytla, Patrick, et al.
Veröffentlicht: (2025)
Accurate and Diverse LLM Mathematical Reasoning via Automated PRM-Guided GFlowNets
von: Younsi, Adam, et al.
Veröffentlicht: (2025)
von: Younsi, Adam, et al.
Veröffentlicht: (2025)
Learnable Similarity and Dissimilarity Guided Symmetric Non-Negative Matrix Factorization
von: Lyu, Wenlong, et al.
Veröffentlicht: (2024)
von: Lyu, Wenlong, et al.
Veröffentlicht: (2024)
Alternating Direction Method of Multipliers for Nonlinear Matrix Decompositions
von: Awari, Atharva, et al.
Veröffentlicht: (2025)
von: Awari, Atharva, et al.
Veröffentlicht: (2025)
On the Learnability of Watermarks for Language Models
von: Gu, Chenchen, et al.
Veröffentlicht: (2023)
von: Gu, Chenchen, et al.
Veröffentlicht: (2023)
Efficient Conformal Prediction under Data Heterogeneity
von: Plassier, Vincent, et al.
Veröffentlicht: (2023)
von: Plassier, Vincent, et al.
Veröffentlicht: (2023)
Scaling Embedding Layers in Language Models
von: Yu, Da, et al.
Veröffentlicht: (2025)
von: Yu, Da, et al.
Veröffentlicht: (2025)
Auto-FlexSwitch: Efficient Dynamic Model Merging via Learnable Task Vector Compression
von: Gao, Junqi, et al.
Veröffentlicht: (2026)
von: Gao, Junqi, et al.
Veröffentlicht: (2026)
Scale-Sensitive Shattering: Learnability and Evaluability at Optimal Scale
von: Aiyer, Shashaank, et al.
Veröffentlicht: (2026)
von: Aiyer, Shashaank, et al.
Veröffentlicht: (2026)
MSPT: Efficient Large-Scale Physical Modeling via Parallelized Multi-Scale Attention
von: Curvo, Pedro M. P., et al.
Veröffentlicht: (2025)
von: Curvo, Pedro M. P., et al.
Veröffentlicht: (2025)
Bit-Accurate Modeling of GPU Matrix Multiply-Accumulate Units: Demystifying Numerical Discrepancy and Accuracy
von: Xie, Peichen, et al.
Veröffentlicht: (2025)
von: Xie, Peichen, et al.
Veröffentlicht: (2025)
AKD : Adversarial Knowledge Distillation For Large Language Models Alignment on Coding tasks
von: Oulkadda, Ilyas, et al.
Veröffentlicht: (2025)
von: Oulkadda, Ilyas, et al.
Veröffentlicht: (2025)
Sparse Layers are Critical to Scaling Looped Language Models
von: Lee, Ryan, et al.
Veröffentlicht: (2026)
von: Lee, Ryan, et al.
Veröffentlicht: (2026)
STRIDE: Learnable Stepwise Language Feedback for LLM Reasoning
von: Zhang, Junjie, et al.
Veröffentlicht: (2026)
von: Zhang, Junjie, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Falcon Mamba: The First Competitive Attention-free 7B Language Model
von: Zuo, Jingwei, et al.
Veröffentlicht: (2024) -
Re-thinking Human Activity Recognition with Hierarchy-aware Label Relationship Modeling
von: Zuo, Jingwei, et al.
Veröffentlicht: (2024) -
MAGNETO: Edge AI for Human Activity Recognition -- Privacy and Personalization
von: Zuo, Jingwei, et al.
Veröffentlicht: (2024) -
SGD with memory: fundamental properties and stochastic acceleration
von: Yarotsky, Dmitry, et al.
Veröffentlicht: (2024) -
Tight Convergence Rate Bounds for Optimization Under Power Law Spectral Conditions
von: Velikanov, Maksim, et al.
Veröffentlicht: (2022)