Learnable Multipliers: Freeing the Scale of Language Model Matrix Layers
Fuente:
arXiv
Guardado en:
| Autores principales: | Velikanov, Maksim, Chahed, Ilyas, Zuo, Jingwei, Rhaiem, Dhia Eddine, Belkada, Younes, Hacid, Hakim |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Falcon Mamba: The First Competitive Attention-free 7B Language Model
por: Zuo, Jingwei, et al.
Publicado: (2024)
por: Zuo, Jingwei, et al.
Publicado: (2024)
Re-thinking Human Activity Recognition with Hierarchy-aware Label Relationship Modeling
por: Zuo, Jingwei, et al.
Publicado: (2024)
por: Zuo, Jingwei, et al.
Publicado: (2024)
MAGNETO: Edge AI for Human Activity Recognition -- Privacy and Personalization
por: Zuo, Jingwei, et al.
Publicado: (2024)
por: Zuo, Jingwei, et al.
Publicado: (2024)
SGD with memory: fundamental properties and stochastic acceleration
por: Yarotsky, Dmitry, et al.
Publicado: (2024)
por: Yarotsky, Dmitry, et al.
Publicado: (2024)
Tight Convergence Rate Bounds for Optimization Under Power Law Spectral Conditions
por: Velikanov, Maksim, et al.
Publicado: (2022)
por: Velikanov, Maksim, et al.
Publicado: (2022)
Generalization error of spectral algorithms
por: Velikanov, Maksim, et al.
Publicado: (2024)
por: Velikanov, Maksim, et al.
Publicado: (2024)
Constrained Online Convex Optimization with Memory and Predictions
por: Abdullah, Mohammed, et al.
Publicado: (2026)
por: Abdullah, Mohammed, et al.
Publicado: (2026)
PORT: Preference Optimization on Reasoning Traces
por: Lahlou, Salem, et al.
Publicado: (2024)
por: Lahlou, Salem, et al.
Publicado: (2024)
Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance
por: Zuo, Jingwei, et al.
Publicado: (2025)
por: Zuo, Jingwei, et al.
Publicado: (2025)
BayesJudge: Bayesian Kernel Language Modelling with Confidence Uncertainty in Legal Judgment Prediction
por: Azam, Ubaid, et al.
Publicado: (2024)
por: Azam, Ubaid, et al.
Publicado: (2024)
LEAP: Learnable End-to-End Adaptive Pruning of Large Language Models
por: Mozaffari, Mohammad, et al.
Publicado: (2026)
por: Mozaffari, Mohammad, et al.
Publicado: (2026)
WavLink: Compact Audio-Text Embeddings with a Global Whisper Token
por: Kumar, Gokul Karthik, et al.
Publicado: (2026)
por: Kumar, Gokul Karthik, et al.
Publicado: (2026)
Maximizing the Potential of Synthetic Data: Insights from Random Matrix Theory
por: Firdoussi, Aymane El, et al.
Publicado: (2024)
por: Firdoussi, Aymane El, et al.
Publicado: (2024)
Competitive Audio-Language Models with Data-Efficient Single-Stage Training on Public Data
por: Kumar, Gokul Karthik, et al.
Publicado: (2025)
por: Kumar, Gokul Karthik, et al.
Publicado: (2025)
When Kernels Multiply, Clusters Unify: Fusing Embeddings with the Kronecker Product
por: Wu, Youqi, et al.
Publicado: (2025)
por: Wu, Youqi, et al.
Publicado: (2025)
Data Quality in Edge Machine Learning: A State-of-the-Art Survey
por: Belgoumri, Mohammed Djameleddine, et al.
Publicado: (2024)
por: Belgoumri, Mohammed Djameleddine, et al.
Publicado: (2024)
LESA: Learnable LLM Layer Scaling-Up
por: Yang, Yifei, et al.
Publicado: (2025)
por: Yang, Yifei, et al.
Publicado: (2025)
From Uncertainty to Trust: Kernel Dropout for AI-Powered Medical Predictions
por: Azam, Ubaid, et al.
Publicado: (2024)
por: Azam, Ubaid, et al.
Publicado: (2024)
Enforcing Consistency and Fairness in Multi-level Hierarchical Classification with a Mask-based Output Layer
por: Chen, Shijing, et al.
Publicado: (2025)
por: Chen, Shijing, et al.
Publicado: (2025)
Towards Fully FP8 GEMM LLM Training at Scale
por: Hernández-Cano, Alejandro, et al.
Publicado: (2025)
por: Hernández-Cano, Alejandro, et al.
Publicado: (2025)
PATCH: Learnable Tile-level Hybrid Sparsity for LLMs
por: Hourri, Younes, et al.
Publicado: (2025)
por: Hourri, Younes, et al.
Publicado: (2025)
Training Machine Learning models at the Edge: A Survey
por: Khouas, Aymen Rayane, et al.
Publicado: (2024)
por: Khouas, Aymen Rayane, et al.
Publicado: (2024)
Gaussian Approximation and Multiplier Bootstrap for Federated Linear Stochastic Approximation
por: Levin, Ilya, et al.
Publicado: (2026)
por: Levin, Ilya, et al.
Publicado: (2026)
Far From Sight, Far From Mind: Inverse Distance Weighting for Graph Federated Recommendation
por: Khouas, Aymen Rayane, et al.
Publicado: (2025)
por: Khouas, Aymen Rayane, et al.
Publicado: (2025)
Alignment with Preference Optimization Is All You Need for LLM Safety
por: Alami, Reda, et al.
Publicado: (2024)
por: Alami, Reda, et al.
Publicado: (2024)
Rolling Ball Optimizer: Learning by ironing out loss landscape wrinkles
por: Belgoumri, Mohammed Djameleddine, et al.
Publicado: (2025)
por: Belgoumri, Mohammed Djameleddine, et al.
Publicado: (2025)
A Linearized Alternating Direction Multiplier Method for Federated Matrix Completion Problems
por: Hytla, Patrick, et al.
Publicado: (2025)
por: Hytla, Patrick, et al.
Publicado: (2025)
Accurate and Diverse LLM Mathematical Reasoning via Automated PRM-Guided GFlowNets
por: Younsi, Adam, et al.
Publicado: (2025)
por: Younsi, Adam, et al.
Publicado: (2025)
Learnable Similarity and Dissimilarity Guided Symmetric Non-Negative Matrix Factorization
por: Lyu, Wenlong, et al.
Publicado: (2024)
por: Lyu, Wenlong, et al.
Publicado: (2024)
Alternating Direction Method of Multipliers for Nonlinear Matrix Decompositions
por: Awari, Atharva, et al.
Publicado: (2025)
por: Awari, Atharva, et al.
Publicado: (2025)
On the Learnability of Watermarks for Language Models
por: Gu, Chenchen, et al.
Publicado: (2023)
por: Gu, Chenchen, et al.
Publicado: (2023)
Efficient Conformal Prediction under Data Heterogeneity
por: Plassier, Vincent, et al.
Publicado: (2023)
por: Plassier, Vincent, et al.
Publicado: (2023)
Scaling Embedding Layers in Language Models
por: Yu, Da, et al.
Publicado: (2025)
por: Yu, Da, et al.
Publicado: (2025)
Auto-FlexSwitch: Efficient Dynamic Model Merging via Learnable Task Vector Compression
por: Gao, Junqi, et al.
Publicado: (2026)
por: Gao, Junqi, et al.
Publicado: (2026)
Scale-Sensitive Shattering: Learnability and Evaluability at Optimal Scale
por: Aiyer, Shashaank, et al.
Publicado: (2026)
por: Aiyer, Shashaank, et al.
Publicado: (2026)
MSPT: Efficient Large-Scale Physical Modeling via Parallelized Multi-Scale Attention
por: Curvo, Pedro M. P., et al.
Publicado: (2025)
por: Curvo, Pedro M. P., et al.
Publicado: (2025)
Bit-Accurate Modeling of GPU Matrix Multiply-Accumulate Units: Demystifying Numerical Discrepancy and Accuracy
por: Xie, Peichen, et al.
Publicado: (2025)
por: Xie, Peichen, et al.
Publicado: (2025)
AKD : Adversarial Knowledge Distillation For Large Language Models Alignment on Coding tasks
por: Oulkadda, Ilyas, et al.
Publicado: (2025)
por: Oulkadda, Ilyas, et al.
Publicado: (2025)
Sparse Layers are Critical to Scaling Looped Language Models
por: Lee, Ryan, et al.
Publicado: (2026)
por: Lee, Ryan, et al.
Publicado: (2026)
STRIDE: Learnable Stepwise Language Feedback for LLM Reasoning
por: Zhang, Junjie, et al.
Publicado: (2026)
por: Zhang, Junjie, et al.
Publicado: (2026)
Ejemplares similares
-
Falcon Mamba: The First Competitive Attention-free 7B Language Model
por: Zuo, Jingwei, et al.
Publicado: (2024) -
Re-thinking Human Activity Recognition with Hierarchy-aware Label Relationship Modeling
por: Zuo, Jingwei, et al.
Publicado: (2024) -
MAGNETO: Edge AI for Human Activity Recognition -- Privacy and Personalization
por: Zuo, Jingwei, et al.
Publicado: (2024) -
SGD with memory: fundamental properties and stochastic acceleration
por: Yarotsky, Dmitry, et al.
Publicado: (2024) -
Tight Convergence Rate Bounds for Optimization Under Power Law Spectral Conditions
por: Velikanov, Maksim, et al.
Publicado: (2022)