Faster Language Models with Better Multi-Token Prediction Using Tensor Decomposition
Fuente:
arXiv
Saved in:
| Main Authors: | Basharin, Artem, Chertkov, Andrei, Oseledets, Ivan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Black-Box Approximation and Optimization with Hierarchical Tucker Decomposition
by: Ryzhakov, Gleb, et al.
Published: (2024)
by: Ryzhakov, Gleb, et al.
Published: (2024)
Low-rank surrogate modeling and stochastic zero-order optimization for training of neural networks with black-box layers
by: Chertkov, Andrei, et al.
Published: (2025)
by: Chertkov, Andrei, et al.
Published: (2025)
Tensor Train Decomposition for Adversarial Attacks on Computer Vision Models
by: Chertkov, Andrei, et al.
Published: (2023)
by: Chertkov, Andrei, et al.
Published: (2023)
Fast gradient-free activation maximization for neurons in spiking neural networks
by: Pospelov, Nikita, et al.
Published: (2023)
by: Pospelov, Nikita, et al.
Published: (2023)
Run LoRA Run: Faster and Lighter LoRA Implementations
by: Cherniuk, Daria, et al.
Published: (2023)
by: Cherniuk, Daria, et al.
Published: (2023)
Inverted Activations: Reducing Memory Footprint in Neural Network Training
by: Novikov, Georgii, et al.
Published: (2024)
by: Novikov, Georgii, et al.
Published: (2024)
Tensor-Train Point Cloud Compression and Efficient Approximate Nearest-Neighbor Search
by: Novikov, Georgii, et al.
Published: (2024)
by: Novikov, Georgii, et al.
Published: (2024)
Global Optimization of Atomic Clusters via Physically-Constrained Tensor Train Decomposition
by: Sozykin, Konstantin, et al.
Published: (2026)
by: Sozykin, Konstantin, et al.
Published: (2026)
Faster Predictive Coding Networks via Better Initialization
by: Pinchetti, Luca, et al.
Published: (2026)
by: Pinchetti, Luca, et al.
Published: (2026)
LoTR: Low Tensor Rank Weight Adaptation
by: Bershatsky, Daniel, et al.
Published: (2024)
by: Bershatsky, Daniel, et al.
Published: (2024)
High-dimensional Optimization with Low Rank Tensor Sampling and Local Search
by: Sozykin, Konstantin, et al.
Published: (2025)
by: Sozykin, Konstantin, et al.
Published: (2025)
Lighter, Better, Faster Multi-Source Domain Adaptation with Gaussian Mixture Models and Optimal Transport
by: Montesuma, Eduardo Fernandes, et al.
Published: (2024)
by: Montesuma, Eduardo Fernandes, et al.
Published: (2024)
The AdEMAMix Optimizer: Better, Faster, Older
by: Pagliardini, Matteo, et al.
Published: (2024)
by: Pagliardini, Matteo, et al.
Published: (2024)
A PID-Controlled Tensor Wheel Decomposition Model for Dynamic Link Prediction
by: Wang, Qu, et al.
Published: (2025)
by: Wang, Qu, et al.
Published: (2025)
Mean-Field Path-Integral Diffusion: From Samples to Interacting Agents
by: Chertkov, Michael
Published: (2026)
by: Chertkov, Michael
Published: (2026)
TUBE: Tangent Upper Bound on Evidence for Discrete Diffusion Language Models
by: Ivanov, Arseny, et al.
Published: (2026)
by: Ivanov, Arseny, et al.
Published: (2026)
Multi-view Graph Condensation via Tensor Decomposition
by: Santos, Nícolas Roque dos, et al.
Published: (2025)
by: Santos, Nícolas Roque dos, et al.
Published: (2025)
ANO : Faster is Better in Noisy Landscape
by: Kegreisz, Adrien
Published: (2025)
by: Kegreisz, Adrien
Published: (2025)
On the Spatial Structure of Mixture-of-Experts in Transformers
by: Bershatsky, Daniel, et al.
Published: (2025)
by: Bershatsky, Daniel, et al.
Published: (2025)
Exploring the Hidden Capacity of LLMs for One-Step Text Generation
by: Mezentsev, Gleb, et al.
Published: (2025)
by: Mezentsev, Gleb, et al.
Published: (2025)
Binding threshold units with artificial oscillatory neurons
by: Fanaskov, Vladimir, et al.
Published: (2025)
by: Fanaskov, Vladimir, et al.
Published: (2025)
Quasi-Random Physics-informed Neural Networks
by: Yu, Tianchi, et al.
Published: (2025)
by: Yu, Tianchi, et al.
Published: (2025)
Space-Time Diffusion Bridge
by: Behjoo, Hamidreza, et al.
Published: (2024)
by: Behjoo, Hamidreza, et al.
Published: (2024)
A Biased Nonnegative Block Term Tensor Decomposition Model for Dynamic QoS Prediction
by: Liu, Wenjing, et al.
Published: (2026)
by: Liu, Wenjing, et al.
Published: (2026)
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
by: Li, Pengyi, et al.
Published: (2025)
by: Li, Pengyi, et al.
Published: (2025)
No-Rank Tensor Decomposition Using Metric Learning
by: Bagherian, Maryam
Published: (2025)
by: Bagherian, Maryam
Published: (2025)
Mixing Artificial and Natural Intelligence: From Statistical Mechanics to AI and Back to Turbulence
by: Chertkov, Michael
Published: (2024)
by: Chertkov, Michael
Published: (2024)
Temporal Memory for Resource-Constrained Agents: Continual Learning via Stochastic Compress-Add-Smooth
by: Chertkov, Michael
Published: (2026)
by: Chertkov, Michael
Published: (2026)
Generative Stochastic Optimal Transport: Guided Harmonic Path-Integral Diffusion
by: Chertkov, Michael
Published: (2025)
by: Chertkov, Michael
Published: (2025)
Harmonic Path Integral Diffusion
by: Behjoo, Hamidreza, et al.
Published: (2024)
by: Behjoo, Hamidreza, et al.
Published: (2024)
Better Models, Faster Training: Sigmoid Attention for single-cell Foundation Models
by: Sadashivaiah, Vijay, et al.
Published: (2026)
by: Sadashivaiah, Vijay, et al.
Published: (2026)
Robust Spatiotemporally Contiguous Anomaly Detection Using Tensor Decomposition
by: Mondal, Rachita, et al.
Published: (2025)
by: Mondal, Rachita, et al.
Published: (2025)
A Multi-resolution Low-rank Tensor Decomposition
by: Rozada, Sergio, et al.
Published: (2024)
by: Rozada, Sergio, et al.
Published: (2024)
Spectral Analysis of the Weighted Frobenius Objective
by: Trifonov, Vladislav, et al.
Published: (2025)
by: Trifonov, Vladislav, et al.
Published: (2025)
Bayesian Inverse Problems Meet Flow Matching: Efficient and Flexible Inference via Transformers
by: Sherki, Daniil, et al.
Published: (2025)
by: Sherki, Daniil, et al.
Published: (2025)
Exact Fractional Inference via Re-Parametrization & Interpolation between Tree-Re-Weighted- and Belief Propagation- Algorithms
by: Behjoo, Hamidreza, et al.
Published: (2023)
by: Behjoo, Hamidreza, et al.
Published: (2023)
Not All Denoising Steps Are Equal: Model Scheduling for Faster Masked Diffusion Language Models
by: Sedykh, Ivan, et al.
Published: (2026)
by: Sedykh, Ivan, et al.
Published: (2026)
Tender: Accelerating Large Language Models via Tensor Decomposition and Runtime Requantization
by: Lee, Jungi, et al.
Published: (2024)
by: Lee, Jungi, et al.
Published: (2024)
Parallel Token Prediction for Language Models
by: Draxler, Felix, et al.
Published: (2025)
by: Draxler, Felix, et al.
Published: (2025)
Multi-Token Residual Prediction
by: Xu, Yufeng, et al.
Published: (2026)
by: Xu, Yufeng, et al.
Published: (2026)
Similar Items
-
Black-Box Approximation and Optimization with Hierarchical Tucker Decomposition
by: Ryzhakov, Gleb, et al.
Published: (2024) -
Low-rank surrogate modeling and stochastic zero-order optimization for training of neural networks with black-box layers
by: Chertkov, Andrei, et al.
Published: (2025) -
Tensor Train Decomposition for Adversarial Attacks on Computer Vision Models
by: Chertkov, Andrei, et al.
Published: (2023) -
Fast gradient-free activation maximization for neurons in spiking neural networks
by: Pospelov, Nikita, et al.
Published: (2023) -
Run LoRA Run: Faster and Lighter LoRA Implementations
by: Cherniuk, Daria, et al.
Published: (2023)