Gespeichert in:
| Hauptverfasser: | Novikov, Georgii, Oseledets, Ivan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2407.15545 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Tensor-Train Point Cloud Compression and Efficient Approximate Nearest-Neighbor Search
von: Novikov, Georgii, et al.
Veröffentlicht: (2024)
von: Novikov, Georgii, et al.
Veröffentlicht: (2024)
Quasi-Random Physics-informed Neural Networks
von: Yu, Tianchi, et al.
Veröffentlicht: (2025)
von: Yu, Tianchi, et al.
Veröffentlicht: (2025)
Spectral Informed Neural Network: An Efficient and Low-Memory PINN
von: Yu, Tianchi, et al.
Veröffentlicht: (2024)
von: Yu, Tianchi, et al.
Veröffentlicht: (2024)
RECE: Reduced Cross-Entropy Loss for Large-Catalogue Sequential Recommenders
von: Gusak, Danil, et al.
Veröffentlicht: (2024)
von: Gusak, Danil, et al.
Veröffentlicht: (2024)
Diagonal Batching Unlocks Parallelism in Recurrent Memory Transformers for Long Contexts
von: Sivtsov, Danil, et al.
Veröffentlicht: (2025)
von: Sivtsov, Danil, et al.
Veröffentlicht: (2025)
Linearly Constrained Weights: Reducing Activation Shift for Faster Training of Neural Networks
von: Kutsuna, Takuro
Veröffentlicht: (2024)
von: Kutsuna, Takuro
Veröffentlicht: (2024)
On the Spatial Structure of Mixture-of-Experts in Transformers
von: Bershatsky, Daniel, et al.
Veröffentlicht: (2025)
von: Bershatsky, Daniel, et al.
Veröffentlicht: (2025)
Exploring the Hidden Capacity of LLMs for One-Step Text Generation
von: Mezentsev, Gleb, et al.
Veröffentlicht: (2025)
von: Mezentsev, Gleb, et al.
Veröffentlicht: (2025)
Binding threshold units with artificial oscillatory neurons
von: Fanaskov, Vladimir, et al.
Veröffentlicht: (2025)
von: Fanaskov, Vladimir, et al.
Veröffentlicht: (2025)
MLPMoE: Zero-Shot Architectural Metamorphosis of Dense LLM MLPs into Static Mixture-of-Experts
von: Novikov, Ivan
Veröffentlicht: (2025)
von: Novikov, Ivan
Veröffentlicht: (2025)
Faster Language Models with Better Multi-Token Prediction Using Tensor Decomposition
von: Basharin, Artem, et al.
Veröffentlicht: (2024)
von: Basharin, Artem, et al.
Veröffentlicht: (2024)
Run LoRA Run: Faster and Lighter LoRA Implementations
von: Cherniuk, Daria, et al.
Veröffentlicht: (2023)
von: Cherniuk, Daria, et al.
Veröffentlicht: (2023)
FreshGNN: Reducing Memory Access via Stable Historical Embeddings for Graph Neural Network Training
von: Huang, Kezhao, et al.
Veröffentlicht: (2023)
von: Huang, Kezhao, et al.
Veröffentlicht: (2023)
The Rogue Scalpel: Activation Steering Compromises LLM Safety
von: Korznikov, Anton, et al.
Veröffentlicht: (2025)
von: Korznikov, Anton, et al.
Veröffentlicht: (2025)
Monitoring Neural Training with Topology: A Footprint-Predictable Collapse Index
von: Kalinowski, Alexander
Veröffentlicht: (2026)
von: Kalinowski, Alexander
Veröffentlicht: (2026)
Learning from Linear Algebra: A Graph Neural Network Approach to Preconditioner Design for Conjugate Gradient Solvers
von: Trifonov, Vladislav, et al.
Veröffentlicht: (2024)
von: Trifonov, Vladislav, et al.
Veröffentlicht: (2024)
Spectral Analysis of the Weighted Frobenius Objective
von: Trifonov, Vladislav, et al.
Veröffentlicht: (2025)
von: Trifonov, Vladislav, et al.
Veröffentlicht: (2025)
Bayesian Inverse Problems Meet Flow Matching: Efficient and Flexible Inference via Transformers
von: Sherki, Daniil, et al.
Veröffentlicht: (2025)
von: Sherki, Daniil, et al.
Veröffentlicht: (2025)
Memory Faults in Activation-sparse Quantized Deep Neural Networks: Analysis and Mitigation using Sharpness-aware Training
von: Malhotra, Akul, et al.
Veröffentlicht: (2024)
von: Malhotra, Akul, et al.
Veröffentlicht: (2024)
Reducing Smoothness with Expressive Memory Enhanced Hierarchical Graph Neural Networks
von: Bailie, Thomas, et al.
Veröffentlicht: (2025)
von: Bailie, Thomas, et al.
Veröffentlicht: (2025)
ConDiff: A Challenging Dataset for Neural Solvers of Partial Differential Equations
von: Trifonov, Vladislav, et al.
Veröffentlicht: (2024)
von: Trifonov, Vladislav, et al.
Veröffentlicht: (2024)
Explicit Flow Matching: On The Theory of Flow Matching Algorithms with Applications
von: Ryzhakov, Gleb, et al.
Veröffentlicht: (2024)
von: Ryzhakov, Gleb, et al.
Veröffentlicht: (2024)
MaxInfo: A Training-Free Key-Frame Selection Method Using Maximum Volume for Enhanced Video Understanding
von: Li, Pengyi, et al.
Veröffentlicht: (2025)
von: Li, Pengyi, et al.
Veröffentlicht: (2025)
Message-Passing GNNs Fail to Approximate Sparse Triangular Factorizations
von: Trifonov, Vladislav, et al.
Veröffentlicht: (2025)
von: Trifonov, Vladislav, et al.
Veröffentlicht: (2025)
DNN Memory Footprint Reduction via Post-Training Intra-Layer Multi-Precision Quantization
von: Ghavami, Behnam, et al.
Veröffentlicht: (2024)
von: Ghavami, Behnam, et al.
Veröffentlicht: (2024)
Topology-based Representative Datasets to Reduce Neural Network Training Resources
von: Gonzalez-Diaz, Rocio, et al.
Veröffentlicht: (2019)
von: Gonzalez-Diaz, Rocio, et al.
Veröffentlicht: (2019)
Framework GNN-AID: Graph Neural Network Analysis Interpretation and Defense
von: Lukyanov, Kirill, et al.
Veröffentlicht: (2025)
von: Lukyanov, Kirill, et al.
Veröffentlicht: (2025)
A case study of spatiotemporal forecasting techniques for weather forecasting
von: Sofi, Shakir Showkat, et al.
Veröffentlicht: (2022)
von: Sofi, Shakir Showkat, et al.
Veröffentlicht: (2022)
Sparse and Transferable Universal Singular Vectors Attack
von: Kuvshinova, Kseniia, et al.
Veröffentlicht: (2024)
von: Kuvshinova, Kseniia, et al.
Veröffentlicht: (2024)
Inverting Non-Injective Functions with Twin Neural Network Regression
von: Wetzel, Sebastian J.
Veröffentlicht: (2026)
von: Wetzel, Sebastian J.
Veröffentlicht: (2026)
Black-Box Approximation and Optimization with Hierarchical Tucker Decomposition
von: Ryzhakov, Gleb, et al.
Veröffentlicht: (2024)
von: Ryzhakov, Gleb, et al.
Veröffentlicht: (2024)
Scalable Cross-Entropy Loss for Sequential Recommendations with Large Item Catalogs
von: Mezentsev, Gleb, et al.
Veröffentlicht: (2024)
von: Mezentsev, Gleb, et al.
Veröffentlicht: (2024)
Back to Basics: Revisiting Exploration in Reinforcement Learning for LLM Reasoning via Generative Probabilities
von: Li, Pengyi, et al.
Veröffentlicht: (2026)
von: Li, Pengyi, et al.
Veröffentlicht: (2026)
NNTile: a machine learning framework capable of training extremely large GPT language models on a single node
von: Mikhalev, Aleksandr, et al.
Veröffentlicht: (2025)
von: Mikhalev, Aleksandr, et al.
Veröffentlicht: (2025)
Training Memory in Deep Neural Networks: Mechanisms, Evidence, and Measurement Gaps
von: Sevetlidis, Vasileios, et al.
Veröffentlicht: (2026)
von: Sevetlidis, Vasileios, et al.
Veröffentlicht: (2026)
Marchuk: Efficient Global Weather Forecasting from Mid-Range to Sub-Seasonal Scales via Flow Matching
von: Kuzhamuratov, Arsen, et al.
Veröffentlicht: (2026)
von: Kuzhamuratov, Arsen, et al.
Veröffentlicht: (2026)
OASIS: Online Activation Subspace Learning for Memory-Efficient Training
von: Choudhary, Sakshi, et al.
Veröffentlicht: (2026)
von: Choudhary, Sakshi, et al.
Veröffentlicht: (2026)
FRUGAL: Memory-Efficient Optimization by Reducing State Overhead for Scalable Training
von: Zmushko, Philip, et al.
Veröffentlicht: (2024)
von: Zmushko, Philip, et al.
Veröffentlicht: (2024)
OUI as a Structural Observable: Towards an Activation-Centric View of Neural Network Training
von: Fernández-Hernández, Alberto, et al.
Veröffentlicht: (2026)
von: Fernández-Hernández, Alberto, et al.
Veröffentlicht: (2026)
Semiring Activation in Neural Networks
von: Smets, Bart M. N., et al.
Veröffentlicht: (2024)
von: Smets, Bart M. N., et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Tensor-Train Point Cloud Compression and Efficient Approximate Nearest-Neighbor Search
von: Novikov, Georgii, et al.
Veröffentlicht: (2024) -
Quasi-Random Physics-informed Neural Networks
von: Yu, Tianchi, et al.
Veröffentlicht: (2025) -
Spectral Informed Neural Network: An Efficient and Low-Memory PINN
von: Yu, Tianchi, et al.
Veröffentlicht: (2024) -
RECE: Reduced Cross-Entropy Loss for Large-Catalogue Sequential Recommenders
von: Gusak, Danil, et al.
Veröffentlicht: (2024) -
Diagonal Batching Unlocks Parallelism in Recurrent Memory Transformers for Long Contexts
von: Sivtsov, Danil, et al.
Veröffentlicht: (2025)