Salvato in:
| Autori principali: | Skorokhodov, Ivan, Burtsev, Mikhail |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2019
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/1910.03867 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Can Stationary Distributions of Scale-Invariant Neural Networks Be Described by the Thermodynamics of an Ideal Gas?
di: Sadrtdinov, Ildus, et al.
Pubblicazione: (2025)
di: Sadrtdinov, Ildus, et al.
Pubblicazione: (2025)
Associative Recurrent Memory Transformer
di: Rodkin, Ivan, et al.
Pubblicazione: (2024)
di: Rodkin, Ivan, et al.
Pubblicazione: (2024)
Cramming 1568 Tokens into a Single Vector and Back Again: Exploring the Limits of Embedding Space Capacity
di: Kuratov, Yuri, et al.
Pubblicazione: (2025)
di: Kuratov, Yuri, et al.
Pubblicazione: (2025)
SRMT: Shared Memory for Multi-agent Lifelong Pathfinding
di: Sagirova, Alsu, et al.
Pubblicazione: (2025)
di: Sagirova, Alsu, et al.
Pubblicazione: (2025)
Wikontic: Constructing Wikidata-Aligned, Ontology-Aware Knowledge Graphs with Large Language Models
di: Chepurova, Alla, et al.
Pubblicazione: (2025)
di: Chepurova, Alla, et al.
Pubblicazione: (2025)
Limitations of Normalization in Attention Mechanism
di: Mudarisov, Timur, et al.
Pubblicazione: (2025)
di: Mudarisov, Timur, et al.
Pubblicazione: (2025)
Scaling Transformer to 1M tokens and beyond with RMT
di: Bulatov, Aydar, et al.
Pubblicazione: (2023)
di: Bulatov, Aydar, et al.
Pubblicazione: (2023)
In Search of Needles in a 11M Haystack: Recurrent Memory Finds What LLMs Miss
di: Kuratov, Yuri, et al.
Pubblicazione: (2024)
di: Kuratov, Yuri, et al.
Pubblicazione: (2024)
Geometric Analysis of Token Selection in Multi-Head Attention
di: Mudarisov, Timur, et al.
Pubblicazione: (2026)
di: Mudarisov, Timur, et al.
Pubblicazione: (2026)
Contextual Bandit Optimization with Pre-Trained Neural Networks
di: Terekhov, Mikhail
Pubblicazione: (2025)
di: Terekhov, Mikhail
Pubblicazione: (2025)
Loss Spike in Training Neural Networks
di: Li, Xiaolong, et al.
Pubblicazione: (2023)
di: Li, Xiaolong, et al.
Pubblicazione: (2023)
Classification with Deep Neural Networks and Logistic Loss
di: Zhang, Zihan, et al.
Pubblicazione: (2023)
di: Zhang, Zihan, et al.
Pubblicazione: (2023)
Disentangling the Causes of Plasticity Loss in Neural Networks
di: Lyle, Clare, et al.
Pubblicazione: (2024)
di: Lyle, Clare, et al.
Pubblicazione: (2024)
Learning Elementary Cellular Automata with Transformers
di: Burtsev, Mikhail
Pubblicazione: (2024)
di: Burtsev, Mikhail
Pubblicazione: (2024)
Neural Network Plasticity and Loss Sharpness
di: Koster, Max, et al.
Pubblicazione: (2024)
di: Koster, Max, et al.
Pubblicazione: (2024)
Neural Tangent Kernel of Neural Networks with Loss Informed by Differential Operators
di: Gan, Weiye, et al.
Pubblicazione: (2025)
di: Gan, Weiye, et al.
Pubblicazione: (2025)
Beyond Memorization: Extending Reasoning Depth with Recurrence, Memory and Test-Time Compute Scaling
di: Rodkin, Ivan, et al.
Pubblicazione: (2025)
di: Rodkin, Ivan, et al.
Pubblicazione: (2025)
LossVal: Efficient Data Valuation for Neural Networks
di: Wibiral, Tim, et al.
Pubblicazione: (2024)
di: Wibiral, Tim, et al.
Pubblicazione: (2024)
Visualization and Analysis of the Loss Landscape in Graph Neural Networks
di: Moustafa, Samir, et al.
Pubblicazione: (2025)
di: Moustafa, Samir, et al.
Pubblicazione: (2025)
Loss Jump During Loss Switch in Solving PDEs with Neural Networks
di: Wang, Zhiwei, et al.
Pubblicazione: (2024)
di: Wang, Zhiwei, et al.
Pubblicazione: (2024)
AlphaFlow: Understanding and Improving MeanFlow Models
di: Zhang, Huijie, et al.
Pubblicazione: (2025)
di: Zhang, Huijie, et al.
Pubblicazione: (2025)
Z-Error Loss for Training Neural Networks
di: Godin, Guillaume
Pubblicazione: (2025)
di: Godin, Guillaume
Pubblicazione: (2025)
A Quasi-Wasserstein Loss for Learning Graph Neural Networks
di: Cheng, Minjie, et al.
Pubblicazione: (2023)
di: Cheng, Minjie, et al.
Pubblicazione: (2023)
Spectral Alignment as Predictor of Loss Explosion in Neural Network Training
di: Qiu, Haiquan, et al.
Pubblicazione: (2025)
di: Qiu, Haiquan, et al.
Pubblicazione: (2025)
Visualizing, Rethinking, and Mining the Loss Landscape of Deep Neural Networks
di: Xu, Yichu, et al.
Pubblicazione: (2024)
di: Xu, Yichu, et al.
Pubblicazione: (2024)
Loss-aware Curriculum Learning for Heterogeneous Graph Neural Networks
di: Wong, Zhen Hao, et al.
Pubblicazione: (2024)
di: Wong, Zhen Hao, et al.
Pubblicazione: (2024)
Fragmentation is Efficiently Learnable by Quantum Neural Networks
di: Mints, Mikhail, et al.
Pubblicazione: (2025)
di: Mints, Mikhail, et al.
Pubblicazione: (2025)
Q-RAG: Long Context Multi-step Retrieval via Value-based Embedder Training
di: Sorokin, Artyom, et al.
Pubblicazione: (2025)
di: Sorokin, Artyom, et al.
Pubblicazione: (2025)
Framework GNN-AID: Graph Neural Network Analysis Interpretation and Defense
di: Lukyanov, Kirill, et al.
Pubblicazione: (2025)
di: Lukyanov, Kirill, et al.
Pubblicazione: (2025)
Online Neural Networks for Change-Point Detection
di: Hushchyn, Mikhail, et al.
Pubblicazione: (2020)
di: Hushchyn, Mikhail, et al.
Pubblicazione: (2020)
Physics-Informed Neural Networks: Minimizing Residual Loss with Wide Networks and Effective Activations
di: Dashtbayaz, Nima Hosseini, et al.
Pubblicazione: (2024)
di: Dashtbayaz, Nima Hosseini, et al.
Pubblicazione: (2024)
Evaluating Loss Functions for Graph Neural Networks: Towards Pretraining and Generalization
di: Abbas, Khushnood, et al.
Pubblicazione: (2025)
di: Abbas, Khushnood, et al.
Pubblicazione: (2025)
Random Linear Projections Loss for Hyperplane-Based Optimization in Neural Networks
di: Venkatasubramanian, Shyam, et al.
Pubblicazione: (2023)
di: Venkatasubramanian, Shyam, et al.
Pubblicazione: (2023)
Convex Loss Functions for Support Vector Machines (SVMs) and Neural Networks
di: Portera, Filippo
Pubblicazione: (2026)
di: Portera, Filippo
Pubblicazione: (2026)
Per-Loss Adapters for Gradient Conflict in Physics-Informed Neural Networks
di: Kim, Bum Jun, et al.
Pubblicazione: (2026)
di: Kim, Bum Jun, et al.
Pubblicazione: (2026)
Embedding Principle in Depth for the Loss Landscape Analysis of Deep Neural Networks
di: Bai, Zhiwei, et al.
Pubblicazione: (2022)
di: Bai, Zhiwei, et al.
Pubblicazione: (2022)
Towards Migrating Neural Network Implementations
di: Daoudi, Nadia, et al.
Pubblicazione: (2025)
di: Daoudi, Nadia, et al.
Pubblicazione: (2025)
Dynamic Concepts Personalization from Single Videos
di: Abdal, Rameen, et al.
Pubblicazione: (2025)
di: Abdal, Rameen, et al.
Pubblicazione: (2025)
Loss Function Considering Dead Zone for Neural Networks
di: Inami, Koki, et al.
Pubblicazione: (2024)
di: Inami, Koki, et al.
Pubblicazione: (2024)
In Search for Architectures and Loss Functions in Multi-Objective Reinforcement Learning
di: Terekhov, Mikhail, et al.
Pubblicazione: (2024)
di: Terekhov, Mikhail, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Can Stationary Distributions of Scale-Invariant Neural Networks Be Described by the Thermodynamics of an Ideal Gas?
di: Sadrtdinov, Ildus, et al.
Pubblicazione: (2025) -
Associative Recurrent Memory Transformer
di: Rodkin, Ivan, et al.
Pubblicazione: (2024) -
Cramming 1568 Tokens into a Single Vector and Back Again: Exploring the Limits of Embedding Space Capacity
di: Kuratov, Yuri, et al.
Pubblicazione: (2025) -
SRMT: Shared Memory for Multi-agent Lifelong Pathfinding
di: Sagirova, Alsu, et al.
Pubblicazione: (2025) -
Wikontic: Constructing Wikidata-Aligned, Ontology-Aware Knowledge Graphs with Large Language Models
di: Chepurova, Alla, et al.
Pubblicazione: (2025)