Salvato in:
| Autori principali: | Sivtsov, Danil, Rodkin, Ivan, Kuzmin, Gleb, Kuratov, Yuri, Oseledets, Ivan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2506.05229 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Associative Recurrent Memory Transformer
di: Rodkin, Ivan, et al.
Pubblicazione: (2024)
di: Rodkin, Ivan, et al.
Pubblicazione: (2024)
Exploring the Hidden Capacity of LLMs for One-Step Text Generation
di: Mezentsev, Gleb, et al.
Pubblicazione: (2025)
di: Mezentsev, Gleb, et al.
Pubblicazione: (2025)
BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack
di: Kuratov, Yuri, et al.
Pubblicazione: (2024)
di: Kuratov, Yuri, et al.
Pubblicazione: (2024)
On the Spatial Structure of Mixture-of-Experts in Transformers
di: Bershatsky, Daniel, et al.
Pubblicazione: (2025)
di: Bershatsky, Daniel, et al.
Pubblicazione: (2025)
In Search of Needles in a 11M Haystack: Recurrent Memory Finds What LLMs Miss
di: Kuratov, Yuri, et al.
Pubblicazione: (2024)
di: Kuratov, Yuri, et al.
Pubblicazione: (2024)
Scaling Transformer to 1M tokens and beyond with RMT
di: Bulatov, Aydar, et al.
Pubblicazione: (2023)
di: Bulatov, Aydar, et al.
Pubblicazione: (2023)
Cluster Topology-Driven Placement of Experts Reduces Network Traffic in MoE Inference
di: Sivtsov, Danil, et al.
Pubblicazione: (2025)
di: Sivtsov, Danil, et al.
Pubblicazione: (2025)
Scalable Cross-Entropy Loss for Sequential Recommendations with Large Item Catalogs
di: Mezentsev, Gleb, et al.
Pubblicazione: (2024)
di: Mezentsev, Gleb, et al.
Pubblicazione: (2024)
RECE: Reduced Cross-Entropy Loss for Large-Catalogue Sequential Recommenders
di: Gusak, Danil, et al.
Pubblicazione: (2024)
di: Gusak, Danil, et al.
Pubblicazione: (2024)
Cramming 1568 Tokens into a Single Vector and Back Again: Exploring the Limits of Embedding Space Capacity
di: Kuratov, Yuri, et al.
Pubblicazione: (2025)
di: Kuratov, Yuri, et al.
Pubblicazione: (2025)
Beyond Memorization: Extending Reasoning Depth with Recurrence, Memory and Test-Time Compute Scaling
di: Rodkin, Ivan, et al.
Pubblicazione: (2025)
di: Rodkin, Ivan, et al.
Pubblicazione: (2025)
Wikontic: Constructing Wikidata-Aligned, Ontology-Aware Knowledge Graphs with Large Language Models
di: Chepurova, Alla, et al.
Pubblicazione: (2025)
di: Chepurova, Alla, et al.
Pubblicazione: (2025)
Back to Basics: Revisiting Exploration in Reinforcement Learning for LLM Reasoning via Generative Probabilities
di: Li, Pengyi, et al.
Pubblicazione: (2026)
di: Li, Pengyi, et al.
Pubblicazione: (2026)
Birch SGD: A Tree Graph Framework for Local and Asynchronous SGD Methods
di: Tyurin, Alexander, et al.
Pubblicazione: (2025)
di: Tyurin, Alexander, et al.
Pubblicazione: (2025)
Your Transformer is Secretly Linear
di: Razzhigaev, Anton, et al.
Pubblicazione: (2024)
di: Razzhigaev, Anton, et al.
Pubblicazione: (2024)
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
di: Li, Pengyi, et al.
Pubblicazione: (2025)
di: Li, Pengyi, et al.
Pubblicazione: (2025)
Logit-KL Flow Matching: Non-Autoregressive Text Generation via Sampling-Hybrid Inference
di: Sevriugov, Egor, et al.
Pubblicazione: (2024)
di: Sevriugov, Egor, et al.
Pubblicazione: (2024)
HMT: Hierarchical Memory Transformer for Efficient Long Context Language Processing
di: He, Zifan, et al.
Pubblicazione: (2024)
di: He, Zifan, et al.
Pubblicazione: (2024)
LoTR: Low Tensor Rank Weight Adaptation
di: Bershatsky, Daniel, et al.
Pubblicazione: (2024)
di: Bershatsky, Daniel, et al.
Pubblicazione: (2024)
The Shape of Learning: Anisotropy and Intrinsic Dimensions in Transformer-Based Models
di: Razzhigaev, Anton, et al.
Pubblicazione: (2023)
di: Razzhigaev, Anton, et al.
Pubblicazione: (2023)
Quantization of Large Language Models with an Overdetermined Basis
di: Merkulov, Daniil, et al.
Pubblicazione: (2024)
di: Merkulov, Daniil, et al.
Pubblicazione: (2024)
Inverted Activations: Reducing Memory Footprint in Neural Network Training
di: Novikov, Georgii, et al.
Pubblicazione: (2024)
di: Novikov, Georgii, et al.
Pubblicazione: (2024)
Core Context Aware Transformers for Long Context Language Modeling
di: Chen, Yaofo, et al.
Pubblicazione: (2024)
di: Chen, Yaofo, et al.
Pubblicazione: (2024)
State Stream Transformer (SST) V2: Parallel Training of Nonlinear Recurrence for Latent Space Reasoning
di: Aviss, Thea
Pubblicazione: (2026)
di: Aviss, Thea
Pubblicazione: (2026)
CART: Context-Anchored Recurrent Transformer -- A Parameter-Efficient Architecture with Learned Stability
di: Capps, Chad A.
Pubblicazione: (2026)
di: Capps, Chad A.
Pubblicazione: (2026)
Structured Recurrent Mixers for Massively Parallelized Sequence Generation
di: Badger, Benjamin L.
Pubblicazione: (2026)
di: Badger, Benjamin L.
Pubblicazione: (2026)
Batch-ICL: Effective, Efficient, and Order-Agnostic In-Context Learning
di: Zhang, Kaiyi, et al.
Pubblicazione: (2024)
di: Zhang, Kaiyi, et al.
Pubblicazione: (2024)
Toward Conversational Agents with Context and Time Sensitive Long-term Memory
di: Alonso, Nick, et al.
Pubblicazione: (2024)
di: Alonso, Nick, et al.
Pubblicazione: (2024)
Learning When to Attend: Conditional Memory Access for Long-Context LLMs
di: Choudhary, Sakshi, et al.
Pubblicazione: (2026)
di: Choudhary, Sakshi, et al.
Pubblicazione: (2026)
Evaluating Memory Structure in LLM Agents
di: Shutova, Alina, et al.
Pubblicazione: (2026)
di: Shutova, Alina, et al.
Pubblicazione: (2026)
Mitigating Copy Bias in In-Context Learning through Neuron Pruning
di: Ali, Ameen, et al.
Pubblicazione: (2024)
di: Ali, Ameen, et al.
Pubblicazione: (2024)
Latent Context Compilation: Distilling Long Context into Compact Portable Memory
di: Li, Zeju, et al.
Pubblicazione: (2026)
di: Li, Zeju, et al.
Pubblicazione: (2026)
LLM-Microscope: Uncovering the Hidden Role of Punctuation in Context Memory of Transformers
di: Razzhigaev, Anton, et al.
Pubblicazione: (2025)
di: Razzhigaev, Anton, et al.
Pubblicazione: (2025)
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference
di: Li, Weizhuo, et al.
Pubblicazione: (2024)
di: Li, Weizhuo, et al.
Pubblicazione: (2024)
Compress, Gather, and Recompute: REFORMing Long-Context Processing in Transformers
di: Song, Woomin, et al.
Pubblicazione: (2025)
di: Song, Woomin, et al.
Pubblicazione: (2025)
Short Data, Long Context: Distilling Positional Knowledge in Transformers
di: Huber, Patrick, et al.
Pubblicazione: (2026)
di: Huber, Patrick, et al.
Pubblicazione: (2026)
S$^3$-Attention:Attention-Aligned Endogenous Retrieval for Memory-Bounded Long-Context Inference
di: Ma, Qingsen, et al.
Pubblicazione: (2026)
di: Ma, Qingsen, et al.
Pubblicazione: (2026)
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
di: Xiao, Chaojun, et al.
Pubblicazione: (2024)
di: Xiao, Chaojun, et al.
Pubblicazione: (2024)
Explicit Flow Matching: On The Theory of Flow Matching Algorithms with Applications
di: Ryzhakov, Gleb, et al.
Pubblicazione: (2024)
di: Ryzhakov, Gleb, et al.
Pubblicazione: (2024)
Spectral Informed Neural Network: An Efficient and Low-Memory PINN
di: Yu, Tianchi, et al.
Pubblicazione: (2024)
di: Yu, Tianchi, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Associative Recurrent Memory Transformer
di: Rodkin, Ivan, et al.
Pubblicazione: (2024) -
Exploring the Hidden Capacity of LLMs for One-Step Text Generation
di: Mezentsev, Gleb, et al.
Pubblicazione: (2025) -
BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack
di: Kuratov, Yuri, et al.
Pubblicazione: (2024) -
On the Spatial Structure of Mixture-of-Experts in Transformers
di: Bershatsky, Daniel, et al.
Pubblicazione: (2025) -
In Search of Needles in a 11M Haystack: Recurrent Memory Finds What LLMs Miss
di: Kuratov, Yuri, et al.
Pubblicazione: (2024)