Associative Recurrent Memory Transformer
Fuente:
arXiv
Guardado en:
| Autores principales: | Rodkin, Ivan, Kuratov, Yuri, Bulatov, Aydar, Burtsev, Mikhail |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Scaling Transformer to 1M tokens and beyond with RMT
por: Bulatov, Aydar, et al.
Publicado: (2023)
por: Bulatov, Aydar, et al.
Publicado: (2023)
Wikontic: Constructing Wikidata-Aligned, Ontology-Aware Knowledge Graphs with Large Language Models
por: Chepurova, Alla, et al.
Publicado: (2025)
por: Chepurova, Alla, et al.
Publicado: (2025)
In Search of Needles in a 11M Haystack: Recurrent Memory Finds What LLMs Miss
por: Kuratov, Yuri, et al.
Publicado: (2024)
por: Kuratov, Yuri, et al.
Publicado: (2024)
Key-Value Means: Transformers with Expandable Block-Recurrent Compressed Memory
por: Goldstein, Daniel, et al.
Publicado: (2026)
por: Goldstein, Daniel, et al.
Publicado: (2026)
Graph Memory Transformer (GMT)
por: Zanarini, Nicola, et al.
Publicado: (2026)
por: Zanarini, Nicola, et al.
Publicado: (2026)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
por: Oketunji, Abiodun Finbarrs
Publicado: (2023)
por: Oketunji, Abiodun Finbarrs
Publicado: (2023)
Cramming 1568 Tokens into a Single Vector and Back Again: Exploring the Limits of Embedding Space Capacity
por: Kuratov, Yuri, et al.
Publicado: (2025)
por: Kuratov, Yuri, et al.
Publicado: (2025)
KV-Fold: One-Step KV-Cache Recurrence for Long-Context Inference
por: Nadali, Alireza, et al.
Publicado: (2026)
por: Nadali, Alireza, et al.
Publicado: (2026)
BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack
por: Kuratov, Yuri, et al.
Publicado: (2024)
por: Kuratov, Yuri, et al.
Publicado: (2024)
SRMT: Shared Memory for Multi-agent Lifelong Pathfinding
por: Sagirova, Alsu, et al.
Publicado: (2025)
por: Sagirova, Alsu, et al.
Publicado: (2025)
Large Language Model (LLM) Bias Index -- LLMBI
por: Oketunji, Abiodun Finbarrs, et al.
Publicado: (2023)
por: Oketunji, Abiodun Finbarrs, et al.
Publicado: (2023)
Learning from Supervision with Semantic and Episodic Memory: A Reflective Approach to Agent Adaptation
por: Hassell, Jackson, et al.
Publicado: (2025)
por: Hassell, Jackson, et al.
Publicado: (2025)
Assisting humans in complex comparisons: automated information comparison at scale
por: Yuen, Truman, et al.
Publicado: (2024)
por: Yuen, Truman, et al.
Publicado: (2024)
Synthius-Mem: Brain-Inspired Hallucination-Resistant Persona Memory Achieving 94.4% Memory Accuracy and 99.6% Adversarial Robustness on LoCoMo
por: Gadzhiev, Artem, et al.
Publicado: (2026)
por: Gadzhiev, Artem, et al.
Publicado: (2026)
Disentangling Direction and Magnitude in Transformer Representations: A Double Dissociation Through L2-Matched Perturbation Analysis
por: Vardhan, Mangadoddi Srikar, et al.
Publicado: (2026)
por: Vardhan, Mangadoddi Srikar, et al.
Publicado: (2026)
Enhancing Transformer RNNs with Multiple Temporal Perspectives
por: Dumitru, Razvan-Gabriel, et al.
Publicado: (2024)
por: Dumitru, Razvan-Gabriel, et al.
Publicado: (2024)
No Free Swap: Protocol-Dependent Layer Redundancy in Transformers
por: Garcia, Gabriel
Publicado: (2026)
por: Garcia, Gabriel
Publicado: (2026)
Weakly Supervised Distillation of Hallucination Signals into Transformer Representations
por: Salehmohamed, Shoaib Sadiq, et al.
Publicado: (2026)
por: Salehmohamed, Shoaib Sadiq, et al.
Publicado: (2026)
Prototype Transformer: Towards Language Model Architectures Interpretable by Design
por: Yordanov, Yordan, et al.
Publicado: (2026)
por: Yordanov, Yordan, et al.
Publicado: (2026)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
por: Fadli, Samih
Publicado: (2025)
por: Fadli, Samih
Publicado: (2025)
$\text{Memory}^3$: Language Modeling with Explicit Memory
por: Yang, Hongkang, et al.
Publicado: (2024)
por: Yang, Hongkang, et al.
Publicado: (2024)
Grokking in the Wild: Data Augmentation for Real-World Multi-Hop Reasoning with Transformers
por: Abramov, Roman, et al.
Publicado: (2025)
por: Abramov, Roman, et al.
Publicado: (2025)
ObfusQAte: A Proposed Framework to Evaluate LLM Robustness on Obfuscated Factual Question Answering
por: Ghosh, Shubhra, et al.
Publicado: (2025)
por: Ghosh, Shubhra, et al.
Publicado: (2025)
Dealing with Annotator Disagreement in Hate Speech Classification
por: Dehghan, Somaiyeh, et al.
Publicado: (2025)
por: Dehghan, Somaiyeh, et al.
Publicado: (2025)
Improving Discrete Diffusion Unmasking Policies Beyond Explicit Reference Policies
por: Hong, Chunsan, et al.
Publicado: (2025)
por: Hong, Chunsan, et al.
Publicado: (2025)
Unveiling Transformer Perception by Exploring Input Manifolds
por: Benfenati, Alessandro, et al.
Publicado: (2024)
por: Benfenati, Alessandro, et al.
Publicado: (2024)
Knowledge Graph Embeddings: A Comprehensive Survey on Capturing Relation Properties
por: Niu, Guanglin
Publicado: (2024)
por: Niu, Guanglin
Publicado: (2024)
Observations on Building RAG Systems for Technical Documents
por: Soman, Sumit, et al.
Publicado: (2024)
por: Soman, Sumit, et al.
Publicado: (2024)
Neural Multimodal Topic Modeling: A Comprehensive Evaluation
por: González-Pizarro, Felipe, et al.
Publicado: (2024)
por: González-Pizarro, Felipe, et al.
Publicado: (2024)
NLP Case Study on Predicting the Before and After of the Ukraine-Russia and Hamas-Israel Conflicts
por: Miner, Jordan, et al.
Publicado: (2024)
por: Miner, Jordan, et al.
Publicado: (2024)
SocraSynth: Multi-LLM Reasoning with Conditional Statistics
por: Chang, Edward Y.
Publicado: (2024)
por: Chang, Edward Y.
Publicado: (2024)
Uncovering Latent Human Wellbeing in Language Model Embeddings
por: Freire, Pedro, et al.
Publicado: (2024)
por: Freire, Pedro, et al.
Publicado: (2024)
Detecting Hallucinations in Large Language Model Generation: A Token Probability Approach
por: Quevedo, Ernesto, et al.
Publicado: (2024)
por: Quevedo, Ernesto, et al.
Publicado: (2024)
Mathador-LM: A Dynamic Benchmark for Mathematical Reasoning on Large Language Models
por: Kurtic, Eldar, et al.
Publicado: (2024)
por: Kurtic, Eldar, et al.
Publicado: (2024)
Enhancing In-Context Learning via Implicit Demonstration Augmentation
por: Zhou, Xiaoling, et al.
Publicado: (2024)
por: Zhou, Xiaoling, et al.
Publicado: (2024)
Adaptive Circuit Behavior and Generalization in Mechanistic Interpretability
por: Nainani, Jatin, et al.
Publicado: (2024)
por: Nainani, Jatin, et al.
Publicado: (2024)
ProSparse: Introducing and Enhancing Intrinsic Activation Sparsity within Large Language Models
por: Song, Chenyang, et al.
Publicado: (2024)
por: Song, Chenyang, et al.
Publicado: (2024)
Encoder vs Decoder: Comparative Analysis of Encoder and Decoder Language Models on Multilingual NLU Tasks
por: Nielsen, Dan Saattrup, et al.
Publicado: (2024)
por: Nielsen, Dan Saattrup, et al.
Publicado: (2024)
Uncovering Biases with Reflective Large Language Models
por: Chang, Edward Y.
Publicado: (2024)
por: Chang, Edward Y.
Publicado: (2024)
Self-Supervised Position Debiasing for Large Language Models
por: Liu, Zhongkun, et al.
Publicado: (2024)
por: Liu, Zhongkun, et al.
Publicado: (2024)
Ejemplares similares
-
Scaling Transformer to 1M tokens and beyond with RMT
por: Bulatov, Aydar, et al.
Publicado: (2023) -
Wikontic: Constructing Wikidata-Aligned, Ontology-Aware Knowledge Graphs with Large Language Models
por: Chepurova, Alla, et al.
Publicado: (2025) -
In Search of Needles in a 11M Haystack: Recurrent Memory Finds What LLMs Miss
por: Kuratov, Yuri, et al.
Publicado: (2024) -
Key-Value Means: Transformers with Expandable Block-Recurrent Compressed Memory
por: Goldstein, Daniel, et al.
Publicado: (2026) -
Graph Memory Transformer (GMT)
por: Zanarini, Nicola, et al.
Publicado: (2026)