Cramming 1568 Tokens into a Single Vector and Back Again: Exploring the Limits of Embedding Space Capacity
Fuente:
arXiv
Salvato in:
| Autori principali: | Kuratov, Yuri, Arkhipov, Mikhail, Bulatov, Aydar, Burtsev, Mikhail |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Wikontic: Constructing Wikidata-Aligned, Ontology-Aware Knowledge Graphs with Large Language Models
di: Chepurova, Alla, et al.
Pubblicazione: (2025)
di: Chepurova, Alla, et al.
Pubblicazione: (2025)
Scaling Transformer to 1M tokens and beyond with RMT
di: Bulatov, Aydar, et al.
Pubblicazione: (2023)
di: Bulatov, Aydar, et al.
Pubblicazione: (2023)
Associative Recurrent Memory Transformer
di: Rodkin, Ivan, et al.
Pubblicazione: (2024)
di: Rodkin, Ivan, et al.
Pubblicazione: (2024)
In Search of Needles in a 11M Haystack: Recurrent Memory Finds What LLMs Miss
di: Kuratov, Yuri, et al.
Pubblicazione: (2024)
di: Kuratov, Yuri, et al.
Pubblicazione: (2024)
BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack
di: Kuratov, Yuri, et al.
Pubblicazione: (2024)
di: Kuratov, Yuri, et al.
Pubblicazione: (2024)
SRMT: Shared Memory for Multi-agent Lifelong Pathfinding
di: Sagirova, Alsu, et al.
Pubblicazione: (2025)
di: Sagirova, Alsu, et al.
Pubblicazione: (2025)
Limitations of Normalization in Attention Mechanism
di: Mudarisov, Timur, et al.
Pubblicazione: (2025)
di: Mudarisov, Timur, et al.
Pubblicazione: (2025)
Beyond Memorization: Extending Reasoning Depth with Recurrence, Memory and Test-Time Compute Scaling
di: Rodkin, Ivan, et al.
Pubblicazione: (2025)
di: Rodkin, Ivan, et al.
Pubblicazione: (2025)
Diagonal Batching Unlocks Parallelism in Recurrent Memory Transformers for Long Contexts
di: Sivtsov, Danil, et al.
Pubblicazione: (2025)
di: Sivtsov, Danil, et al.
Pubblicazione: (2025)
Learning Elementary Cellular Automata with Transformers
di: Burtsev, Mikhail
Pubblicazione: (2024)
di: Burtsev, Mikhail
Pubblicazione: (2024)
Complexity of Symbolic Representation in Working Memory of Transformer Correlates with the Complexity of a Task
di: Sagirova, Alsu, et al.
Pubblicazione: (2024)
di: Sagirova, Alsu, et al.
Pubblicazione: (2024)
Long Input Benchmark for Russian Analysis
di: Churin, Igor, et al.
Pubblicazione: (2024)
di: Churin, Igor, et al.
Pubblicazione: (2024)
Cram Less to Fit More: Training Data Pruning Improves Memorization of Facts
di: Ye, Jiayuan, et al.
Pubblicazione: (2026)
di: Ye, Jiayuan, et al.
Pubblicazione: (2026)
Loss Patterns of Neural Networks
di: Skorokhodov, Ivan, et al.
Pubblicazione: (2019)
di: Skorokhodov, Ivan, et al.
Pubblicazione: (2019)
FoNE: Precise Single-Token Number Embeddings via Fourier Features
di: Zhou, Tianyi, et al.
Pubblicazione: (2025)
di: Zhou, Tianyi, et al.
Pubblicazione: (2025)
Token Distillation: Attention-aware Input Embeddings For New Tokens
di: Dobler, Konstantin, et al.
Pubblicazione: (2025)
di: Dobler, Konstantin, et al.
Pubblicazione: (2025)
Measuring Intrinsic Dimension of Token Embeddings
di: Kataiwa, Takuya, et al.
Pubblicazione: (2025)
di: Kataiwa, Takuya, et al.
Pubblicazione: (2025)
Adaptive Retrieval Without Self-Knowledge? Bringing Uncertainty Back Home
di: Moskvoretskii, Viktor, et al.
Pubblicazione: (2025)
di: Moskvoretskii, Viktor, et al.
Pubblicazione: (2025)
Transfer of Structural Knowledge from Synthetic Languages
di: Budnikov, Mikhail, et al.
Pubblicazione: (2025)
di: Budnikov, Mikhail, et al.
Pubblicazione: (2025)
Iterative Layer-wise Distillation for Efficient Compression of Large Language Models
di: Kovalev, Grigory, et al.
Pubblicazione: (2025)
di: Kovalev, Grigory, et al.
Pubblicazione: (2025)
Beyond Early-Token Bias: Model-Specific and Language-Specific Position Effects in Multilingual LLMs
di: Menschikov, Mikhail, et al.
Pubblicazione: (2025)
di: Menschikov, Mikhail, et al.
Pubblicazione: (2025)
Attention with Trained Embeddings Provably Selects Important Tokens
di: Wu, Diyuan, et al.
Pubblicazione: (2025)
di: Wu, Diyuan, et al.
Pubblicazione: (2025)
The Geometry of Tokens in Internal Representations of Large Language Models
di: Viswanathan, Karthik, et al.
Pubblicazione: (2025)
di: Viswanathan, Karthik, et al.
Pubblicazione: (2025)
Rethinking Token Reduction for State Space Models
di: Zhan, Zheng, et al.
Pubblicazione: (2024)
di: Zhan, Zheng, et al.
Pubblicazione: (2024)
ProcrustesGPT: Compressing LLMs with Structured Matrices and Orthogonal Transformations
di: Grishina, Ekaterina, et al.
Pubblicazione: (2025)
di: Grishina, Ekaterina, et al.
Pubblicazione: (2025)
Reasoning Capacity in Multi-Agent Systems: Limitations, Challenges and Human-Centered Solutions
di: Pezeshkpour, Pouya, et al.
Pubblicazione: (2024)
di: Pezeshkpour, Pouya, et al.
Pubblicazione: (2024)
Dysarthria Normalization via Local Lie Group Transformations for Robust ASR
di: Osipov, Mikhail
Pubblicazione: (2025)
di: Osipov, Mikhail
Pubblicazione: (2025)
Large Language Models in the Task of Automatic Validation of Text Classifier Predictions
di: Tsymbalov, Aleksandr, et al.
Pubblicazione: (2025)
di: Tsymbalov, Aleksandr, et al.
Pubblicazione: (2025)
Hierarchical Token Prepending: Enhancing Information Flow in Decoder-based LLM Embeddings
di: Ding, Xueying, et al.
Pubblicazione: (2025)
di: Ding, Xueying, et al.
Pubblicazione: (2025)
Catching rationalization in the act: detecting motivated reasoning before and after CoT via activation probing
di: Mirtaheri, Parsa, et al.
Pubblicazione: (2026)
di: Mirtaheri, Parsa, et al.
Pubblicazione: (2026)
RaguTeam at SemEval-2026 Task 8: Meno and Friends in a Judge-Orchestrated LLM Ensemble for Faithful Multi-Turn Response Generation
di: Bondarenko, Ivan, et al.
Pubblicazione: (2026)
di: Bondarenko, Ivan, et al.
Pubblicazione: (2026)
Future Lens: Anticipating Subsequent Tokens from a Single Hidden State
di: Pal, Koyena, et al.
Pubblicazione: (2023)
di: Pal, Koyena, et al.
Pubblicazione: (2023)
MambaByte: Token-free Selective State Space Model
di: Wang, Junxiong, et al.
Pubblicazione: (2024)
di: Wang, Junxiong, et al.
Pubblicazione: (2024)
Analogical Reasoning Inside Large Language Models: Concept Vectors and the Limits of Abstraction
di: Opiełka, Gustaw, et al.
Pubblicazione: (2025)
di: Opiełka, Gustaw, et al.
Pubblicazione: (2025)
Understanding Token Probability Encoding in Output Embeddings
di: Cho, Hakaze, et al.
Pubblicazione: (2024)
di: Cho, Hakaze, et al.
Pubblicazione: (2024)
Prediction is not Explanation: Revisiting the Explanatory Capacity of Mapping Embeddings
di: Herasimchyk, Hanna, et al.
Pubblicazione: (2025)
di: Herasimchyk, Hanna, et al.
Pubblicazione: (2025)
WavLink: Compact Audio-Text Embeddings with a Global Whisper Token
di: Kumar, Gokul Karthik, et al.
Pubblicazione: (2026)
di: Kumar, Gokul Karthik, et al.
Pubblicazione: (2026)
Interchangeable Token Embeddings for Extendable Vocabulary and Alpha-Equivalence
di: Işık, İlker, et al.
Pubblicazione: (2024)
di: Işık, İlker, et al.
Pubblicazione: (2024)
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer
di: Dong, Yihe, et al.
Pubblicazione: (2025)
di: Dong, Yihe, et al.
Pubblicazione: (2025)
Prompt Exploration with Prompt Regression
di: Feffer, Michael, et al.
Pubblicazione: (2024)
di: Feffer, Michael, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Wikontic: Constructing Wikidata-Aligned, Ontology-Aware Knowledge Graphs with Large Language Models
di: Chepurova, Alla, et al.
Pubblicazione: (2025) -
Scaling Transformer to 1M tokens and beyond with RMT
di: Bulatov, Aydar, et al.
Pubblicazione: (2023) -
Associative Recurrent Memory Transformer
di: Rodkin, Ivan, et al.
Pubblicazione: (2024) -
In Search of Needles in a 11M Haystack: Recurrent Memory Finds What LLMs Miss
di: Kuratov, Yuri, et al.
Pubblicazione: (2024) -
BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack
di: Kuratov, Yuri, et al.
Pubblicazione: (2024)