Effective Context in Transformers: An Analysis of Fragmentation and Tokenization
Fuente:
arXiv
Guardado en:
| Autores principales: | Fesharaki, Amirmehdi Jafari, Rami, Mohammadamin, Tchamkerten, Aslan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Feedback Increases the Capacity of Queues with Bounded Service Times
por: Sahasranand, K. R., et al.
Publicado: (2023)
por: Sahasranand, K. R., et al.
Publicado: (2023)
Attention with Markov: A Framework for Principled Analysis of Transformers via Markov Chains
por: Makkuva, Ashok Vardhan, et al.
Publicado: (2024)
por: Makkuva, Ashok Vardhan, et al.
Publicado: (2024)
In-Context Learning with Representations: Contextual Generalization of Trained Transformers
por: Yang, Tong, et al.
Publicado: (2024)
por: Yang, Tong, et al.
Publicado: (2024)
Measuring Uncertainty in Transformer Circuits with Effective Information Consistency
por: Krasnovsky, Anatoly A.
Publicado: (2025)
por: Krasnovsky, Anatoly A.
Publicado: (2025)
MultiTok: Variable-Length Tokenization for Efficient LLMs Adapted from LZW Compression
por: Elias, Noel, et al.
Publicado: (2024)
por: Elias, Noel, et al.
Publicado: (2024)
Filtering Beats Fine Tuning: A Bayesian Kalman View of In Context Learning in LLMs
por: Kiruluta, Andrew
Publicado: (2026)
por: Kiruluta, Andrew
Publicado: (2026)
Transformers on Markov Data: Constant Depth Suffices
por: Rajaraman, Nived, et al.
Publicado: (2024)
por: Rajaraman, Nived, et al.
Publicado: (2024)
Understanding Factual Recall in Transformers via Associative Memories
por: Nichani, Eshaan, et al.
Publicado: (2024)
por: Nichani, Eshaan, et al.
Publicado: (2024)
An Enhanced Text Compression Approach Using Transformer-based Language Models
por: Rahman, Chowdhury Mofizur, et al.
Publicado: (2024)
por: Rahman, Chowdhury Mofizur, et al.
Publicado: (2024)
Quantifying Logical Consistency in Transformers via Query-Key Alignment
por: Tulchinskii, Eduard, et al.
Publicado: (2025)
por: Tulchinskii, Eduard, et al.
Publicado: (2025)
NAZM: Network Analysis of Zonal Metrics in Persian Poetic Tradition
por: Shahnazari, Kourosh, et al.
Publicado: (2025)
por: Shahnazari, Kourosh, et al.
Publicado: (2025)
Subjective Depth and Timescale Transformers: Learning Where and When to Compute
por: Wieser, Frederico, et al.
Publicado: (2025)
por: Wieser, Frederico, et al.
Publicado: (2025)
Compression with Privacy-Preserving Random Access
por: Chandar, Venkat, et al.
Publicado: (2025)
por: Chandar, Venkat, et al.
Publicado: (2025)
Proposal and study of statistical features for string similarity computation and classification
por: Rodrigues, E. O., et al.
Publicado: (2026)
por: Rodrigues, E. O., et al.
Publicado: (2026)
Efficient Learned Data Compression via Dual-Stream Feature Decoupling
por: Ma, Huidong, et al.
Publicado: (2026)
por: Ma, Huidong, et al.
Publicado: (2026)
RateQuant: Optimal Mixed-Precision KV Cache Quantization via Rate-Distortion Theory
por: Zuo, Fei, et al.
Publicado: (2026)
por: Zuo, Fei, et al.
Publicado: (2026)
Speculative Decoding Scaling Laws (SDSL): Throughput Optimization Made Simple
por: Bozorgkhoo, Amirhossein, et al.
Publicado: (2026)
por: Bozorgkhoo, Amirhossein, et al.
Publicado: (2026)
A Rate-Distortion Framework for Summarization
por: Arda, Enes, et al.
Publicado: (2025)
por: Arda, Enes, et al.
Publicado: (2025)
Theoretical guarantees on the best-of-n alignment policy
por: Beirami, Ahmad, et al.
Publicado: (2024)
por: Beirami, Ahmad, et al.
Publicado: (2024)
Multi-turn Training with Basic Human Feedback Helps Little on LLM Reasoning
por: Liu, Qiang, et al.
Publicado: (2025)
por: Liu, Qiang, et al.
Publicado: (2025)
An Information-theoretic Multi-task Representation Learning Framework for Natural Language Understanding
por: Hu, Dou, et al.
Publicado: (2025)
por: Hu, Dou, et al.
Publicado: (2025)
InfAlign: Inference-aware language model alignment
por: Balashankar, Ananth, et al.
Publicado: (2024)
por: Balashankar, Ananth, et al.
Publicado: (2024)
A Mathematical Theory for Learning Semantic Languages by Abstract Learners
por: Liao, Kuo-Yu, et al.
Publicado: (2024)
por: Liao, Kuo-Yu, et al.
Publicado: (2024)
What Makes the Preferred Thinking Direction for LLMs in Multiple-choice Questions?
por: Zhang, Yizhe, et al.
Publicado: (2025)
por: Zhang, Yizhe, et al.
Publicado: (2025)
Fundamental Limits of Prompt Compression: A Rate-Distortion Framework for Black-Box Language Models
por: Nagle, Alliot, et al.
Publicado: (2024)
por: Nagle, Alliot, et al.
Publicado: (2024)
Cost-aware LLM-based Online Dataset Annotation
por: Elumar, Eray Can, et al.
Publicado: (2025)
por: Elumar, Eray Can, et al.
Publicado: (2025)
Iterative Counterfactual Data Augmentation
por: Plyler, Mitchell, et al.
Publicado: (2025)
por: Plyler, Mitchell, et al.
Publicado: (2025)
Latent Space Alignment for Semantic Channel Equalization
por: Hüttebräucker, Tomás, et al.
Publicado: (2024)
por: Hüttebräucker, Tomás, et al.
Publicado: (2024)
Information-Theoretic Generative Clustering of Documents
por: Du, Xin, et al.
Publicado: (2024)
por: Du, Xin, et al.
Publicado: (2024)
Context-Free Recognition with Transformers
por: Jerad, Selim, et al.
Publicado: (2026)
por: Jerad, Selim, et al.
Publicado: (2026)
Semantic Soft Bootstrapping: Long Context Reasoning in LLMs without Reinforcement Learning
por: Mitra, Purbesh, et al.
Publicado: (2025)
por: Mitra, Purbesh, et al.
Publicado: (2025)
Retrieval-Augmented Generation as Noisy In-Context Learning: A Unified Theory and Risk Bounds
por: Guo, Yang, et al.
Publicado: (2025)
por: Guo, Yang, et al.
Publicado: (2025)
Theoretical Limits of Language Model Alignment
por: Paes, Lucas Monteiro, et al.
Publicado: (2026)
por: Paes, Lucas Monteiro, et al.
Publicado: (2026)
HeavyWater and SimplexWater: Distortion-Free LLM Watermarks for Low-Entropy Next-Token Predictions
por: Tsur, Dor, et al.
Publicado: (2025)
por: Tsur, Dor, et al.
Publicado: (2025)
The Shape of Learning: Anisotropy and Intrinsic Dimensions in Transformer-Based Models
por: Razzhigaev, Anton, et al.
Publicado: (2023)
por: Razzhigaev, Anton, et al.
Publicado: (2023)
NDT: Non-Differential Transformer and Its Application to Sentiment Analysis
por: Ghoshal, Soudeep, et al.
Publicado: (2026)
por: Ghoshal, Soudeep, et al.
Publicado: (2026)
DTRNet: Dynamic Token Routing Network to Reduce Quadratic Costs in Transformers
por: Sharma, Aman, et al.
Publicado: (2025)
por: Sharma, Aman, et al.
Publicado: (2025)
The Atlas of In-Context Learning: How Attention Heads Shape In-Context Retrieval Augmentation
por: Kahardipraja, Patrick, et al.
Publicado: (2025)
por: Kahardipraja, Patrick, et al.
Publicado: (2025)
New Directions in Text Classification Research: Maximizing The Performance of Sentiment Classification from Limited Data
por: Agustian, Surya, et al.
Publicado: (2024)
por: Agustian, Surya, et al.
Publicado: (2024)
Optimizing Small Transformer-Based Language Models for Multi-Label Sentiment Analysis in Short Texts
por: Neumann, Julius, et al.
Publicado: (2025)
por: Neumann, Julius, et al.
Publicado: (2025)
Ejemplares similares
-
Feedback Increases the Capacity of Queues with Bounded Service Times
por: Sahasranand, K. R., et al.
Publicado: (2023) -
Attention with Markov: A Framework for Principled Analysis of Transformers via Markov Chains
por: Makkuva, Ashok Vardhan, et al.
Publicado: (2024) -
In-Context Learning with Representations: Contextual Generalization of Trained Transformers
por: Yang, Tong, et al.
Publicado: (2024) -
Measuring Uncertainty in Transformer Circuits with Effective Information Consistency
por: Krasnovsky, Anatoly A.
Publicado: (2025) -
MultiTok: Variable-Length Tokenization for Efficient LLMs Adapted from LZW Compression
por: Elias, Noel, et al.
Publicado: (2024)