Short Data, Long Context: Distilling Positional Knowledge in Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Huber, Patrick, Chang, Ernie, Sankar, Chinnadhurai, Conway, Rylan, Fedorov, Igor, Arefin, Md Rifat, Sagar, Adithya |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CoSMoEs: Compact Sparse Mixture of Experts
by: Huber, Patrick, et al.
Published: (2025)
by: Huber, Patrick, et al.
Published: (2025)
CoDi: Conversational Distillation for Grounded Question Answering
by: Huber, Patrick, et al.
Published: (2024)
by: Huber, Patrick, et al.
Published: (2024)
MobileLLM-Flash: Latency-Guided On-Device LLM Design for Industry Scale Deployment
by: Huang, Hanxian, et al.
Published: (2026)
by: Huang, Hanxian, et al.
Published: (2026)
PrE-Text: Training Language Models on Private Federated Data in the Age of LLMs
by: Hou, Charlie, et al.
Published: (2024)
by: Hou, Charlie, et al.
Published: (2024)
Does Representation Matter? Exploring Intermediate Layers in Large Language Models
by: Skean, Oscar, et al.
Published: (2024)
by: Skean, Oscar, et al.
Published: (2024)
Seq-VCR: Preventing Collapse in Intermediate Transformer Representations for Enhanced Reasoning
by: Arefin, Md Rifat, et al.
Published: (2024)
by: Arefin, Md Rifat, et al.
Published: (2024)
MobileLLM-Pro Technical Report
by: Huber, Patrick, et al.
Published: (2025)
by: Huber, Patrick, et al.
Published: (2025)
Black-box Context-free Grammar Inference for Readable & Natural Grammars
by: Arefin, Mohammad Rifat, et al.
Published: (2025)
by: Arefin, Mohammad Rifat, et al.
Published: (2025)
Transformer-Driven Triple Fusion Framework for Enhanced Multimodal Author Intent Classification in Low-Resource Bangla
by: Islam, Ariful, et al.
Published: (2025)
by: Islam, Ariful, et al.
Published: (2025)
LongReD: Mitigating Short-Text Degradation of Long-Context Large Language Models via Restoration Distillation
by: Dong, Zican, et al.
Published: (2025)
by: Dong, Zican, et al.
Published: (2025)
Fast Deterministic Black-box Context-free Grammar Inference
by: Arefin, Mohammad Rifat, et al.
Published: (2023)
by: Arefin, Mohammad Rifat, et al.
Published: (2023)
Brewing Knowledge in Context: Distillation Perspectives on In-Context Learning
by: Li, Chengye, et al.
Published: (2025)
by: Li, Chengye, et al.
Published: (2025)
Long-Short Alignment for Effective Long-Context Modeling in LLMs
by: Du, Tianqi, et al.
Published: (2025)
by: Du, Tianqi, et al.
Published: (2025)
Layer by Layer: Uncovering Hidden Representations in Language Models
by: Skean, Oscar, et al.
Published: (2025)
by: Skean, Oscar, et al.
Published: (2025)
BanglaSentNet: An Explainable Hybrid Deep Learning Framework for Multi-Aspect Sentiment Analysis with Cross-Domain Transfer Learning
by: Islam, Ariful, et al.
Published: (2025)
by: Islam, Ariful, et al.
Published: (2025)
Latent Context Compilation: Distilling Long Context into Compact Portable Memory
by: Li, Zeju, et al.
Published: (2026)
by: Li, Zeju, et al.
Published: (2026)
Scavenging Hyena: Distilling Transformers into Long Convolution Models
by: Ralambomihanta, Tokiniaina Raharison, et al.
Published: (2024)
by: Ralambomihanta, Tokiniaina Raharison, et al.
Published: (2024)
Core Context Aware Transformers for Long Context Language Modeling
by: Chen, Yaofo, et al.
Published: (2024)
by: Chen, Yaofo, et al.
Published: (2024)
LongPO: Long Context Self-Evolution of Large Language Models through Short-to-Long Preference Optimization
by: Chen, Guanzheng, et al.
Published: (2025)
by: Chen, Guanzheng, et al.
Published: (2025)
Long Context Alignment with Short Instructions and Synthesized Positions
by: Wu, Wenhao, et al.
Published: (2024)
by: Wu, Wenhao, et al.
Published: (2024)
AutoMixer: Checkpoint Artifacts as Automatic Data Mixers
by: Chang, Ernie, et al.
Published: (2025)
by: Chang, Ernie, et al.
Published: (2025)
How Different Tokenization Algorithms Impact LLMs and Transformer Models for Binary Code Analysis
by: Mostafa, Ahmed, et al.
Published: (2025)
by: Mostafa, Ahmed, et al.
Published: (2025)
Navigating the Minefield of MT Beam Search in Cascaded Streaming Speech Translation
by: Rabatin, Rastislav, et al.
Published: (2024)
by: Rabatin, Rastislav, et al.
Published: (2024)
Rotary Positional Embeddings as Phase Modulation: Theoretical Bounds on the RoPE Base for Long-Context Transformers
by: Liu, Feilong
Published: (2026)
by: Liu, Feilong
Published: (2026)
CoCA: Fusing Position Embedding with Collinear Constrained Attention in Transformers for Long Context Window Extending
by: Zhu, Shiyi, et al.
Published: (2023)
by: Zhu, Shiyi, et al.
Published: (2023)
Extracting Rule-based Descriptions of Attention Features in Transformers
by: Friedman, Dan, et al.
Published: (2025)
by: Friedman, Dan, et al.
Published: (2025)
Self-Vocabularizing Training for Neural Machine Translation
by: Lin, Pin-Jie, et al.
Published: (2025)
by: Lin, Pin-Jie, et al.
Published: (2025)
MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases
by: Liu, Zechun, et al.
Published: (2024)
by: Liu, Zechun, et al.
Published: (2024)
Multi-modal Anchor Gated Transformer with Knowledge Distillation for Emotion Recognition in Conversation
by: Li, Jie, et al.
Published: (2025)
by: Li, Jie, et al.
Published: (2025)
Compress, Gather, and Recompute: REFORMing Long-Context Processing in Transformers
by: Song, Woomin, et al.
Published: (2025)
by: Song, Woomin, et al.
Published: (2025)
Context Biasing for Pronunciation-Orthography Mismatch in Automatic Speech Recognition
by: Huber, Christian, et al.
Published: (2025)
by: Huber, Christian, et al.
Published: (2025)
HMT: Hierarchical Memory Transformer for Efficient Long Context Language Processing
by: He, Zifan, et al.
Published: (2024)
by: He, Zifan, et al.
Published: (2024)
Diagonal Batching Unlocks Parallelism in Recurrent Memory Transformers for Long Contexts
by: Sivtsov, Danil, et al.
Published: (2025)
by: Sivtsov, Danil, et al.
Published: (2025)
Retrieval-Augmented Generation Meets Data-Driven Tabula Rasa Approach for Temporal Knowledge Graph Forecasting
by: Sannidhi, Geethan, et al.
Published: (2024)
by: Sannidhi, Geethan, et al.
Published: (2024)
Min-p, Max Exaggeration: A Critical Analysis of Min-p Sampling in Language Models
by: Schaeffer, Rylan, et al.
Published: (2025)
by: Schaeffer, Rylan, et al.
Published: (2025)
LongAct: Harnessing Intrinsic Activation Patterns for Long-Context Reinforcement Learning
by: Ping, Bowen, et al.
Published: (2026)
by: Ping, Bowen, et al.
Published: (2026)
Knowledge Distillation with Training Wheels
by: Liu, Guanlin, et al.
Published: (2025)
by: Liu, Guanlin, et al.
Published: (2025)
Position: The Turing-Completeness of Autoregressive Transformers Relies Heavily on Context Management
by: Cui, Guanyu, et al.
Published: (2026)
by: Cui, Guanyu, et al.
Published: (2026)
HarmAug: Effective Data Augmentation for Knowledge Distillation of Safety Guard Models
by: Lee, Seanie, et al.
Published: (2024)
by: Lee, Seanie, et al.
Published: (2024)
BanglaASTE: A Novel Framework for Aspect-Sentiment-Opinion Extraction in Bangla E-commerce Reviews Using Ensemble Deep Learning
by: Islam, Ariful, et al.
Published: (2025)
by: Islam, Ariful, et al.
Published: (2025)
Similar Items
-
CoSMoEs: Compact Sparse Mixture of Experts
by: Huber, Patrick, et al.
Published: (2025) -
CoDi: Conversational Distillation for Grounded Question Answering
by: Huber, Patrick, et al.
Published: (2024) -
MobileLLM-Flash: Latency-Guided On-Device LLM Design for Industry Scale Deployment
by: Huang, Hanxian, et al.
Published: (2026) -
PrE-Text: Training Language Models on Private Federated Data in the Age of LLMs
by: Hou, Charlie, et al.
Published: (2024) -
Does Representation Matter? Exploring Intermediate Layers in Large Language Models
by: Skean, Oscar, et al.
Published: (2024)