Long Context Pre-Training with Lighthouse Attention
Fuente:
arXiv
Salvato in:
| Autori principali: | Peng, Bowen, Ghosh, Subho, Quesnelle, Jeffrey |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Efficient Pre-Training with Token Superposition
di: Peng, Bowen, et al.
Pubblicazione: (2026)
di: Peng, Bowen, et al.
Pubblicazione: (2026)
Decoupling the Benefits of Subword Tokenization for Language Model Training via Byte-level Simulation
di: Gigant, Théo, et al.
Pubblicazione: (2026)
di: Gigant, Théo, et al.
Pubblicazione: (2026)
YaRN: Efficient Context Window Extension of Large Language Models
di: Peng, Bowen, et al.
Pubblicazione: (2023)
di: Peng, Bowen, et al.
Pubblicazione: (2023)
Hermes 3 Technical Report
di: Teknium, Ryan, et al.
Pubblicazione: (2024)
di: Teknium, Ryan, et al.
Pubblicazione: (2024)
Training-free Context-adaptive Attention for Efficient Long Context Modeling
di: You, Zeng, et al.
Pubblicazione: (2025)
di: You, Zeng, et al.
Pubblicazione: (2025)
Lag-Relative Sparse Attention In Long Context Training
di: Liang, Manlai, et al.
Pubblicazione: (2025)
di: Liang, Manlai, et al.
Pubblicazione: (2025)
garak: A Framework for Security Probing Large Language Models
di: Derczynski, Leon, et al.
Pubblicazione: (2024)
di: Derczynski, Leon, et al.
Pubblicazione: (2024)
Attention Reveals More Than Tokens: Training-Free Long-Context Reasoning with Attention-guided Retrieval
di: Zhang, Yuwei, et al.
Pubblicazione: (2025)
di: Zhang, Yuwei, et al.
Pubblicazione: (2025)
A Little Goes a Long Way: Efficient Long Context Training and Inference with Partial Contexts
di: Ge, Suyu, et al.
Pubblicazione: (2024)
di: Ge, Suyu, et al.
Pubblicazione: (2024)
HyLRA: Hybrid Layer Reuse Attention for Efficient Long-Context Inference
di: Ai, Xuan, et al.
Pubblicazione: (2026)
di: Ai, Xuan, et al.
Pubblicazione: (2026)
MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
di: Jiang, Huiqiang, et al.
Pubblicazione: (2024)
di: Jiang, Huiqiang, et al.
Pubblicazione: (2024)
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training
di: Li, Wenxuan, et al.
Pubblicazione: (2025)
di: Li, Wenxuan, et al.
Pubblicazione: (2025)
The Lighthouse of Language: Enhancing LLM Agents via Critique-Guided Improvement
di: Yang, Ruihan, et al.
Pubblicazione: (2025)
di: Yang, Ruihan, et al.
Pubblicazione: (2025)
Untie the Knots: An Efficient Data Augmentation Strategy for Long-Context Pre-Training in Language Models
di: Tian, Junfeng, et al.
Pubblicazione: (2024)
di: Tian, Junfeng, et al.
Pubblicazione: (2024)
Long-Context Generalization with Sparse Attention
di: Vasylenko, Pavlo, et al.
Pubblicazione: (2025)
di: Vasylenko, Pavlo, et al.
Pubblicazione: (2025)
Revealing the Learning Dynamics of Long-Context Continual Pre-training
di: Liang, Yupu, et al.
Pubblicazione: (2026)
di: Liang, Yupu, et al.
Pubblicazione: (2026)
Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data
di: Guo, Xu, et al.
Pubblicazione: (2026)
di: Guo, Xu, et al.
Pubblicazione: (2026)
Elias in the Lighthouse, Again? Diagnosing Low Diversity in LLM Stories
di: Hamilton, Sil, et al.
Pubblicazione: (2026)
di: Hamilton, Sil, et al.
Pubblicazione: (2026)
Adamas: Hadamard Sparse Attention for Efficient Long-Context Inference
di: Yan, Siyuan, et al.
Pubblicazione: (2025)
di: Yan, Siyuan, et al.
Pubblicazione: (2025)
Infinite Retrieval: Attention Enhanced LLMs in Long-Context Processing
di: Ye, Xiaoju, et al.
Pubblicazione: (2025)
di: Ye, Xiaoju, et al.
Pubblicazione: (2025)
Squeezed Attention: Accelerating Long Context Length LLM Inference
di: Hooper, Coleman, et al.
Pubblicazione: (2024)
di: Hooper, Coleman, et al.
Pubblicazione: (2024)
Multipole Attention for Efficient Long Context Reasoning
di: Hooper, Coleman, et al.
Pubblicazione: (2025)
di: Hooper, Coleman, et al.
Pubblicazione: (2025)
SPECTRUM: Speaker-Enhanced Pre-Training for Long Dialogue Summarization
di: Cho, Sangwoo, et al.
Pubblicazione: (2024)
di: Cho, Sangwoo, et al.
Pubblicazione: (2024)
GRKV: Global Regression for Training-Free KV Cache Compression in Long-Context LLMs
di: Peng, Junjie, et al.
Pubblicazione: (2026)
di: Peng, Junjie, et al.
Pubblicazione: (2026)
Ltri-LLM: Streaming Long Context Inference for LLMs with Training-Free Dynamic Triangular Attention Pattern
di: Tang, Hongyin, et al.
Pubblicazione: (2024)
di: Tang, Hongyin, et al.
Pubblicazione: (2024)
QwenLong-L1.5: Post-Training Recipe for Long-Context Reasoning and Memory Management
di: Shen, Weizhou, et al.
Pubblicazione: (2025)
di: Shen, Weizhou, et al.
Pubblicazione: (2025)
ReAttention: Training-Free Infinite Context with Finite Attention Scope
di: Liu, Xiaoran, et al.
Pubblicazione: (2024)
di: Liu, Xiaoran, et al.
Pubblicazione: (2024)
LongR: Unleashing Long-Context Reasoning via Reinforcement Learning with Dense Utility Rewards
di: Ping, Bowen, et al.
Pubblicazione: (2026)
di: Ping, Bowen, et al.
Pubblicazione: (2026)
SPLA: Block Sparse Plus Linear Attention for Long Context Modeling
di: Wang, Bailin, et al.
Pubblicazione: (2026)
di: Wang, Bailin, et al.
Pubblicazione: (2026)
Understanding the RoPE Extensions of Long-Context LLMs: An Attention Perspective
di: Zhong, Meizhi, et al.
Pubblicazione: (2024)
di: Zhong, Meizhi, et al.
Pubblicazione: (2024)
Lighthouse: A User-Friendly Library for Reproducible Video Moment Retrieval and Highlight Detection
di: Nishimura, Taichi, et al.
Pubblicazione: (2024)
di: Nishimura, Taichi, et al.
Pubblicazione: (2024)
LongAttn: Selecting Long-context Training Data via Token-level Attention
di: Wu, Longyun, et al.
Pubblicazione: (2025)
di: Wu, Longyun, et al.
Pubblicazione: (2025)
Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters
di: Shyam, Vasudev, et al.
Pubblicazione: (2024)
di: Shyam, Vasudev, et al.
Pubblicazione: (2024)
MuDAF: Long-Context Multi-Document Attention Focusing through Contrastive Learning on Attention Heads
di: Liu, Weihao, et al.
Pubblicazione: (2025)
di: Liu, Weihao, et al.
Pubblicazione: (2025)
LLMSteer: Improving Long-Context LLM Inference by Steering Attention on Reused Contexts
di: Gu, Zhuohan, et al.
Pubblicazione: (2024)
di: Gu, Zhuohan, et al.
Pubblicazione: (2024)
Structured Packing in LLM Training Improves Long Context Utilization
di: Staniszewski, Konrad, et al.
Pubblicazione: (2023)
di: Staniszewski, Konrad, et al.
Pubblicazione: (2023)
Training-Free Long-Context Scaling of Large Language Models
di: An, Chenxin, et al.
Pubblicazione: (2024)
di: An, Chenxin, et al.
Pubblicazione: (2024)
NExtLong: Toward Effective Long-Context Training without Long Documents
di: Gao, Chaochen, et al.
Pubblicazione: (2025)
di: Gao, Chaochen, et al.
Pubblicazione: (2025)
LongHeads: Multi-Head Attention is Secretly a Long Context Processor
di: Lu, Yi, et al.
Pubblicazione: (2024)
di: Lu, Yi, et al.
Pubblicazione: (2024)
DySCO: Dynamic Attention-Scaling Decoding for Long-Context Language Models
di: Ye, Xi, et al.
Pubblicazione: (2026)
di: Ye, Xi, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Efficient Pre-Training with Token Superposition
di: Peng, Bowen, et al.
Pubblicazione: (2026) -
Decoupling the Benefits of Subword Tokenization for Language Model Training via Byte-level Simulation
di: Gigant, Théo, et al.
Pubblicazione: (2026) -
YaRN: Efficient Context Window Extension of Large Language Models
di: Peng, Bowen, et al.
Pubblicazione: (2023) -
Hermes 3 Technical Report
di: Teknium, Ryan, et al.
Pubblicazione: (2024) -
Training-free Context-adaptive Attention for Efficient Long Context Modeling
di: You, Zeng, et al.
Pubblicazione: (2025)