DySCO: Dynamic Attention-Scaling Decoding for Long-Context Language Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Ye, Xi, Zhang, Wuwei, Yin, Fangcong, Yen, Howard, Chen, Danqi |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Query-Focused Retrieval Heads Improve Long-Context Reasoning and Re-ranking
par: Zhang, Wuwei, et autres
Publié: (2025)
par: Zhang, Wuwei, et autres
Publié: (2025)
LongProc: Benchmarking Long-Context Language Models on Long Procedural Generation
par: Ye, Xi, et autres
Publié: (2025)
par: Ye, Xi, et autres
Publié: (2025)
Long-Context Language Modeling with Parallel Context Encoding
par: Yen, Howard, et autres
Publié: (2024)
par: Yen, Howard, et autres
Publié: (2024)
Agentic Aggregation for Parallel Scaling of Long-Horizon Agentic Tasks
par: Lee, Yoonsang, et autres
Publié: (2026)
par: Lee, Yoonsang, et autres
Publié: (2026)
How to Train Long-Context Language Models (Effectively)
par: Gao, Tianyu, et autres
Publié: (2024)
par: Gao, Tianyu, et autres
Publié: (2024)
LoFiT: Localized Fine-tuning on LLM Representations
par: Yin, Fangcong, et autres
Publié: (2024)
par: Yin, Fangcong, et autres
Publié: (2024)
HELMET: How to Evaluate Long-Context Language Models Effectively and Thoroughly
par: Yen, Howard, et autres
Publié: (2024)
par: Yen, Howard, et autres
Publié: (2024)
Lost in the Maze: Overcoming Context Limitations in Long-Horizon Agentic Search
par: Yen, Howard, et autres
Publié: (2025)
par: Yen, Howard, et autres
Publié: (2025)
Language Models that Think, Chat Better
par: Bhaskar, Adithya, et autres
Publié: (2025)
par: Bhaskar, Adithya, et autres
Publié: (2025)
Understanding Synthetic Context Extension via Retrieval Heads
par: Zhao, Xinyu, et autres
Publié: (2024)
par: Zhao, Xinyu, et autres
Publié: (2024)
Continual Memorization of Factoids in Language Models
par: Chen, Howard, et autres
Publié: (2024)
par: Chen, Howard, et autres
Publié: (2024)
Dynamic Attention-Guided Context Decoding for Mitigating Context Faithfulness Hallucinations in Large Language Models
par: Huang, Yanwen, et autres
Publié: (2025)
par: Huang, Yanwen, et autres
Publié: (2025)
Learning Composable Chains-of-Thought
par: Yin, Fangcong, et autres
Publié: (2025)
par: Yin, Fangcong, et autres
Publié: (2025)
Scaling Context, Not Parameters: Training a Compact 7B Language Model for Efficient Long-Context Processing
par: Wu, Chen, et autres
Publié: (2025)
par: Wu, Chen, et autres
Publié: (2025)
AdaST: Dynamically Adapting Encoder States in the Decoder for End-to-End Speech-to-Text Translation
par: Huang, Wuwei, et autres
Publié: (2025)
par: Huang, Wuwei, et autres
Publié: (2025)
DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration
par: Zhang, Hanzhi, et autres
Publié: (2025)
par: Zhang, Hanzhi, et autres
Publié: (2025)
LongVILA: Scaling Long-Context Visual Language Models for Long Videos
par: Chen, Yukang, et autres
Publié: (2024)
par: Chen, Yukang, et autres
Publié: (2024)
Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters
par: Shyam, Vasudev, et autres
Publié: (2024)
par: Shyam, Vasudev, et autres
Publié: (2024)
Training-free Context-adaptive Attention for Efficient Long Context Modeling
par: You, Zeng, et autres
Publié: (2025)
par: You, Zeng, et autres
Publié: (2025)
Training-Free Long-Context Scaling of Large Language Models
par: An, Chenxin, et autres
Publié: (2024)
par: An, Chenxin, et autres
Publié: (2024)
Cache Me If You Can: How Many KVs Do You Need for Effective Long-Context LMs?
par: Bhaskar, Adithya, et autres
Publié: (2025)
par: Bhaskar, Adithya, et autres
Publié: (2025)
MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding
par: Sadhukhan, Ranajoy, et autres
Publié: (2024)
par: Sadhukhan, Ranajoy, et autres
Publié: (2024)
Efficient Context Scaling with LongCat ZigZag Attention
par: Zhang, Chen, et autres
Publié: (2025)
par: Zhang, Chen, et autres
Publié: (2025)
Precise Information Control in Long-Form Text Generation
par: He, Jacqueline, et autres
Publié: (2025)
par: He, Jacqueline, et autres
Publié: (2025)
HERA: Improving Long Document Summarization using Large Language Models with Context Packaging and Reordering
par: Li, Taiji, et autres
Publié: (2025)
par: Li, Taiji, et autres
Publié: (2025)
Detecting and Suppressing Reward Hacking with Gradient Fingerprints
par: Wang, Songtao, et autres
Publié: (2026)
par: Wang, Songtao, et autres
Publié: (2026)
Infinite Retrieval: Attention Enhanced LLMs in Long-Context Processing
par: Ye, Xiaoju, et autres
Publié: (2025)
par: Ye, Xiaoju, et autres
Publié: (2025)
Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting
par: Chen, Howard, et autres
Publié: (2025)
par: Chen, Howard, et autres
Publié: (2025)
The Heuristic Core: Understanding Subnetwork Generalization in Pretrained Language Models
par: Bhaskar, Adithya, et autres
Publié: (2024)
par: Bhaskar, Adithya, et autres
Publié: (2024)
Unveiling Reasoning Thresholds in Language Models: Scaling, Fine-Tuning, and Interpretability through Attention Maps
par: Hsiao, Yen-Che, et autres
Publié: (2025)
par: Hsiao, Yen-Che, et autres
Publié: (2025)
ScaleFormer: Span Representation Cumulation for Long-Context Transformer
par: Du, Jiangshu, et autres
Publié: (2025)
par: Du, Jiangshu, et autres
Publié: (2025)
Beyond In-Context Learning: Aligning Long-form Generation of Large Language Models via Task-Inherent Attribute Guidelines
par: Long, Do Xuan, et autres
Publié: (2025)
par: Long, Do Xuan, et autres
Publié: (2025)
Linear Representation Transferability Hypothesis: Leveraging Small Models to Steer Large Models
par: Bello, Femi, et autres
Publié: (2025)
par: Bello, Femi, et autres
Publié: (2025)
Understanding the RoPE Extensions of Long-Context LLMs: An Attention Perspective
par: Zhong, Meizhi, et autres
Publié: (2024)
par: Zhong, Meizhi, et autres
Publié: (2024)
Extracting Rule-based Descriptions of Attention Features in Transformers
par: Friedman, Dan, et autres
Publié: (2025)
par: Friedman, Dan, et autres
Publié: (2025)
SEAL: Scaling to Emphasize Attention for Long-Context Retrieval
par: Lee, Changhun, et autres
Publié: (2025)
par: Lee, Changhun, et autres
Publié: (2025)
RAPID: Long-Context Inference with Retrieval-Augmented Speculative Decoding
par: Chen, Guanzheng, et autres
Publié: (2025)
par: Chen, Guanzheng, et autres
Publié: (2025)
Dynamic Chunking and Selection for Reading Comprehension of Ultra-Long Context in Large Language Models
par: Sheng, Boheng, et autres
Publié: (2025)
par: Sheng, Boheng, et autres
Publié: (2025)
LycheeDecode: Accelerating Long-Context LLM Inference via Hybrid-Head Sparse Decoding
par: Lin, Gang, et autres
Publié: (2026)
par: Lin, Gang, et autres
Publié: (2026)
P1SCO: Social Dimensions from a Perspectivist Lens
par: Curry, Amanda Cercas, et autres
Publié: (2026)
par: Curry, Amanda Cercas, et autres
Publié: (2026)
Documents similaires
-
Query-Focused Retrieval Heads Improve Long-Context Reasoning and Re-ranking
par: Zhang, Wuwei, et autres
Publié: (2025) -
LongProc: Benchmarking Long-Context Language Models on Long Procedural Generation
par: Ye, Xi, et autres
Publié: (2025) -
Long-Context Language Modeling with Parallel Context Encoding
par: Yen, Howard, et autres
Publié: (2024) -
Agentic Aggregation for Parallel Scaling of Long-Horizon Agentic Tasks
par: Lee, Yoonsang, et autres
Publié: (2026) -
How to Train Long-Context Language Models (Effectively)
par: Gao, Tianyu, et autres
Publié: (2024)