Rope to Nope and Back Again: A New Hybrid Attention Strategy
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Bowen, Venkitesh, Bharat, Talupuru, Dwarak, Lin, Hangyu, Cairuz, David, Blunsom, Phil, Locatelli, Acyr |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Aya 23: Open Weight Releases to Further Multilingual Progress
by: Aryabumi, Viraat, et al.
Published: (2024)
by: Aryabumi, Viraat, et al.
Published: (2024)
SnapKV: LLM Knows What You are Looking for Before Generation
by: Li, Yuhong, et al.
Published: (2024)
by: Li, Yuhong, et al.
Published: (2024)
BAM! Just Like That: Simple and Efficient Parameter Upcycling for Mixture of Experts
by: Zhang, Qizhen, et al.
Published: (2024)
by: Zhang, Qizhen, et al.
Published: (2024)
Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models
by: Ruis, Laura, et al.
Published: (2024)
by: Ruis, Laura, et al.
Published: (2024)
Aya Expanse: Combining Research Breakthroughs for a New Multilingual Frontier
by: Dang, John, et al.
Published: (2024)
by: Dang, John, et al.
Published: (2024)
Human Feedback is not Gold Standard
by: Hosking, Tom, et al.
Published: (2023)
by: Hosking, Tom, et al.
Published: (2023)
Efficient Benchmarking Is Just Feature Selection and Multiple Regression
by: Bowyer, Sam, et al.
Published: (2026)
by: Bowyer, Sam, et al.
Published: (2026)
One Tokenizer To Rule Them All: Emergent Language Plasticity via Multilingual Tokenizers
by: Abagyan, Diana, et al.
Published: (2025)
by: Abagyan, Diana, et al.
Published: (2025)
Nexus: Specialization meets Adaptability for Efficiently Training Mixture of Experts
by: Gritsch, Nikolas, et al.
Published: (2024)
by: Gritsch, Nikolas, et al.
Published: (2024)
Understanding Likelihood Over-optimisation in Direct Alignment Algorithms
by: Shi, Zhengyan, et al.
Published: (2024)
by: Shi, Zhengyan, et al.
Published: (2024)
Aya Vision: Advancing the Frontier of Multilingual Multimodality
by: Dash, Saurabh, et al.
Published: (2025)
by: Dash, Saurabh, et al.
Published: (2025)
Uncertainty-Aware Step-wise Verification with Generative Reward Models
by: Ye, Zihuiwen, et al.
Published: (2025)
by: Ye, Zihuiwen, et al.
Published: (2025)
Tiny Aya: Bridging Scale and Multilingual Depth
by: Salamanca, Alejandro R., et al.
Published: (2026)
by: Salamanca, Alejandro R., et al.
Published: (2026)
Improving Reward Models with Synthetic Critiques
by: Ye, Zihuiwen, et al.
Published: (2024)
by: Ye, Zihuiwen, et al.
Published: (2024)
To Code, or Not To Code? Exploring Impact of Code in Pre-training
by: Aryabumi, Viraat, et al.
Published: (2024)
by: Aryabumi, Viraat, et al.
Published: (2024)
Asking Again and Again: Exploring LLM Robustness to Repeated Questions
by: Shaier, Sagi, et al.
Published: (2024)
by: Shaier, Sagi, et al.
Published: (2024)
A Hybrid Defense Strategy for Boosting Adversarial Robustness in Vision-Language Models
by: Liang, Yuhan, et al.
Published: (2024)
by: Liang, Yuhan, et al.
Published: (2024)
Cramming 1568 Tokens into a Single Vector and Back Again: Exploring the Limits of Embedding Space Capacity
by: Kuratov, Yuri, et al.
Published: (2025)
by: Kuratov, Yuri, et al.
Published: (2025)
There and Back Again: A Netlist's Tale with Much Egraphin'
by: Smith, Gus Henry, et al.
Published: (2024)
by: Smith, Gus Henry, et al.
Published: (2024)
From Practice to Theory and Back Again.
by: Kramsch, Claire
Published: (2002)
by: Kramsch, Claire
Published: (2002)
WuNeng: Hybrid State with Attention
by: Xiao, Liu, et al.
Published: (2025)
by: Xiao, Liu, et al.
Published: (2025)
A Hybrid Attention Framework for Fake News Detection with Large Language Models
by: Xu, Xiaochuan, et al.
Published: (2025)
by: Xu, Xiaochuan, et al.
Published: (2025)
In-person, Online and Back Again -- A Tale of Three Hybrid Hackathons
by: Affia-Jomants, Abasi-amefon Obot, et al.
Published: (2025)
by: Affia-Jomants, Abasi-amefon Obot, et al.
Published: (2025)
Full Attention Strikes Back: Transferring Full Attention into Sparse within Hundred Training Steps
by: Zhou, Yanke, et al.
Published: (2026)
by: Zhou, Yanke, et al.
Published: (2026)
Olmo Hybrid: From Theory to Practice and Back
by: Merrill, William, et al.
Published: (2026)
by: Merrill, William, et al.
Published: (2026)
Hello Again! LLM-powered Personalized Agent for Long-term Dialogue
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
HyLRA: Hybrid Layer Reuse Attention for Efficient Long-Context Inference
by: Ai, Xuan, et al.
Published: (2026)
by: Ai, Xuan, et al.
Published: (2026)
Elias in the Lighthouse, Again? Diagnosing Low Diversity in LLM Stories
by: Hamilton, Sil, et al.
Published: (2026)
by: Hamilton, Sil, et al.
Published: (2026)
Long Context Pre-Training with Lighthouse Attention
by: Peng, Bowen, et al.
Published: (2026)
by: Peng, Bowen, et al.
Published: (2026)
Controlling Copatterns: There and Back Again (Extended Version)
by: Downen, Paul
Published: (2025)
by: Downen, Paul
Published: (2025)
Blockwise SFT for Diffusion Language Models: Reconciling Bidirectional Attention and Autoregressive Decoding
by: Sun, Bowen, et al.
Published: (2025)
by: Sun, Bowen, et al.
Published: (2025)
Back Attention: Understanding and Enhancing Multi-Hop Reasoning in Large Language Models
by: Yu, Zeping, et al.
Published: (2025)
by: Yu, Zeping, et al.
Published: (2025)
Losing our Tail, Again: (Un)Natural Selection & Multilingual LLMs
by: Vanmassenhove, Eva
Published: (2025)
by: Vanmassenhove, Eva
Published: (2025)
A Systematic Analysis of Hybrid Linear Attention
by: Wang, Dustin, et al.
Published: (2025)
by: Wang, Dustin, et al.
Published: (2025)
Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference
by: Qiu, Quantong, et al.
Published: (2026)
by: Qiu, Quantong, et al.
Published: (2026)
Before Generation, Align it! A Novel and Effective Strategy for Mitigating Hallucinations in Text-to-SQL Generation
by: Qu, Ge, et al.
Published: (2024)
by: Qu, Ge, et al.
Published: (2024)
Positional Attention for Efficient BERT-Based Named Entity Recognition
by: Sun, Mo, et al.
Published: (2025)
by: Sun, Mo, et al.
Published: (2025)
Optimizing Native Sparse Attention with Latent Attention and Local Global Alternating Strategies
by: Hu, Yuxuan, et al.
Published: (2025)
by: Hu, Yuxuan, et al.
Published: (2025)
TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization
by: Yao, Dingyu, et al.
Published: (2025)
by: Yao, Dingyu, et al.
Published: (2025)
Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement Learning
by: Qin, Yulei, et al.
Published: (2025)
by: Qin, Yulei, et al.
Published: (2025)
Similar Items
-
Aya 23: Open Weight Releases to Further Multilingual Progress
by: Aryabumi, Viraat, et al.
Published: (2024) -
SnapKV: LLM Knows What You are Looking for Before Generation
by: Li, Yuhong, et al.
Published: (2024) -
BAM! Just Like That: Simple and Efficient Parameter Upcycling for Mixture of Experts
by: Zhang, Qizhen, et al.
Published: (2024) -
Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models
by: Ruis, Laura, et al.
Published: (2024) -
Aya Expanse: Combining Research Breakthroughs for a New Multilingual Frontier
by: Dang, John, et al.
Published: (2024)