WuNeng: Hybrid State with Attention
Fuente:
arXiv
Saved in:
| Main Authors: | Xiao, Liu, Zhiyuan, Li, Yueyu, Lin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
State Tuning: State-based Test-Time Scaling on RWKV-7
by: Xiao, Liu, et al.
Published: (2025)
by: Xiao, Liu, et al.
Published: (2025)
Cross-attention for State-based model RWKV-7
by: Xiao, Liu, et al.
Published: (2025)
by: Xiao, Liu, et al.
Published: (2025)
Millions of States: Designing a Scalable MoE Architecture with RWKV-7 Meta-learner
by: Xiao, Liu, et al.
Published: (2025)
by: Xiao, Liu, et al.
Published: (2025)
ARWKV: Pretrain is not what we need, an RNN-Attention-Based Language Model Born from Transformer
by: Yueyu, Lin, et al.
Published: (2025)
by: Yueyu, Lin, et al.
Published: (2025)
Learning to Focus: Causal Attention Distillation via Gradient-Guided Token Pruning
by: Guo, Yiju, et al.
Published: (2025)
by: Guo, Yiju, et al.
Published: (2025)
States Hidden in Hidden States: LLMs Emerge Discrete State Representations Implicitly
by: Chen, Junhao, et al.
Published: (2024)
by: Chen, Junhao, et al.
Published: (2024)
Hybrid Linear Attention Done Right: Efficient Distillation and Effective Architectures for Extremely Long Contexts
by: Chen, Yingfa, et al.
Published: (2026)
by: Chen, Yingfa, et al.
Published: (2026)
Machine Translation Evaluation Benchmark for Wu Chinese: Workflow and Analysis
by: Yu, Hongjian, et al.
Published: (2024)
by: Yu, Hongjian, et al.
Published: (2024)
Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference
by: Qiu, Quantong, et al.
Published: (2026)
by: Qiu, Quantong, et al.
Published: (2026)
A Systematic Analysis of Hybrid Linear Attention
by: Wang, Dustin, et al.
Published: (2025)
by: Wang, Dustin, et al.
Published: (2025)
Rope to Nope and Back Again: A New Hybrid Attention Strategy
by: Yang, Bowen, et al.
Published: (2025)
by: Yang, Bowen, et al.
Published: (2025)
HyLRA: Hybrid Layer Reuse Attention for Efficient Long-Context Inference
by: Ai, Xuan, et al.
Published: (2026)
by: Ai, Xuan, et al.
Published: (2026)
Switch Attention: Towards Dynamic and Fine-grained Hybrid Transformers
by: Zhao, Yusheng, et al.
Published: (2026)
by: Zhao, Yusheng, et al.
Published: (2026)
Punctuation-aware Hybrid Trainable Sparse Attention for Large Language Models
by: Qiu, Junxiang, et al.
Published: (2026)
by: Qiu, Junxiang, et al.
Published: (2026)
Reinforced Attention Learning
by: Li, Bangzheng, et al.
Published: (2026)
by: Li, Bangzheng, et al.
Published: (2026)
MiniCPM-SALA: Hybridizing Sparse and Linear Attention for Efficient Long-Context Modeling
by: MiniCPM Team, et al.
Published: (2026)
by: MiniCPM Team, et al.
Published: (2026)
Dynamic Adaptive Attention and Supervised Contrastive Learning: A Novel Hybrid Framework for Text Sentiment Classification
by: Li, Qingyang
Published: (2026)
by: Li, Qingyang
Published: (2026)
DREAMSTATE: Diffusing States and Parameters for Recurrent Large Language Models
by: Xiao, Liu
Published: (2026)
by: Xiao, Liu
Published: (2026)
Robust Audio-Text Retrieval via Cross-Modal Attention and Hybrid Loss
by: Liu, Meizhu, et al.
Published: (2026)
by: Liu, Meizhu, et al.
Published: (2026)
RouteLMT: Learned Sample Routing for Hybrid LLM Translation Deployment
by: Luo, Yingfeng, et al.
Published: (2026)
by: Luo, Yingfeng, et al.
Published: (2026)
MossNet: Mixture of State-Space Experts is a Multi-Head Attention
by: Tuli, Shikhar, et al.
Published: (2025)
by: Tuli, Shikhar, et al.
Published: (2025)
Scaling Linear Attention with Sparse State Expansion
by: Pan, Yuqi, et al.
Published: (2025)
by: Pan, Yuqi, et al.
Published: (2025)
Resolving the Robustness-Precision Trade-off in Financial RAG through Hybrid Document-Routed Retrieval
by: Cheng, Zhiyuan, et al.
Published: (2026)
by: Cheng, Zhiyuan, et al.
Published: (2026)
Rodimus*: Breaking the Accuracy-Efficiency Trade-Off with Efficient Attentions
by: He, Zhihao, et al.
Published: (2024)
by: He, Zhihao, et al.
Published: (2024)
What Training Data Teaches RL Memory Agents: An Empirical Study of Curriculum Effects in Memory-Augmented QA
by: He, Xinjie, et al.
Published: (2026)
by: He, Xinjie, et al.
Published: (2026)
APB-V: Accelerating Long-Video Understanding via Sequence-Parallelism-aware Approximate Attention
by: Huang, Yuxiang, et al.
Published: (2026)
by: Huang, Yuxiang, et al.
Published: (2026)
SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature Space
by: Shen, Zhenyi, et al.
Published: (2025)
by: Shen, Zhenyi, et al.
Published: (2025)
Alleviating Forgetfulness of Linear Attention by Hybrid Sparse Attention and Contextualized Learnable Token Eviction
by: He, Mutian, et al.
Published: (2025)
by: He, Mutian, et al.
Published: (2025)
Neural Attention Search Linear: Towards Adaptive Token-Level Hybrid Attention Models
by: Deng, Difan, et al.
Published: (2026)
by: Deng, Difan, et al.
Published: (2026)
Locret: Enhancing Eviction in Long-Context LLM Inference with Trained Retaining Heads on Consumer-Grade Devices
by: Huang, Yuxiang, et al.
Published: (2024)
by: Huang, Yuxiang, et al.
Published: (2024)
Distilling to Hybrid Attention Models via KL-Guided Layer Selection
by: Li, Yanhong, et al.
Published: (2025)
by: Li, Yanhong, et al.
Published: (2025)
A Hybrid Attention Framework for Fake News Detection with Large Language Models
by: Xu, Xiaochuan, et al.
Published: (2025)
by: Xu, Xiaochuan, et al.
Published: (2025)
Attention2Probability: Attention-Driven Terminology Probability Estimation for Robust Speech-to-Text System
by: Du, Yanfan, et al.
Published: (2025)
by: Du, Yanfan, et al.
Published: (2025)
Hybrid Alignment Training for Large Language Models
by: Wang, Chenglong, et al.
Published: (2024)
by: Wang, Chenglong, et al.
Published: (2024)
Octopus v2: On-device language model for super agent
by: Chen, Wei, et al.
Published: (2024)
by: Chen, Wei, et al.
Published: (2024)
Octopus v4: Graph of language models
by: Chen, Wei, et al.
Published: (2024)
by: Chen, Wei, et al.
Published: (2024)
Monocle: Hybrid Local-Global In-Context Evaluation for Long-Text Generation with Uncertainty-Based Active Learning
by: Wang, Xiaorong, et al.
Published: (2025)
by: Wang, Xiaorong, et al.
Published: (2025)
Tiny Recursive Reasoning with Mamba-2 Attention Hybrid
by: Wang, Wenlong, et al.
Published: (2026)
by: Wang, Wenlong, et al.
Published: (2026)
NOSA: Native and Offloadable Sparse Attention
by: Huang, Yuxiang, et al.
Published: (2025)
by: Huang, Yuxiang, et al.
Published: (2025)
Full Attention Strikes Back: Transferring Full Attention into Sparse within Hundred Training Steps
by: Zhou, Yanke, et al.
Published: (2026)
by: Zhou, Yanke, et al.
Published: (2026)
Similar Items
-
State Tuning: State-based Test-Time Scaling on RWKV-7
by: Xiao, Liu, et al.
Published: (2025) -
Cross-attention for State-based model RWKV-7
by: Xiao, Liu, et al.
Published: (2025) -
Millions of States: Designing a Scalable MoE Architecture with RWKV-7 Meta-learner
by: Xiao, Liu, et al.
Published: (2025) -
ARWKV: Pretrain is not what we need, an RNN-Attention-Based Language Model Born from Transformer
by: Yueyu, Lin, et al.
Published: (2025) -
Learning to Focus: Causal Attention Distillation via Gradient-Guided Token Pruning
by: Guo, Yiju, et al.
Published: (2025)