Making Every Head Count: Sparse Attention Without the Speed-Performance Trade-off
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhao, Mingkuan, Hu, Wentao, Wang, Jiayin, Lai, Xin, Huang, Tianchen, Min, Yuheng, Yan, Rui, Zhu, Xiaoyan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Fast Quiet-STaR: Thinking Without Thought Tokens
por: Huang, Wei, et al.
Publicado: (2025)
por: Huang, Wei, et al.
Publicado: (2025)
Train-Attention: Meta-Learning Where to Focus in Continual Knowledge Learning
por: Seo, Yeongbin, et al.
Publicado: (2024)
por: Seo, Yeongbin, et al.
Publicado: (2024)
WeDLM: Reconciling Diffusion Language Models with Standard Causal Attention for Fast Inference
por: Liu, Aiwei, et al.
Publicado: (2025)
por: Liu, Aiwei, et al.
Publicado: (2025)
D-SMART: Enhancing LLM Dialogue Consistency via Dynamic Structured Memory And Reasoning Tree
por: Lei, Xiang, et al.
Publicado: (2025)
por: Lei, Xiang, et al.
Publicado: (2025)
From Brazilian Portuguese to European Portuguese
por: Sanches, João, et al.
Publicado: (2024)
por: Sanches, João, et al.
Publicado: (2024)
Towards Effective and Efficient Continual Pre-training of Large Language Models
por: Chen, Jie, et al.
Publicado: (2024)
por: Chen, Jie, et al.
Publicado: (2024)
When Does Content-Based Routing Work? Representation Requirements for Selective Attention in Hybrid Sequence Models
por: Basu, Abhinaba
Publicado: (2026)
por: Basu, Abhinaba
Publicado: (2026)
Prompt Engineering and the Effectiveness of Large Language Models in Enhancing Human Productivity
por: Anam, Rizal Khoirul
Publicado: (2025)
por: Anam, Rizal Khoirul
Publicado: (2025)
Transactional Attention: Semantic Sponsorship for KV-Cache Retention
por: Basu, Abhinaba
Publicado: (2026)
por: Basu, Abhinaba
Publicado: (2026)
How Pruning Reshapes Features: Sparse Autoencoder Analysis of Weight-Pruned Language Models
por: Borobia, Hector, et al.
Publicado: (2026)
por: Borobia, Hector, et al.
Publicado: (2026)
CLMN: Concept based Language Models via Neural Symbolic Reasoning
por: Yang, Yibo
Publicado: (2025)
por: Yang, Yibo
Publicado: (2025)
Softmax Linear Attention: Reclaiming Global Competition
por: Xu, Mingwei, et al.
Publicado: (2026)
por: Xu, Mingwei, et al.
Publicado: (2026)
Co-NAML-LSTUR: A Combined Model with Attentive Multi-View Learning and Long- and Short-term User Representations for News Recommendation
por: Nguyen, Minh Hoang, et al.
Publicado: (2025)
por: Nguyen, Minh Hoang, et al.
Publicado: (2025)
Fact Grounded Attention: Eliminating Hallucination in Large Language Models Through Attention Level Knowledge Integration
por: Gupta, Aayush
Publicado: (2025)
por: Gupta, Aayush
Publicado: (2025)
Align-to-Distill: Trainable Attention Alignment for Knowledge Distillation in Neural Machine Translation
por: Jin, Heegon, et al.
Publicado: (2024)
por: Jin, Heegon, et al.
Publicado: (2024)
Towards Probabilistic Question Answering Over Tabular Data
por: Shen, Chen, et al.
Publicado: (2025)
por: Shen, Chen, et al.
Publicado: (2025)
Mitigating LLM Hallucinations through Domain-Grounded Tiered Retrieval
por: Haque, Md. Asraful, et al.
Publicado: (2026)
por: Haque, Md. Asraful, et al.
Publicado: (2026)
The Knesset Corpus: An Annotated Corpus of Hebrew Parliamentary Proceedings
por: Goldin, Gili, et al.
Publicado: (2024)
por: Goldin, Gili, et al.
Publicado: (2024)
Math Natural Language Inference: this should be easy!
por: de Paiva, Valeria, et al.
Publicado: (2025)
por: de Paiva, Valeria, et al.
Publicado: (2025)
New Skills or Sharper Primitives? A Probabilistic Perspective on the Emergence of Reasoning in RLVR
por: Wang, Zhilin, et al.
Publicado: (2026)
por: Wang, Zhilin, et al.
Publicado: (2026)
Pitfalls in Evaluating Interpretability Agents
por: Haklay, Tal, et al.
Publicado: (2026)
por: Haklay, Tal, et al.
Publicado: (2026)
d-TreeRPO: Towards More Reliable Policy Optimization for Diffusion Language Models
por: Pan, Leyi, et al.
Publicado: (2025)
por: Pan, Leyi, et al.
Publicado: (2025)
An Unforgeable Publicly Verifiable Watermark for Large Language Models
por: Liu, Aiwei, et al.
Publicado: (2023)
por: Liu, Aiwei, et al.
Publicado: (2023)
The Unlikely Duel: Evaluating Creative Writing in LLMs through a Unique Scenario
por: Gómez-Rodríguez, Carlos, et al.
Publicado: (2024)
por: Gómez-Rodríguez, Carlos, et al.
Publicado: (2024)
Direct Large Language Model Alignment Through Self-Rewarding Contrastive Prompt Distillation
por: Liu, Aiwei, et al.
Publicado: (2024)
por: Liu, Aiwei, et al.
Publicado: (2024)
SentiCSE: A Sentiment-aware Contrastive Sentence Embedding Framework with Sentiment-guided Textual Similarity
por: Kim, Jaemin, et al.
Publicado: (2024)
por: Kim, Jaemin, et al.
Publicado: (2024)
Distractor Injection Attacks on Large Reasoning Models: Characterization and Defense
por: Zhang, Zhehao, et al.
Publicado: (2025)
por: Zhang, Zhehao, et al.
Publicado: (2025)
Exploiting Pre-trained Encoder-Decoder Transformers for Sequence-to-Sequence Constituent Parsing
por: Fernández-González, Daniel, et al.
Publicado: (2026)
por: Fernández-González, Daniel, et al.
Publicado: (2026)
Parametric Social Identity Injection and Diversification in Public Opinion Simulation
por: Wang, Hexi, et al.
Publicado: (2026)
por: Wang, Hexi, et al.
Publicado: (2026)
Trusted Uncertainty in Large Language Models: A Unified Framework for Confidence Calibration and Risk-Controlled Refusal
por: Oehri, Markus, et al.
Publicado: (2025)
por: Oehri, Markus, et al.
Publicado: (2025)
Beyond Cosine Similarity
por: Ai, Xinbo
Publicado: (2026)
por: Ai, Xinbo
Publicado: (2026)
AI-assisted German Employment Contract Review: A Benchmark Dataset
por: Wardas, Oliver, et al.
Publicado: (2025)
por: Wardas, Oliver, et al.
Publicado: (2025)
The Superalignment of Superhuman Intelligence with Large Language Models
por: Huang, Minlie, et al.
Publicado: (2024)
por: Huang, Minlie, et al.
Publicado: (2024)
ScoreRAG: A Retrieval-Augmented Generation Framework with Consistency-Relevance Scoring and Structured Summarization for News Generation
por: Lin, Pei-Yun, et al.
Publicado: (2025)
por: Lin, Pei-Yun, et al.
Publicado: (2025)
Unifying Uniform and Binary-coding Quantization for Accurate Compression of Large Language Models
por: Park, Seungcheol, et al.
Publicado: (2025)
por: Park, Seungcheol, et al.
Publicado: (2025)
When Retrieval Succeeds and Fails: Rethinking Retrieval-Augmented Generation for LLMs
por: Wang, Yongjie, et al.
Publicado: (2025)
por: Wang, Yongjie, et al.
Publicado: (2025)
GATE: Graph-based Adaptive Tool Evolution Across Diverse Tasks
por: Luo, Jianwen, et al.
Publicado: (2025)
por: Luo, Jianwen, et al.
Publicado: (2025)
Omni-SafetyBench: A Benchmark for Safety Evaluation of Audio-Visual Large Language Models
por: Pan, Leyi, et al.
Publicado: (2025)
por: Pan, Leyi, et al.
Publicado: (2025)
Accurate Retraining-free Pruning for Pretrained Encoder-based Language Models
por: Park, Seungcheol, et al.
Publicado: (2023)
por: Park, Seungcheol, et al.
Publicado: (2023)
Compression Method Matters: Benchmark-Dependent Output Dynamics in LLM Prompt Compression
por: Johnson, Warren
Publicado: (2026)
por: Johnson, Warren
Publicado: (2026)
Ejemplares similares
-
Fast Quiet-STaR: Thinking Without Thought Tokens
por: Huang, Wei, et al.
Publicado: (2025) -
Train-Attention: Meta-Learning Where to Focus in Continual Knowledge Learning
por: Seo, Yeongbin, et al.
Publicado: (2024) -
WeDLM: Reconciling Diffusion Language Models with Standard Causal Attention for Fast Inference
por: Liu, Aiwei, et al.
Publicado: (2025) -
D-SMART: Enhancing LLM Dialogue Consistency via Dynamic Structured Memory And Reasoning Tree
por: Lei, Xiang, et al.
Publicado: (2025) -
From Brazilian Portuguese to European Portuguese
por: Sanches, João, et al.
Publicado: (2024)