Top-Theta Attention: Sparsifying Transformers by Compensated Thresholding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Berestizshevsky, Konstantin, Andri, Renzo, Cavigelli, Lukas |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Compressible Softmax-Attended Language under Incompressible Attention
von: Lee, Wonsuk
Veröffentlicht: (2026)
von: Lee, Wonsuk
Veröffentlicht: (2026)
Transformadores: Fundamentos teoricos y Aplicaciones
von: de la Torre, Jordi
Veröffentlicht: (2023)
von: de la Torre, Jordi
Veröffentlicht: (2023)
Can Out-of-Distribution Evaluations Uncover Reliance on Shortcuts? A Case Study in Question Answering
von: Štefánik, Michal, et al.
Veröffentlicht: (2025)
von: Štefánik, Michal, et al.
Veröffentlicht: (2025)
SATA-BENCH: Select All That Apply Benchmark for Multiple Choice Questions
von: Xu, Weijie, et al.
Veröffentlicht: (2025)
von: Xu, Weijie, et al.
Veröffentlicht: (2025)
Controlling Long-Horizon Behavior in Language Model Agents with Explicit State Dynamics
von: Subaharan, Sukesh
Veröffentlicht: (2026)
von: Subaharan, Sukesh
Veröffentlicht: (2026)
An Automatic Text Classification Method Based on Hierarchical Taxonomies, Neural Networks and Document Embedding: The NETHIC Tool
von: Lomasto, Luigi, et al.
Veröffentlicht: (2026)
von: Lomasto, Luigi, et al.
Veröffentlicht: (2026)
Data and AI governance: Promoting equity, ethics, and fairness in large language models
von: Abhishek, Alok, et al.
Veröffentlicht: (2025)
von: Abhishek, Alok, et al.
Veröffentlicht: (2025)
BEATS: Bias Evaluation and Assessment Test Suite for Large Language Models
von: Abhishek, Alok, et al.
Veröffentlicht: (2025)
von: Abhishek, Alok, et al.
Veröffentlicht: (2025)
SHARP: Social Harm Analysis via Risk Profiles for Measuring Inequities in Large Language Models
von: Abhishek, Alok, et al.
Veröffentlicht: (2026)
von: Abhishek, Alok, et al.
Veröffentlicht: (2026)
TSDS: Data Selection for Task-Specific Model Finetuning
von: Liu, Zifan, et al.
Veröffentlicht: (2024)
von: Liu, Zifan, et al.
Veröffentlicht: (2024)
Reasoning Promotes Robustness in Theory of Mind Tasks
von: de Haan, Ian B., et al.
Veröffentlicht: (2026)
von: de Haan, Ian B., et al.
Veröffentlicht: (2026)
ConSensus: Multi-Agent Collaboration for Multimodal Sensing
von: Yoon, Hyungjun, et al.
Veröffentlicht: (2026)
von: Yoon, Hyungjun, et al.
Veröffentlicht: (2026)
Morphological Synthesizer for Ge'ez Language: Addressing Morphological Complexity and Resource Limitations
von: Gebremariam, Gebrearegawi, et al.
Veröffentlicht: (2025)
von: Gebremariam, Gebrearegawi, et al.
Veröffentlicht: (2025)
Large Language Models are Inconsistent and Biased Evaluators
von: Stureborg, Rickard, et al.
Veröffentlicht: (2024)
von: Stureborg, Rickard, et al.
Veröffentlicht: (2024)
On the Invariants of Softmax Attention
von: Lee, Wonsuk
Veröffentlicht: (2026)
von: Lee, Wonsuk
Veröffentlicht: (2026)
Attention Please: What Transformer Models Really Learn for Process Prediction
von: Käppel, Martin, et al.
Veröffentlicht: (2024)
von: Käppel, Martin, et al.
Veröffentlicht: (2024)
On Adversarial Examples for Text Classification by Perturbing Latent Representations
von: Sooksatra, Korn, et al.
Veröffentlicht: (2024)
von: Sooksatra, Korn, et al.
Veröffentlicht: (2024)
Tailoring Vaccine Messaging with Common-Ground Opinions
von: Stureborg, Rickard, et al.
Veröffentlicht: (2024)
von: Stureborg, Rickard, et al.
Veröffentlicht: (2024)
An Epidemiological Knowledge Graph extracted from the World Health Organization's Disease Outbreak News
von: Consoli, Sergio, et al.
Veröffentlicht: (2025)
von: Consoli, Sergio, et al.
Veröffentlicht: (2025)
HAAP: Vision-context Hierarchical Attention Autoregressive with Adaptive Permutation for Scene Text Recognition
von: Chen, Honghui, et al.
Veröffentlicht: (2024)
von: Chen, Honghui, et al.
Veröffentlicht: (2024)
Instilling Organisational Values in Firefighters through Simulation-Based Training
von: Osman, Nardine, et al.
Veröffentlicht: (2025)
von: Osman, Nardine, et al.
Veröffentlicht: (2025)
Feature Relevancy, Necessity and Usefulness: Complexity and Algorithms
von: Capdevielle, Tomás, et al.
Veröffentlicht: (2025)
von: Capdevielle, Tomás, et al.
Veröffentlicht: (2025)
A Taxonomy of Omnicidal Futures Involving Artificial Intelligence
von: Critch, Andrew, et al.
Veröffentlicht: (2025)
von: Critch, Andrew, et al.
Veröffentlicht: (2025)
A Theoretical Framework for Adaptive Utility-Weighted Benchmarking
von: Waggoner, Philip
Veröffentlicht: (2026)
von: Waggoner, Philip
Veröffentlicht: (2026)
From Language Models to Practical Self-Improving Computer Agents
von: Sheng, Alex
Veröffentlicht: (2024)
von: Sheng, Alex
Veröffentlicht: (2024)
Hilbert-Geo: Solving Solid Geometric Problems by Neural-Symbolic Reasoning
von: Xu, Ruoran, et al.
Veröffentlicht: (2026)
von: Xu, Ruoran, et al.
Veröffentlicht: (2026)
Revenue-Sharing as Infrastructure: A Distributed Business Model for Generative AI Platforms
von: Mondjo, Ghislain Dorian Tchuente
Veröffentlicht: (2026)
von: Mondjo, Ghislain Dorian Tchuente
Veröffentlicht: (2026)
Associative memory inspires improvements for in-context learning using a novel attention residual stream architecture
von: Burns, Thomas F, et al.
Veröffentlicht: (2024)
von: Burns, Thomas F, et al.
Veröffentlicht: (2024)
CAG: Chunked Augmented Generation for Google Chrome's Built-in Gemini Nano
von: Surulimuthu, Vivek Vellaiyappan, et al.
Veröffentlicht: (2024)
von: Surulimuthu, Vivek Vellaiyappan, et al.
Veröffentlicht: (2024)
Graph Transformers: A Survey
von: Shehzad, Ahsan, et al.
Veröffentlicht: (2024)
von: Shehzad, Ahsan, et al.
Veröffentlicht: (2024)
Abductive explanations of classifiers under constraints: Complexity and properties
von: Cooper, Martin, et al.
Veröffentlicht: (2024)
von: Cooper, Martin, et al.
Veröffentlicht: (2024)
AI Education in Higher Education: A Taxonomy for Curriculum Reform and the Mission of Knowledge
von: Zheng, Tian
Veröffentlicht: (2025)
von: Zheng, Tian
Veröffentlicht: (2025)
Reducing Instability in Synthetic Data Evaluation with a Super-Metric in MalDataGen
von: da Silva, Anna Luiza Gomes, et al.
Veröffentlicht: (2025)
von: da Silva, Anna Luiza Gomes, et al.
Veröffentlicht: (2025)
AI Consciousness is Inevitable: A Theoretical Computer Science Perspective
von: Blum, Lenore, et al.
Veröffentlicht: (2024)
von: Blum, Lenore, et al.
Veröffentlicht: (2024)
Intelligence as Computation
von: Brock, Oliver
Veröffentlicht: (2024)
von: Brock, Oliver
Veröffentlicht: (2024)
Quantifying Behavioral Dissimilarity Between Mathematical Expressions
von: Mežnar, Sebastian, et al.
Veröffentlicht: (2024)
von: Mežnar, Sebastian, et al.
Veröffentlicht: (2024)
ORPHEAS: A Cross-Lingual Greek-English Embedding Model for Retrieval-Augmented Generation
von: Livieris, Ioannis E., et al.
Veröffentlicht: (2026)
von: Livieris, Ioannis E., et al.
Veröffentlicht: (2026)
On Privacy Leakage in Tabular Diffusion Models: Influential Factors, Attacker Knowledge, and Metrics
von: Shafieinejad, Masoumeh, et al.
Veröffentlicht: (2026)
von: Shafieinejad, Masoumeh, et al.
Veröffentlicht: (2026)
On measuring grounding and generalizing grounding problems
von: Quigley, Daniel, et al.
Veröffentlicht: (2025)
von: Quigley, Daniel, et al.
Veröffentlicht: (2025)
Project Synapse: A Hierarchical Multi-Agent Framework with Hybrid Memory for Autonomous Resolution of Last-Mile Delivery Disruptions
von: Yadav, Arin Gopalan, et al.
Veröffentlicht: (2026)
von: Yadav, Arin Gopalan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Compressible Softmax-Attended Language under Incompressible Attention
von: Lee, Wonsuk
Veröffentlicht: (2026) -
Transformadores: Fundamentos teoricos y Aplicaciones
von: de la Torre, Jordi
Veröffentlicht: (2023) -
Can Out-of-Distribution Evaluations Uncover Reliance on Shortcuts? A Case Study in Question Answering
von: Štefánik, Michal, et al.
Veröffentlicht: (2025) -
SATA-BENCH: Select All That Apply Benchmark for Multiple Choice Questions
von: Xu, Weijie, et al.
Veröffentlicht: (2025) -
Controlling Long-Horizon Behavior in Language Model Agents with Explicit State Dynamics
von: Subaharan, Sukesh
Veröffentlicht: (2026)