QiMeng-Attention: SOTA Attention Operator is generated by SOTA Attention Algorithm
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Qirui, Peng, Shaohui, Xiong, Weiqiang, Chen, Haixin, Wen, Yuanbo, Li, Haochen, Li, Ling, Guo, Qi, Zhao, Yongwei, Gao, Ke, Chen, Ruizhi, Wu, Yanjun, Zhao, Chen, Chen, Yunji |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
QiMeng-TensorOp: Automatically Generating High-Performance Tensor Operators with Hardware Primitives
von: Zhang, Xuzhi, et al.
Veröffentlicht: (2025)
von: Zhang, Xuzhi, et al.
Veröffentlicht: (2025)
QiMeng-Kernel: Macro-Thinking Micro-Coding Paradigm for LLM-Based High-Performance GPU Kernel Generation
von: Zhu, Xinguo, et al.
Veröffentlicht: (2025)
von: Zhu, Xinguo, et al.
Veröffentlicht: (2025)
QiMeng: Fully Automated Hardware and Software Design for Processor Chip
von: Zhang, Rui, et al.
Veröffentlicht: (2025)
von: Zhang, Rui, et al.
Veröffentlicht: (2025)
QiMeng-MuPa: Mutual-Supervised Learning for Sequential-to-Parallel Code Translation
von: Ke, Changxin, et al.
Veröffentlicht: (2025)
von: Ke, Changxin, et al.
Veröffentlicht: (2025)
QiMeng-PRepair: Precise Code Repair via Edit-Aware Reward Optimization
von: Ke, Changxin, et al.
Veröffentlicht: (2026)
von: Ke, Changxin, et al.
Veröffentlicht: (2026)
QiMeng-CPU-v2: Automated Superscalar Processor Design by Learning Data Dependencies
von: Cheng, Shuyao, et al.
Veröffentlicht: (2025)
von: Cheng, Shuyao, et al.
Veröffentlicht: (2025)
QiMeng-SALV: Signal-Aware Learning for Verilog Code Generation
von: Zhang, Yang, et al.
Veröffentlicht: (2025)
von: Zhang, Yang, et al.
Veröffentlicht: (2025)
QiMeng-NeuComBack: Self-Evolving Translation from IR to Assembly Code
von: Fang, Hainan, et al.
Veröffentlicht: (2025)
von: Fang, Hainan, et al.
Veröffentlicht: (2025)
QiMeng-Xpiler: Transcompiling Tensor Programs for Deep Learning Systems with a Neural-Symbolic Approach
von: Dong, Shouyang, et al.
Veröffentlicht: (2025)
von: Dong, Shouyang, et al.
Veröffentlicht: (2025)
QiMeng-CodeV-R1: Reasoning-Enhanced Verilog Generation
von: Zhu, Yaoyu, et al.
Veröffentlicht: (2025)
von: Zhu, Yaoyu, et al.
Veröffentlicht: (2025)
QiMeng-CRUX: Narrowing the Gap Between Natural Language and Verilog via Core Refined Understanding eXpression for Circuit Design
von: Huang, Lei, et al.
Veröffentlicht: (2025)
von: Huang, Lei, et al.
Veröffentlicht: (2025)
An Experimental Study of SOTA LiDAR Segmentation Models
von: Chen, Bike, et al.
Veröffentlicht: (2025)
von: Chen, Bike, et al.
Veröffentlicht: (2025)
Resonant Context Anchoring: Decoupling Attention Routing and Signal Gain at Inference Time
von: Zhao, Mingkuan, et al.
Veröffentlicht: (2026)
von: Zhao, Mingkuan, et al.
Veröffentlicht: (2026)
The Hidden Attention of Mamba Models
von: Ali, Ameen, et al.
Veröffentlicht: (2024)
von: Ali, Ameen, et al.
Veröffentlicht: (2024)
Jasper and Stella: distillation of SOTA embedding models
von: Zhang, Dun, et al.
Veröffentlicht: (2024)
von: Zhang, Dun, et al.
Veröffentlicht: (2024)
Structured-Sparse Attention for Entity Tracking with Subquadratic Sequence Complexity
von: Zhao, Hangyue, et al.
Veröffentlicht: (2026)
von: Zhao, Hangyue, et al.
Veröffentlicht: (2026)
Turbo Sparse: Achieving LLM SOTA Performance with Minimal Activated Parameters
von: Song, Yixin, et al.
Veröffentlicht: (2024)
von: Song, Yixin, et al.
Veröffentlicht: (2024)
RADLADS: Rapid Attention Distillation to Linear Attention Decoders at Scale
von: Goldstein, Daniel, et al.
Veröffentlicht: (2025)
von: Goldstein, Daniel, et al.
Veröffentlicht: (2025)
Text Summarization With Graph Attention Networks
von: Ardestani, Mohammadreza, et al.
Veröffentlicht: (2026)
von: Ardestani, Mohammadreza, et al.
Veröffentlicht: (2026)
ATTNChecker: Highly-Optimized Fault Tolerant Attention for Large Language Model Training
von: Liang, Yuhang, et al.
Veröffentlicht: (2024)
von: Liang, Yuhang, et al.
Veröffentlicht: (2024)
Forget Attention: Importance-Aware Attention Is All You Need
von: Shin, Soohyeong, et al.
Veröffentlicht: (2026)
von: Shin, Soohyeong, et al.
Veröffentlicht: (2026)
Overcoming Long-Context Limitations of State-Space Models via Context-Dependent Sparse Attention
von: Zhan, Zhihao, et al.
Veröffentlicht: (2025)
von: Zhan, Zhihao, et al.
Veröffentlicht: (2025)
Lean Attention: Hardware-Aware Scalable Attention Mechanism for the Decode-Phase of Transformers
von: Sanovar, Rya, et al.
Veröffentlicht: (2024)
von: Sanovar, Rya, et al.
Veröffentlicht: (2024)
QiMeng-CodeV-SVA: Training Specialized LLMs for Hardware Assertion Generation via RTL-Grounded Bidirectional Data Synthesis
von: Wu, Yutong, et al.
Veröffentlicht: (2026)
von: Wu, Yutong, et al.
Veröffentlicht: (2026)
OnlySportsLM: Optimizing Sports-Domain Language Models with SOTA Performance under Billion Parameters
von: Chen, Zexin, et al.
Veröffentlicht: (2024)
von: Chen, Zexin, et al.
Veröffentlicht: (2024)
Hierarchical Shift Mixing -- Beyond Dense Attention in Transformers
von: Forchheimer, Robert
Veröffentlicht: (2026)
von: Forchheimer, Robert
Veröffentlicht: (2026)
Mixture of Attention Spans: Optimizing LLM Inference Efficiency with Heterogeneous Sliding-Window Lengths
von: Fu, Tianyu, et al.
Veröffentlicht: (2024)
von: Fu, Tianyu, et al.
Veröffentlicht: (2024)
Chain and Causal Attention for Efficient Entity Tracking
von: Fagnou, Erwan, et al.
Veröffentlicht: (2024)
von: Fagnou, Erwan, et al.
Veröffentlicht: (2024)
On Explaining with Attention Matrices
von: Naim, Omar, et al.
Veröffentlicht: (2024)
von: Naim, Omar, et al.
Veröffentlicht: (2024)
How LLMs Are Persuaded: A Few Attention Heads, Rerouted
von: Sun, Xiangkun, et al.
Veröffentlicht: (2026)
von: Sun, Xiangkun, et al.
Veröffentlicht: (2026)
Making Every Head Count: Sparse Attention Without the Speed-Performance Trade-off
von: Zhao, Mingkuan, et al.
Veröffentlicht: (2025)
von: Zhao, Mingkuan, et al.
Veröffentlicht: (2025)
Beyond Prefixes: Graph-as-Memory Cross-Attention for Knowledge Graph Completion with Large Language Models
von: Liu, Ruitong, et al.
Veröffentlicht: (2025)
von: Liu, Ruitong, et al.
Veröffentlicht: (2025)
Transparency Distortion Robustness for SOTA Image Segmentation Tasks
von: Knauthe, Volker, et al.
Veröffentlicht: (2024)
von: Knauthe, Volker, et al.
Veröffentlicht: (2024)
Sentinel: SOTA model to protect against prompt injections
von: Ivry, Dror, et al.
Veröffentlicht: (2025)
von: Ivry, Dror, et al.
Veröffentlicht: (2025)
ELÍAS RAMÓN DE LA SOTA (1932-2014)
von: Esteban Ismael Meza Torres
Veröffentlicht: (2014)
von: Esteban Ismael Meza Torres
Veröffentlicht: (2014)
Softmax Linear Attention: Reclaiming Global Competition
von: Xu, Mingwei, et al.
Veröffentlicht: (2026)
von: Xu, Mingwei, et al.
Veröffentlicht: (2026)
System-Mediated Attention Imbalances Make Vision-Language Models Say Yes
von: Chan, Tsan Tsai, et al.
Veröffentlicht: (2026)
von: Chan, Tsan Tsai, et al.
Veröffentlicht: (2026)
Detecting Hallucinations in Graph Retrieval-Augmented Generation via Attention Patterns and Semantic Alignment
von: Li, Shanghao, et al.
Veröffentlicht: (2025)
von: Li, Shanghao, et al.
Veröffentlicht: (2025)
Attention Drift: What Autoregressive Speculative Decoding Models Learn
von: Eldenk, Doğaç, et al.
Veröffentlicht: (2026)
von: Eldenk, Doğaç, et al.
Veröffentlicht: (2026)
$\rm SP^3$: Enhancing Structured Pruning via PCA Projection
von: Hu, Yuxuan, et al.
Veröffentlicht: (2023)
von: Hu, Yuxuan, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
QiMeng-TensorOp: Automatically Generating High-Performance Tensor Operators with Hardware Primitives
von: Zhang, Xuzhi, et al.
Veröffentlicht: (2025) -
QiMeng-Kernel: Macro-Thinking Micro-Coding Paradigm for LLM-Based High-Performance GPU Kernel Generation
von: Zhu, Xinguo, et al.
Veröffentlicht: (2025) -
QiMeng: Fully Automated Hardware and Software Design for Processor Chip
von: Zhang, Rui, et al.
Veröffentlicht: (2025) -
QiMeng-MuPa: Mutual-Supervised Learning for Sequential-to-Parallel Code Translation
von: Ke, Changxin, et al.
Veröffentlicht: (2025) -
QiMeng-PRepair: Precise Code Repair via Edit-Aware Reward Optimization
von: Ke, Changxin, et al.
Veröffentlicht: (2026)