QiMeng-Attention: SOTA Attention Operator is generated by SOTA Attention Algorithm
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhou, Qirui, Peng, Shaohui, Xiong, Weiqiang, Chen, Haixin, Wen, Yuanbo, Li, Haochen, Li, Ling, Guo, Qi, Zhao, Yongwei, Gao, Ke, Chen, Ruizhi, Wu, Yanjun, Zhao, Chen, Chen, Yunji |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
QiMeng-TensorOp: Automatically Generating High-Performance Tensor Operators with Hardware Primitives
di: Zhang, Xuzhi, et al.
Pubblicazione: (2025)
di: Zhang, Xuzhi, et al.
Pubblicazione: (2025)
QiMeng-Kernel: Macro-Thinking Micro-Coding Paradigm for LLM-Based High-Performance GPU Kernel Generation
di: Zhu, Xinguo, et al.
Pubblicazione: (2025)
di: Zhu, Xinguo, et al.
Pubblicazione: (2025)
QiMeng: Fully Automated Hardware and Software Design for Processor Chip
di: Zhang, Rui, et al.
Pubblicazione: (2025)
di: Zhang, Rui, et al.
Pubblicazione: (2025)
QiMeng-MuPa: Mutual-Supervised Learning for Sequential-to-Parallel Code Translation
di: Ke, Changxin, et al.
Pubblicazione: (2025)
di: Ke, Changxin, et al.
Pubblicazione: (2025)
QiMeng-PRepair: Precise Code Repair via Edit-Aware Reward Optimization
di: Ke, Changxin, et al.
Pubblicazione: (2026)
di: Ke, Changxin, et al.
Pubblicazione: (2026)
QiMeng-CPU-v2: Automated Superscalar Processor Design by Learning Data Dependencies
di: Cheng, Shuyao, et al.
Pubblicazione: (2025)
di: Cheng, Shuyao, et al.
Pubblicazione: (2025)
QiMeng-SALV: Signal-Aware Learning for Verilog Code Generation
di: Zhang, Yang, et al.
Pubblicazione: (2025)
di: Zhang, Yang, et al.
Pubblicazione: (2025)
QiMeng-NeuComBack: Self-Evolving Translation from IR to Assembly Code
di: Fang, Hainan, et al.
Pubblicazione: (2025)
di: Fang, Hainan, et al.
Pubblicazione: (2025)
QiMeng-Xpiler: Transcompiling Tensor Programs for Deep Learning Systems with a Neural-Symbolic Approach
di: Dong, Shouyang, et al.
Pubblicazione: (2025)
di: Dong, Shouyang, et al.
Pubblicazione: (2025)
QiMeng-CodeV-R1: Reasoning-Enhanced Verilog Generation
di: Zhu, Yaoyu, et al.
Pubblicazione: (2025)
di: Zhu, Yaoyu, et al.
Pubblicazione: (2025)
QiMeng-CRUX: Narrowing the Gap Between Natural Language and Verilog via Core Refined Understanding eXpression for Circuit Design
di: Huang, Lei, et al.
Pubblicazione: (2025)
di: Huang, Lei, et al.
Pubblicazione: (2025)
An Experimental Study of SOTA LiDAR Segmentation Models
di: Chen, Bike, et al.
Pubblicazione: (2025)
di: Chen, Bike, et al.
Pubblicazione: (2025)
Resonant Context Anchoring: Decoupling Attention Routing and Signal Gain at Inference Time
di: Zhao, Mingkuan, et al.
Pubblicazione: (2026)
di: Zhao, Mingkuan, et al.
Pubblicazione: (2026)
The Hidden Attention of Mamba Models
di: Ali, Ameen, et al.
Pubblicazione: (2024)
di: Ali, Ameen, et al.
Pubblicazione: (2024)
Jasper and Stella: distillation of SOTA embedding models
di: Zhang, Dun, et al.
Pubblicazione: (2024)
di: Zhang, Dun, et al.
Pubblicazione: (2024)
Structured-Sparse Attention for Entity Tracking with Subquadratic Sequence Complexity
di: Zhao, Hangyue, et al.
Pubblicazione: (2026)
di: Zhao, Hangyue, et al.
Pubblicazione: (2026)
Turbo Sparse: Achieving LLM SOTA Performance with Minimal Activated Parameters
di: Song, Yixin, et al.
Pubblicazione: (2024)
di: Song, Yixin, et al.
Pubblicazione: (2024)
RADLADS: Rapid Attention Distillation to Linear Attention Decoders at Scale
di: Goldstein, Daniel, et al.
Pubblicazione: (2025)
di: Goldstein, Daniel, et al.
Pubblicazione: (2025)
Text Summarization With Graph Attention Networks
di: Ardestani, Mohammadreza, et al.
Pubblicazione: (2026)
di: Ardestani, Mohammadreza, et al.
Pubblicazione: (2026)
ATTNChecker: Highly-Optimized Fault Tolerant Attention for Large Language Model Training
di: Liang, Yuhang, et al.
Pubblicazione: (2024)
di: Liang, Yuhang, et al.
Pubblicazione: (2024)
Forget Attention: Importance-Aware Attention Is All You Need
di: Shin, Soohyeong, et al.
Pubblicazione: (2026)
di: Shin, Soohyeong, et al.
Pubblicazione: (2026)
Overcoming Long-Context Limitations of State-Space Models via Context-Dependent Sparse Attention
di: Zhan, Zhihao, et al.
Pubblicazione: (2025)
di: Zhan, Zhihao, et al.
Pubblicazione: (2025)
Lean Attention: Hardware-Aware Scalable Attention Mechanism for the Decode-Phase of Transformers
di: Sanovar, Rya, et al.
Pubblicazione: (2024)
di: Sanovar, Rya, et al.
Pubblicazione: (2024)
QiMeng-CodeV-SVA: Training Specialized LLMs for Hardware Assertion Generation via RTL-Grounded Bidirectional Data Synthesis
di: Wu, Yutong, et al.
Pubblicazione: (2026)
di: Wu, Yutong, et al.
Pubblicazione: (2026)
OnlySportsLM: Optimizing Sports-Domain Language Models with SOTA Performance under Billion Parameters
di: Chen, Zexin, et al.
Pubblicazione: (2024)
di: Chen, Zexin, et al.
Pubblicazione: (2024)
Hierarchical Shift Mixing -- Beyond Dense Attention in Transformers
di: Forchheimer, Robert
Pubblicazione: (2026)
di: Forchheimer, Robert
Pubblicazione: (2026)
Mixture of Attention Spans: Optimizing LLM Inference Efficiency with Heterogeneous Sliding-Window Lengths
di: Fu, Tianyu, et al.
Pubblicazione: (2024)
di: Fu, Tianyu, et al.
Pubblicazione: (2024)
Chain and Causal Attention for Efficient Entity Tracking
di: Fagnou, Erwan, et al.
Pubblicazione: (2024)
di: Fagnou, Erwan, et al.
Pubblicazione: (2024)
On Explaining with Attention Matrices
di: Naim, Omar, et al.
Pubblicazione: (2024)
di: Naim, Omar, et al.
Pubblicazione: (2024)
How LLMs Are Persuaded: A Few Attention Heads, Rerouted
di: Sun, Xiangkun, et al.
Pubblicazione: (2026)
di: Sun, Xiangkun, et al.
Pubblicazione: (2026)
Making Every Head Count: Sparse Attention Without the Speed-Performance Trade-off
di: Zhao, Mingkuan, et al.
Pubblicazione: (2025)
di: Zhao, Mingkuan, et al.
Pubblicazione: (2025)
Beyond Prefixes: Graph-as-Memory Cross-Attention for Knowledge Graph Completion with Large Language Models
di: Liu, Ruitong, et al.
Pubblicazione: (2025)
di: Liu, Ruitong, et al.
Pubblicazione: (2025)
Transparency Distortion Robustness for SOTA Image Segmentation Tasks
di: Knauthe, Volker, et al.
Pubblicazione: (2024)
di: Knauthe, Volker, et al.
Pubblicazione: (2024)
Sentinel: SOTA model to protect against prompt injections
di: Ivry, Dror, et al.
Pubblicazione: (2025)
di: Ivry, Dror, et al.
Pubblicazione: (2025)
ELÍAS RAMÓN DE LA SOTA (1932-2014)
di: Esteban Ismael Meza Torres
Pubblicazione: (2014)
di: Esteban Ismael Meza Torres
Pubblicazione: (2014)
Softmax Linear Attention: Reclaiming Global Competition
di: Xu, Mingwei, et al.
Pubblicazione: (2026)
di: Xu, Mingwei, et al.
Pubblicazione: (2026)
System-Mediated Attention Imbalances Make Vision-Language Models Say Yes
di: Chan, Tsan Tsai, et al.
Pubblicazione: (2026)
di: Chan, Tsan Tsai, et al.
Pubblicazione: (2026)
Detecting Hallucinations in Graph Retrieval-Augmented Generation via Attention Patterns and Semantic Alignment
di: Li, Shanghao, et al.
Pubblicazione: (2025)
di: Li, Shanghao, et al.
Pubblicazione: (2025)
Attention Drift: What Autoregressive Speculative Decoding Models Learn
di: Eldenk, Doğaç, et al.
Pubblicazione: (2026)
di: Eldenk, Doğaç, et al.
Pubblicazione: (2026)
$\rm SP^3$: Enhancing Structured Pruning via PCA Projection
di: Hu, Yuxuan, et al.
Pubblicazione: (2023)
di: Hu, Yuxuan, et al.
Pubblicazione: (2023)
Documenti analoghi
-
QiMeng-TensorOp: Automatically Generating High-Performance Tensor Operators with Hardware Primitives
di: Zhang, Xuzhi, et al.
Pubblicazione: (2025) -
QiMeng-Kernel: Macro-Thinking Micro-Coding Paradigm for LLM-Based High-Performance GPU Kernel Generation
di: Zhu, Xinguo, et al.
Pubblicazione: (2025) -
QiMeng: Fully Automated Hardware and Software Design for Processor Chip
di: Zhang, Rui, et al.
Pubblicazione: (2025) -
QiMeng-MuPa: Mutual-Supervised Learning for Sequential-to-Parallel Code Translation
di: Ke, Changxin, et al.
Pubblicazione: (2025) -
QiMeng-PRepair: Precise Code Repair via Edit-Aware Reward Optimization
di: Ke, Changxin, et al.
Pubblicazione: (2026)