Tensor Product Attention Is All You Need
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Yifan, Liu, Yifeng, Yuan, Huizhuo, Qin, Zhen, Yuan, Yang, Gu, Quanquan, Yao, Andrew Chi-Chih |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Design of KL-Regularized Policy Gradient Algorithms for LLM Reasoning
by: Zhang, Yifan, et al.
Published: (2025)
by: Zhang, Yifan, et al.
Published: (2025)
Group Representational Position Encoding
by: Zhang, Yifan, et al.
Published: (2025)
by: Zhang, Yifan, et al.
Published: (2025)
Accelerated Preference Optimization for Large Language Model Alignment
by: He, Jiafan, et al.
Published: (2024)
by: He, Jiafan, et al.
Published: (2024)
Higher-order Linear Attention
by: Zhang, Yifan, et al.
Published: (2025)
by: Zhang, Yifan, et al.
Published: (2025)
Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
by: Chen, Zixiang, et al.
Published: (2024)
by: Chen, Zixiang, et al.
Published: (2024)
Self-Play Fine-Tuning of Diffusion Models for Text-to-Image Generation
by: Yuan, Huizhuo, et al.
Published: (2024)
by: Yuan, Huizhuo, et al.
Published: (2024)
Self-Play Preference Optimization for Language Model Alignment
by: Wu, Yue, et al.
Published: (2024)
by: Wu, Yue, et al.
Published: (2024)
On the Diagram of Thought
by: Zhang, Yifan, et al.
Published: (2024)
by: Zhang, Yifan, et al.
Published: (2024)
Attention Smoothing Is All You Need For Unlearning
by: Zade, Saleh Zare, et al.
Published: (2026)
by: Zade, Saleh Zare, et al.
Published: (2026)
Causal Attention with Lookahead Keys
by: Song, Zhuoqing, et al.
Published: (2025)
by: Song, Zhuoqing, et al.
Published: (2025)
Deep Delta Learning
by: Zhang, Yifan, et al.
Published: (2026)
by: Zhang, Yifan, et al.
Published: (2026)
More Agents Is All You Need
by: Li, Junyou, et al.
Published: (2024)
by: Li, Junyou, et al.
Published: (2024)
Meta Prompting for AI Systems
by: Zhang, Yifan, et al.
Published: (2023)
by: Zhang, Yifan, et al.
Published: (2023)
Augmenting Math Word Problems via Iterative Question Composing
by: Liu, Haoxiong, et al.
Published: (2024)
by: Liu, Haoxiong, et al.
Published: (2024)
Attention Is All You Need for KV Cache in Diffusion LLMs
by: Nguyen-Tri, Quan, et al.
Published: (2025)
by: Nguyen-Tri, Quan, et al.
Published: (2025)
Synthetic Data RL: Task Definition Is All You Need
by: Guo, Yiduo, et al.
Published: (2025)
by: Guo, Yiduo, et al.
Published: (2025)
Fast Sampling via Discrete Non-Markov Diffusion Models with Predetermined Transition Time
by: Chen, Zixiang, et al.
Published: (2023)
by: Chen, Zixiang, et al.
Published: (2023)
Forget Attention: Importance-Aware Attention Is All You Need
by: Shin, Soohyeong, et al.
Published: (2026)
by: Shin, Soohyeong, et al.
Published: (2026)
Similarity is Not All You Need: Endowing Retrieval Augmented Generation with Multi Layered Thoughts
by: Gan, Chunjing, et al.
Published: (2024)
by: Gan, Chunjing, et al.
Published: (2024)
RSPO: Regularized Self-Play Alignment of Large Language Models
by: Tang, Xiaohang, et al.
Published: (2025)
by: Tang, Xiaohang, et al.
Published: (2025)
Autonomous Data Selection with Zero-shot Generative Classifiers for Mathematical Texts
by: Zhang, Yifan, et al.
Published: (2024)
by: Zhang, Yifan, et al.
Published: (2024)
Monadic Context Engineering
by: Zhang, Yifan, et al.
Published: (2025)
by: Zhang, Yifan, et al.
Published: (2025)
More Compute Is What You Need
by: Guo, Zhen
Published: (2024)
by: Guo, Zhen
Published: (2024)
All You Need is One: Capsule Prompt Tuning with a Single Vector
by: Liu, Yiyang, et al.
Published: (2025)
by: Liu, Yiyang, et al.
Published: (2025)
Beyond Bradley-Terry Models: A General Preference Model for Language Model Alignment
by: Zhang, Yifan, et al.
Published: (2024)
by: Zhang, Yifan, et al.
Published: (2024)
Attention is All You Need Until You Need Retention
by: Yaslioglu, M. Murat
Published: (2025)
by: Yaslioglu, M. Murat
Published: (2025)
Rethinking Data Selection at Scale: Random Selection is Almost All You Need
by: Xia, Tingyu, et al.
Published: (2024)
by: Xia, Tingyu, et al.
Published: (2024)
ConfRover: Simultaneous Modeling of Protein Conformation and Dynamics via Autoregression
by: Shen, Yuning, et al.
Published: (2025)
by: Shen, Yuning, et al.
Published: (2025)
What Matters in Transformers? Not All Attention is Needed
by: He, Shwai, et al.
Published: (2024)
by: He, Shwai, et al.
Published: (2024)
RecurFormer: Not All Transformer Heads Need Self-Attention
by: Yan, Ruiqing, et al.
Published: (2024)
by: Yan, Ruiqing, et al.
Published: (2024)
Probing the Lack of Stable Internal Beliefs in LLMs
by: Luo, Yifan, et al.
Published: (2026)
by: Luo, Yifan, et al.
Published: (2026)
SecEncoder: Logs are All You Need in Security
by: Bulut, Muhammed Fatih, et al.
Published: (2024)
by: Bulut, Muhammed Fatih, et al.
Published: (2024)
TransMLA: Multi-Head Latent Attention Is All You Need
by: Meng, Fanxu, et al.
Published: (2025)
by: Meng, Fanxu, et al.
Published: (2025)
Robust Layerwise Scaling Rules by Proper Weight Decay Tuning
by: Fan, Zhiyuan, et al.
Published: (2025)
by: Fan, Zhiyuan, et al.
Published: (2025)
Not All Documents Are What You Need for Extracting Instruction Tuning Data
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
Element-wise Attention Is All You Need
by: Feng, Guoxin
Published: (2025)
by: Feng, Guoxin
Published: (2025)
Evaluation is All You Need: Strategic Overclaiming of LLM Reasoning Capabilities Through Evaluation Design
by: Sun, Lin, et al.
Published: (2025)
by: Sun, Lin, et al.
Published: (2025)
An Extra RMSNorm is All You Need for Fine Tuning to 1.58 Bits
by: Steinmetz, Cody, et al.
Published: (2025)
by: Steinmetz, Cody, et al.
Published: (2025)
Guidance is All You Need: Temperature-Guided Reasoning in Large Language Models
by: Gomaa, Eyad, et al.
Published: (2024)
by: Gomaa, Eyad, et al.
Published: (2024)
English is Not All You Need: Systematically Exploring the Role of Multilinguality in LLM Post-Training
by: Dhaliwal, Mehak, et al.
Published: (2026)
by: Dhaliwal, Mehak, et al.
Published: (2026)
Similar Items
-
On the Design of KL-Regularized Policy Gradient Algorithms for LLM Reasoning
by: Zhang, Yifan, et al.
Published: (2025) -
Group Representational Position Encoding
by: Zhang, Yifan, et al.
Published: (2025) -
Accelerated Preference Optimization for Large Language Model Alignment
by: He, Jiafan, et al.
Published: (2024) -
Higher-order Linear Attention
by: Zhang, Yifan, et al.
Published: (2025) -
Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
by: Chen, Zixiang, et al.
Published: (2024)