Dragutinović, S., Saxe, A. M., & Singh, A. K. (2025). Softmax $\geq$ Linear: Transformers may learn to classify in-context by kernel gradient descent.
Chicago Style (17th ed.) CitationDragutinović, Sara, Andrew M. Saxe, and Aaditya K. Singh. Softmax $\geq$ Linear: Transformers May Learn to Classify In-context by Kernel Gradient Descent. 2025.
MLA (9th ed.) CitationDragutinović, Sara, et al. Softmax $\geq$ Linear: Transformers May Learn to Classify In-context by Kernel Gradient Descent. 2025.
Warning: These citations may not always be 100% accurate.