TaylorShift: Shifting the Complexity of Self-Attention from Squared to Linear (and Back) using Taylor-Softmax
Fuente:
arXiv
Guardado en:
| Autores principales: | Nauen, Tobias Christian, Palacio, Sebastian, Dengel, Andreas |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Which Transformer to Favor: A Comparative Analysis of Efficiency in Vision Transformers
por: Nauen, Tobias Christian, et al.
Publicado: (2023)
por: Nauen, Tobias Christian, et al.
Publicado: (2023)
HuMoCon: Concept Discovery for Human Motion Understanding
por: Fang, Qihang, et al.
Publicado: (2025)
por: Fang, Qihang, et al.
Publicado: (2025)
Anonymization-Enhanced Privacy Protection for Mobile GUI Agents: Available but Invisible
por: Zhao, Lepeng, et al.
Publicado: (2026)
por: Zhao, Lepeng, et al.
Publicado: (2026)
Correspondence of high-dimensional emotion structures elicited by video clips between humans and Multimodal LLMs
por: Asanuma, Haruka, et al.
Publicado: (2025)
por: Asanuma, Haruka, et al.
Publicado: (2025)
Unpacking Hateful Memes: Presupposed Context and False Claims
por: Cai, Weibin, et al.
Publicado: (2025)
por: Cai, Weibin, et al.
Publicado: (2025)
Perception-Consistency Multimodal Large Language Models Reasoning via Caption-Regularized Policy Optimization
por: Tu, Songjun, et al.
Publicado: (2025)
por: Tu, Songjun, et al.
Publicado: (2025)
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
por: Viveiros, André G., et al.
Publicado: (2025)
por: Viveiros, André G., et al.
Publicado: (2025)
Complex-Valued Phase-Coherent Transformer
por: Hioki, Leona
Publicado: (2026)
por: Hioki, Leona
Publicado: (2026)
Predicting When to Trust Vision-Language Models for Spatial Reasoning
por: Imran, Muhammad, et al.
Publicado: (2026)
por: Imran, Muhammad, et al.
Publicado: (2026)
GradES: Significantly Faster Training in Transformers with Gradient-Based Early Stopping
por: Wen, Qifu, et al.
Publicado: (2025)
por: Wen, Qifu, et al.
Publicado: (2025)
Think, Act, Learn: A Framework for Autonomous Robotic Agents using Closed-Loop Large Language Models
por: Menon, Anjali R., et al.
Publicado: (2025)
por: Menon, Anjali R., et al.
Publicado: (2025)
Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests
por: Sáez, Arnau Igualde, et al.
Publicado: (2025)
por: Sáez, Arnau Igualde, et al.
Publicado: (2025)
Inducing Causal World Models in LLMs for Zero-Shot Physical Reasoning
por: Sharma, Aditya, et al.
Publicado: (2025)
por: Sharma, Aditya, et al.
Publicado: (2025)
Optimized Gradient Clipping for Noisy Label Learning
por: Ye, Xichen, et al.
Publicado: (2024)
por: Ye, Xichen, et al.
Publicado: (2024)
Beyond Subtokens: A Rich Character Embedding for Low-resource and Morphologically Complex Languages
por: Schneider, Felix, et al.
Publicado: (2026)
por: Schneider, Felix, et al.
Publicado: (2026)
Diverse capability and scaling of diffusion and auto-regressive models when learning abstract rules
por: Wang, Binxu, et al.
Publicado: (2024)
por: Wang, Binxu, et al.
Publicado: (2024)
TableMoE: Neuro-Symbolic Routing for Structured Expert Reasoning in Multimodal Table Understanding
por: Zhang, Junwen, et al.
Publicado: (2025)
por: Zhang, Junwen, et al.
Publicado: (2025)
When Does Content-Based Routing Work? Representation Requirements for Selective Attention in Hybrid Sequence Models
por: Basu, Abhinaba
Publicado: (2026)
por: Basu, Abhinaba
Publicado: (2026)
Exploring the Capabilities of Large Language Model Encoders for Image-Text Retrieval in Chest X-rays
por: Ko, Hanbin, et al.
Publicado: (2025)
por: Ko, Hanbin, et al.
Publicado: (2025)
TensLoRA: Tensor Alternatives for Low-Rank Adaptation
por: Marmoret, Axel, et al.
Publicado: (2025)
por: Marmoret, Axel, et al.
Publicado: (2025)
Predictive Modeling of Maritime Radar Data Using Transformer Architecture
por: Qesaraku, Bjorna, et al.
Publicado: (2025)
por: Qesaraku, Bjorna, et al.
Publicado: (2025)
Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining
por: Mitchell, Rupert, et al.
Publicado: (2025)
por: Mitchell, Rupert, et al.
Publicado: (2025)
The Origin of Self-Attention: Pairwise Affinity Matrices in Feature Selection and the Emergence of Self-Attention
por: Roffo, Giorgio
Publicado: (2025)
por: Roffo, Giorgio
Publicado: (2025)
PhysNote: Self-Knowledge Notes for Evolvable Physical Reasoning in Vision-Language Model
por: Zhang, Sinin, et al.
Publicado: (2026)
por: Zhang, Sinin, et al.
Publicado: (2026)
The Curious Case of In-Training Compression of State Space Models
por: Chahine, Makram, et al.
Publicado: (2025)
por: Chahine, Makram, et al.
Publicado: (2025)
NeuronSpark: A Spiking Neural Network Language Model with Selective State Space Dynamics
por: Tang, Zhengzheng
Publicado: (2026)
por: Tang, Zhengzheng
Publicado: (2026)
ABot-Claw: A Foundation for Persistent, Cooperative, and Self-Evolving Robotic Agents
por: Huo, Dongjie, et al.
Publicado: (2026)
por: Huo, Dongjie, et al.
Publicado: (2026)
A Survey on Vision-Language-Action Models for Embodied AI
por: Ma, Yueen, et al.
Publicado: (2024)
por: Ma, Yueen, et al.
Publicado: (2024)
K-MetBench: A Multi-Dimensional Benchmark for Fine-Grained Evaluation of Expert Reasoning, Locality, and Multimodality in Meteorology
por: Kim, Soyeon, et al.
Publicado: (2026)
por: Kim, Soyeon, et al.
Publicado: (2026)
Don't Look Back in Anger: MAGIC Net for Streaming Continual Learning with Temporal Dependence
por: Giannini, Federico, et al.
Publicado: (2026)
por: Giannini, Federico, et al.
Publicado: (2026)
Thinking Machines: Mathematical Reasoning in the Age of LLMs
por: Asperti, Andrea, et al.
Publicado: (2025)
por: Asperti, Andrea, et al.
Publicado: (2025)
A Comparative Study of Feature Selection in Tsetlin Machines
por: Halenka, Vojtech, et al.
Publicado: (2025)
por: Halenka, Vojtech, et al.
Publicado: (2025)
Generative AI Models: Opportunities and Risks for Industry and Authorities
por: Alt, Tobias, et al.
Publicado: (2024)
por: Alt, Tobias, et al.
Publicado: (2024)
Rethinking Visual Intelligence: Insights from Video Pretraining
por: Acuaviva, Pablo, et al.
Publicado: (2025)
por: Acuaviva, Pablo, et al.
Publicado: (2025)
MedMemoryBench: Benchmarking Agent Memory in Personalized Healthcare
por: Wang, Yihao, et al.
Publicado: (2026)
por: Wang, Yihao, et al.
Publicado: (2026)
MAcPNN: Mutual Assisted Learning on Data Streams with Temporal Dependence
por: Giannini, Federico, et al.
Publicado: (2026)
por: Giannini, Federico, et al.
Publicado: (2026)
NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution
por: Breneur, Oleksandr Marchenko, et al.
Publicado: (2026)
por: Breneur, Oleksandr Marchenko, et al.
Publicado: (2026)
Rethinking the Multilingual Reasoning Gap with Layer Swap
por: Lasbordes, Maxence, et al.
Publicado: (2026)
por: Lasbordes, Maxence, et al.
Publicado: (2026)
Transparent but Powerful: Explainability, Accuracy, and Generalizability in ADHD Detection from Social Media Data
por: Wiechmann, D., et al.
Publicado: (2024)
por: Wiechmann, D., et al.
Publicado: (2024)
OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping
por: Li, Danyang, et al.
Publicado: (2025)
por: Li, Danyang, et al.
Publicado: (2025)
Ejemplares similares
-
Which Transformer to Favor: A Comparative Analysis of Efficiency in Vision Transformers
por: Nauen, Tobias Christian, et al.
Publicado: (2023) -
HuMoCon: Concept Discovery for Human Motion Understanding
por: Fang, Qihang, et al.
Publicado: (2025) -
Anonymization-Enhanced Privacy Protection for Mobile GUI Agents: Available but Invisible
por: Zhao, Lepeng, et al.
Publicado: (2026) -
Correspondence of high-dimensional emotion structures elicited by video clips between humans and Multimodal LLMs
por: Asanuma, Haruka, et al.
Publicado: (2025) -
Unpacking Hateful Memes: Presupposed Context and False Claims
por: Cai, Weibin, et al.
Publicado: (2025)