Hierarchical Shift Mixing -- Beyond Dense Attention in Transformers
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Forchheimer, Robert |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TensorLens: End-to-End Transformer Analysis via High-Order Attention Tensors
von: Atad, Ido Andrew, et al.
Veröffentlicht: (2026)
von: Atad, Ido Andrew, et al.
Veröffentlicht: (2026)
The Hidden Attention of Mamba Models
von: Ali, Ameen, et al.
Veröffentlicht: (2024)
von: Ali, Ameen, et al.
Veröffentlicht: (2024)
Merge-Bench: Resolve Merge Conflicts with Large Language Models
von: Schesch, Benedikt, et al.
Veröffentlicht: (2026)
von: Schesch, Benedikt, et al.
Veröffentlicht: (2026)
RAPTOR-AI for Disaster OODA Loop: Hierarchical Multimodal RAG with Experience-Driven Agentic Decision-Making
von: Yasuno, Takato
Veröffentlicht: (2026)
von: Yasuno, Takato
Veröffentlicht: (2026)
Recurrent Memory-Augmented Transformers with Chunked Attention for Long-Context Language Modeling
von: Kashyap, Ankit
Veröffentlicht: (2025)
von: Kashyap, Ankit
Veröffentlicht: (2025)
Revisiting LRP: Positional Attribution as the Missing Ingredient for Transformer Explainability
von: Bakish, Yarden, et al.
Veröffentlicht: (2025)
von: Bakish, Yarden, et al.
Veröffentlicht: (2025)
Variance Is Not Importance: Structural Analysis of Transformer Compressibility Across Model Scales
von: Salfati, Samuel
Veröffentlicht: (2026)
von: Salfati, Samuel
Veröffentlicht: (2026)
Transformer Scalability Crisis: The First Comprehensive Empirical Analysis of Performance Walls in Modern Language Models
von: Moghadasi, Mahdi Naser, et al.
Veröffentlicht: (2026)
von: Moghadasi, Mahdi Naser, et al.
Veröffentlicht: (2026)
QiMeng-Attention: SOTA Attention Operator is generated by SOTA Attention Algorithm
von: Zhou, Qirui, et al.
Veröffentlicht: (2025)
von: Zhou, Qirui, et al.
Veröffentlicht: (2025)
Relating Misfit to Gain in Weak-to-Strong Generalization Beyond the Squared Loss
von: Mulgund, Abhijeet, et al.
Veröffentlicht: (2025)
von: Mulgund, Abhijeet, et al.
Veröffentlicht: (2025)
OCRR: A Benchmark for Online Correction Recovery under Distribution Shift
von: Grassi, Adrian
Veröffentlicht: (2026)
von: Grassi, Adrian
Veröffentlicht: (2026)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2023)
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2023)
TaylorShift: Shifting the Complexity of Self-Attention from Squared to Linear (and Back) using Taylor-Softmax
von: Nauen, Tobias Christian, et al.
Veröffentlicht: (2024)
von: Nauen, Tobias Christian, et al.
Veröffentlicht: (2024)
Listwise Direct Preference Optimization with Multi-Dimensional Preference Mixing
von: Sun, Yuhui, et al.
Veröffentlicht: (2025)
von: Sun, Yuhui, et al.
Veröffentlicht: (2025)
Combining Language and Topic Models for Hierarchical Text Classification
von: Toit, Jaco du, et al.
Veröffentlicht: (2025)
von: Toit, Jaco du, et al.
Veröffentlicht: (2025)
Chain and Causal Attention for Efficient Entity Tracking
von: Fagnou, Erwan, et al.
Veröffentlicht: (2024)
von: Fagnou, Erwan, et al.
Veröffentlicht: (2024)
Introducing Three New Benchmark Datasets for Hierarchical Text Classification
von: Toit, Jaco du, et al.
Veröffentlicht: (2024)
von: Toit, Jaco du, et al.
Veröffentlicht: (2024)
Random Heterogeneous Neurochaos Learning Architecture for Data Classification
von: S, Remya Ajai A, et al.
Veröffentlicht: (2024)
von: S, Remya Ajai A, et al.
Veröffentlicht: (2024)
Trading Complexity for Expressivity Through Structured Generalized Linear Token Mixing
von: Fagnou, Erwan, et al.
Veröffentlicht: (2026)
von: Fagnou, Erwan, et al.
Veröffentlicht: (2026)
Structured-Sparse Attention for Entity Tracking with Subquadratic Sequence Complexity
von: Zhao, Hangyue, et al.
Veröffentlicht: (2026)
von: Zhao, Hangyue, et al.
Veröffentlicht: (2026)
Attention Drift: What Autoregressive Speculative Decoding Models Learn
von: Eldenk, Doğaç, et al.
Veröffentlicht: (2026)
von: Eldenk, Doğaç, et al.
Veröffentlicht: (2026)
Solving the Granularity Mismatch: Hierarchical Preference Learning for Long-Horizon LLM Agents
von: Gao, Heyang, et al.
Veröffentlicht: (2025)
von: Gao, Heyang, et al.
Veröffentlicht: (2025)
Explaining Modern Gated-Linear RNNs via a Unified Implicit Attention Formulation
von: Zimerman, Itamar, et al.
Veröffentlicht: (2024)
von: Zimerman, Itamar, et al.
Veröffentlicht: (2024)
On Robustness and Chain-of-Thought Consistency of RL-Finetuned VLMs
von: Zhao, Rosie, et al.
Veröffentlicht: (2026)
von: Zhao, Rosie, et al.
Veröffentlicht: (2026)
LLM Unlearning on Noisy Forget Sets: A Study of Incomplete, Rewritten, and Watermarked Data
von: Wang, Changsheng, et al.
Veröffentlicht: (2025)
von: Wang, Changsheng, et al.
Veröffentlicht: (2025)
Procedural Environment Generation for Tool-Use Agents
von: Sullivan, Michael, et al.
Veröffentlicht: (2025)
von: Sullivan, Michael, et al.
Veröffentlicht: (2025)
EvilGenie: A Reward Hacking Benchmark
von: Gabor, Jonathan, et al.
Veröffentlicht: (2025)
von: Gabor, Jonathan, et al.
Veröffentlicht: (2025)
On the Challenges of Creating Datasets for Analyzing Commercial Sex Advertisements to Assess Human Trafficking Risk and Organized Activity
von: Rivas, Pablo, et al.
Veröffentlicht: (2024)
von: Rivas, Pablo, et al.
Veröffentlicht: (2024)
Unsupervised Evaluation of Multi-Turn Objective-Driven Interactions
von: Soroka, Emi, et al.
Veröffentlicht: (2025)
von: Soroka, Emi, et al.
Veröffentlicht: (2025)
Random Scaling of Emergent Capabilities
von: Zhao, Rosie, et al.
Veröffentlicht: (2025)
von: Zhao, Rosie, et al.
Veröffentlicht: (2025)
Spectral Clustering in Convex and Constrained Settings
von: Behera, Swarup Ranjan, et al.
Veröffentlicht: (2024)
von: Behera, Swarup Ranjan, et al.
Veröffentlicht: (2024)
Mr. Snuffleupagus at SemEval-2025 Task 4: Unlearning Factual Knowledge from LLMs Using Adaptive RMU
von: Dosajh, Arjun, et al.
Veröffentlicht: (2025)
von: Dosajh, Arjun, et al.
Veröffentlicht: (2025)
Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
von: Ahmadian, Arash, et al.
Veröffentlicht: (2024)
von: Ahmadian, Arash, et al.
Veröffentlicht: (2024)
Guided Speculative Inference for Efficient Test-Time Alignment of LLMs
von: Geuter, Jonathan, et al.
Veröffentlicht: (2025)
von: Geuter, Jonathan, et al.
Veröffentlicht: (2025)
Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
von: Zhao, Rosie, et al.
Veröffentlicht: (2025)
von: Zhao, Rosie, et al.
Veröffentlicht: (2025)
Interpreto: An Explainability Library for Transformers
von: Poché, Antonin, et al.
Veröffentlicht: (2025)
von: Poché, Antonin, et al.
Veröffentlicht: (2025)
Resonant Context Anchoring: Decoupling Attention Routing and Signal Gain at Inference Time
von: Zhao, Mingkuan, et al.
Veröffentlicht: (2026)
von: Zhao, Mingkuan, et al.
Veröffentlicht: (2026)
Beyond Memorization: Violating Privacy Via Inference with Large Language Models
von: Staab, Robin, et al.
Veröffentlicht: (2023)
von: Staab, Robin, et al.
Veröffentlicht: (2023)
Beyond the Black Box: A Statistical Model for LLM Reasoning and Inference
von: Dalal, Siddhartha, et al.
Veröffentlicht: (2024)
von: Dalal, Siddhartha, et al.
Veröffentlicht: (2024)
Self-Pruned Key-Value Attention: Learning When to Write by Predicting Future Utility
von: Szilvasy, Gergely, et al.
Veröffentlicht: (2026)
von: Szilvasy, Gergely, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
TensorLens: End-to-End Transformer Analysis via High-Order Attention Tensors
von: Atad, Ido Andrew, et al.
Veröffentlicht: (2026) -
The Hidden Attention of Mamba Models
von: Ali, Ameen, et al.
Veröffentlicht: (2024) -
Merge-Bench: Resolve Merge Conflicts with Large Language Models
von: Schesch, Benedikt, et al.
Veröffentlicht: (2026) -
RAPTOR-AI for Disaster OODA Loop: Hierarchical Multimodal RAG with Experience-Driven Agentic Decision-Making
von: Yasuno, Takato
Veröffentlicht: (2026) -
Recurrent Memory-Augmented Transformers with Chunked Attention for Long-Context Language Modeling
von: Kashyap, Ankit
Veröffentlicht: (2025)