Chain and Causal Attention for Efficient Entity Tracking
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fagnou, Erwan, Caillon, Paul, Delattre, Blaise, Allauzen, Alexandre |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Trading Complexity for Expressivity Through Structured Generalized Linear Token Mixing
von: Fagnou, Erwan, et al.
Veröffentlicht: (2026)
von: Fagnou, Erwan, et al.
Veröffentlicht: (2026)
Structured-Sparse Attention for Entity Tracking with Subquadratic Sequence Complexity
von: Zhao, Hangyue, et al.
Veröffentlicht: (2026)
von: Zhao, Hangyue, et al.
Veröffentlicht: (2026)
Accelerated Training through Iterative Gradient Propagation Along the Residual Path
von: Fagnou, Erwan, et al.
Veröffentlicht: (2025)
von: Fagnou, Erwan, et al.
Veröffentlicht: (2025)
QiMeng-Attention: SOTA Attention Operator is generated by SOTA Attention Algorithm
von: Zhou, Qirui, et al.
Veröffentlicht: (2025)
von: Zhou, Qirui, et al.
Veröffentlicht: (2025)
Structure-Guided Entity Resolution: Fine-Tuning LLMs for Robust Name Matching in Complex Linguistic Contexts
von: Chourasia, Shivam, et al.
Veröffentlicht: (2026)
von: Chourasia, Shivam, et al.
Veröffentlicht: (2026)
Representation-Aware Unlearning via Activation Signatures: From Suppression to Entity-Signature Erasure
von: Mahmood, Syed Naveed, et al.
Veröffentlicht: (2026)
von: Mahmood, Syed Naveed, et al.
Veröffentlicht: (2026)
Bridging the Theoretical Gap in Randomized Smoothing
von: Delattre, Blaise, et al.
Veröffentlicht: (2025)
von: Delattre, Blaise, et al.
Veröffentlicht: (2025)
Resonant Context Anchoring: Decoupling Attention Routing and Signal Gain at Inference Time
von: Zhao, Mingkuan, et al.
Veröffentlicht: (2026)
von: Zhao, Mingkuan, et al.
Veröffentlicht: (2026)
Overcoming Long-Context Limitations of State-Space Models via Context-Dependent Sparse Attention
von: Zhan, Zhihao, et al.
Veröffentlicht: (2025)
von: Zhan, Zhihao, et al.
Veröffentlicht: (2025)
Forward Only Learning for Orthogonal Neural Networks of any Depth
von: Caillon, Paul, et al.
Veröffentlicht: (2025)
von: Caillon, Paul, et al.
Veröffentlicht: (2025)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2023)
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2023)
QUAD: Quantization and Parameter-Efficient Tuning of LLM with Activation Decomposition
von: Hu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Hu, Yuxuan, et al.
Veröffentlicht: (2025)
Future Token Prediction -- Causal Language Modelling with Per-Token Semantic State Vector for Multi-Token Prediction
von: Walker, Nicholas
Veröffentlicht: (2024)
von: Walker, Nicholas
Veröffentlicht: (2024)
Self-Pruned Key-Value Attention: Learning When to Write by Predicting Future Utility
von: Szilvasy, Gergely, et al.
Veröffentlicht: (2026)
von: Szilvasy, Gergely, et al.
Veröffentlicht: (2026)
Bayesian Attention Mechanism: A Probabilistic Framework for Positional Encoding and Context Length Extrapolation
von: Bianchessi, Arthur S., et al.
Veröffentlicht: (2025)
von: Bianchessi, Arthur S., et al.
Veröffentlicht: (2025)
Major Entity Identification: A Generalizable Alternative to Coreference Resolution
von: Manikantan, Kawshik, et al.
Veröffentlicht: (2024)
von: Manikantan, Kawshik, et al.
Veröffentlicht: (2024)
Named Entity Recognition and Classification on Historical Documents: A Survey
von: Ehrmann, Maud, et al.
Veröffentlicht: (2021)
von: Ehrmann, Maud, et al.
Veröffentlicht: (2021)
Efficient Solutions For An Intriguing Failure of LLMs: Long Context Window Does Not Mean LLMs Can Analyze Long Sequences Flawlessly
von: Hosseini, Peyman, et al.
Veröffentlicht: (2024)
von: Hosseini, Peyman, et al.
Veröffentlicht: (2024)
Preserving Empirical Probabilities in BERT for Small-sample Clinical Entity Recognition
von: Rehman, Abdul, et al.
Veröffentlicht: (2024)
von: Rehman, Abdul, et al.
Veröffentlicht: (2024)
Disease Entity Recognition and Normalization is Improved with Large Language Model Derived Synthetic Normalized Mentions
von: Sasse, Kuleen, et al.
Veröffentlicht: (2024)
von: Sasse, Kuleen, et al.
Veröffentlicht: (2024)
SPHERICAL KV: Angle-Domain Attention and Rate-Distortion Retention for Efficient Long-Context Inference
von: Chauhan, Anay, et al.
Veröffentlicht: (2026)
von: Chauhan, Anay, et al.
Veröffentlicht: (2026)
Kronecker Embeddings: Byte-Level Structured Token Representations for Parameter-Efficient Language Models
von: Shravan, Rohan
Veröffentlicht: (2026)
von: Shravan, Rohan
Veröffentlicht: (2026)
Engineering A Large Language Model From Scratch
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2024)
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2024)
Large Language Model (LLM) Bias Index -- LLMBI
von: Oketunji, Abiodun Finbarrs, et al.
Veröffentlicht: (2023)
von: Oketunji, Abiodun Finbarrs, et al.
Veröffentlicht: (2023)
RADLADS: Rapid Attention Distillation to Linear Attention Decoders at Scale
von: Goldstein, Daniel, et al.
Veröffentlicht: (2025)
von: Goldstein, Daniel, et al.
Veröffentlicht: (2025)
Forget Attention: Importance-Aware Attention Is All You Need
von: Shin, Soohyeong, et al.
Veröffentlicht: (2026)
von: Shin, Soohyeong, et al.
Veröffentlicht: (2026)
Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms
von: Hanna, Michael, et al.
Veröffentlicht: (2024)
von: Hanna, Michael, et al.
Veröffentlicht: (2024)
Integrating Expert Labels into LLM-based Emission Goal Detection: Example Selection vs Automatic Prompt Design
von: Wrzalik, Marco, et al.
Veröffentlicht: (2024)
von: Wrzalik, Marco, et al.
Veröffentlicht: (2024)
Efficacy of ByT5 in Multilingual Translation of Biblical Texts for Underrepresented Languages
von: Aars, Corinne, et al.
Veröffentlicht: (2024)
von: Aars, Corinne, et al.
Veröffentlicht: (2024)
Human-interpretable clustering of short-text using large language models
von: Miller, Justin K., et al.
Veröffentlicht: (2024)
von: Miller, Justin K., et al.
Veröffentlicht: (2024)
Scaling Laws for Forgetting When Fine-Tuning Large Language Models
von: Kalajdzievski, Damjan
Veröffentlicht: (2024)
von: Kalajdzievski, Damjan
Veröffentlicht: (2024)
A Framework for Fine-Tuning LLMs using Heterogeneous Feedback
von: Aponte, Ryan, et al.
Veröffentlicht: (2024)
von: Aponte, Ryan, et al.
Veröffentlicht: (2024)
The Role of Language Imbalance in Cross-lingual Generalisation: Insights from Cloned Language Experiments
von: Schäfer, Anton, et al.
Veröffentlicht: (2024)
von: Schäfer, Anton, et al.
Veröffentlicht: (2024)
Ensemble Language Models for Multilingual Sentiment Analysis
von: Hasan, Md Arid
Veröffentlicht: (2024)
von: Hasan, Md Arid
Veröffentlicht: (2024)
Linguistically-Informed Multilingual Instruction Tuning: Is There an Optimal Set of Languages to Tune?
von: Soykan, Gürkan, et al.
Veröffentlicht: (2024)
von: Soykan, Gürkan, et al.
Veröffentlicht: (2024)
MedFuzz: Exploring the Robustness of Large Language Models in Medical Question Answering
von: Ness, Robert Osazuwa, et al.
Veröffentlicht: (2024)
von: Ness, Robert Osazuwa, et al.
Veröffentlicht: (2024)
Sparsing Law: Towards Large Language Models with Greater Activation Sparsity
von: Luo, Yuqi, et al.
Veröffentlicht: (2024)
von: Luo, Yuqi, et al.
Veröffentlicht: (2024)
$S^3$ -- Semantic Signal Separation
von: Kardos, Márton, et al.
Veröffentlicht: (2024)
von: Kardos, Márton, et al.
Veröffentlicht: (2024)
MALoRA: Mixture of Asymmetric Low-Rank Adaptation for Enhanced Multi-Task Learning
von: Wang, Xujia, et al.
Veröffentlicht: (2024)
von: Wang, Xujia, et al.
Veröffentlicht: (2024)
On the Effect of (Near) Duplicate Subwords in Language Modelling
von: Schäfer, Anton, et al.
Veröffentlicht: (2024)
von: Schäfer, Anton, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Trading Complexity for Expressivity Through Structured Generalized Linear Token Mixing
von: Fagnou, Erwan, et al.
Veröffentlicht: (2026) -
Structured-Sparse Attention for Entity Tracking with Subquadratic Sequence Complexity
von: Zhao, Hangyue, et al.
Veröffentlicht: (2026) -
Accelerated Training through Iterative Gradient Propagation Along the Residual Path
von: Fagnou, Erwan, et al.
Veröffentlicht: (2025) -
QiMeng-Attention: SOTA Attention Operator is generated by SOTA Attention Algorithm
von: Zhou, Qirui, et al.
Veröffentlicht: (2025) -
Structure-Guided Entity Resolution: Fine-Tuning LLMs for Robust Name Matching in Complex Linguistic Contexts
von: Chourasia, Shivam, et al.
Veröffentlicht: (2026)