Emergence of Primacy and Recency Effect in Mamba: A Mechanistic Point of View
Fuente:
arXiv
Saved in:
| Main Authors: | Airlangga, Muhammad Cendekia, AlQuabeh, Hilal, Nwadike, Munachiso S, Inui, Kentaro |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mechanistic Insights into Grokking from the Embedding Layer
by: AlquBoj, H. V., et al.
Published: (2025)
by: AlquBoj, H. V., et al.
Published: (2025)
Number Representations in LLMs: A Computational Parallel to Human Perception
by: AlquBoj, H. V., et al.
Published: (2025)
by: AlquBoj, H. V., et al.
Published: (2025)
The Geometry of Numerical Reasoning: Language Models Compare Numeric Properties in Linear Subspaces
by: El-Shangiti, Ahmed Oumar, et al.
Published: (2024)
by: El-Shangiti, Ahmed Oumar, et al.
Published: (2024)
Sycophancy Hides Linearly in the Attention Heads
by: Genadi, Rifo, et al.
Published: (2026)
by: Genadi, Rifo, et al.
Published: (2026)
Measuring AI Reasoning: A Guide for Researchers
by: Nwadike, Munachiso Samuel, et al.
Published: (2026)
by: Nwadike, Munachiso Samuel, et al.
Published: (2026)
Uncovering the Spectral Bias in Diagonal State Space Models
by: Solozabal, Ruben, et al.
Published: (2025)
by: Solozabal, Ruben, et al.
Published: (2025)
RECALL: Library-Like Behavior In Language Models is Enhanced by Self-Referencing Causal Cycles
by: Nwadike, Munachiso, et al.
Published: (2025)
by: Nwadike, Munachiso, et al.
Published: (2025)
ASR Under Noise: Exploring Robustness for Sundanese and Javanese
by: Pranida, Salsabila Zahirah, et al.
Published: (2025)
by: Pranida, Salsabila Zahirah, et al.
Published: (2025)
Limited Memory Online Gradient Descent for Kernelized Pairwise Learning with Dynamic Averaging
by: AlQuabeh, Hilal, et al.
Published: (2024)
by: AlQuabeh, Hilal, et al.
Published: (2024)
The AI Data Scientist
by: Akimov, Farkhad, et al.
Published: (2025)
by: Akimov, Farkhad, et al.
Published: (2025)
Spelling-out is not Straightforward: LLMs' Capability of Tokenization from Token to Characters
by: Hiraoka, Tatsuya, et al.
Published: (2025)
by: Hiraoka, Tatsuya, et al.
Published: (2025)
Repetition Neurons: How Do Language Models Produce Repetitions?
by: Hiraoka, Tatsuya, et al.
Published: (2024)
by: Hiraoka, Tatsuya, et al.
Published: (2024)
Monotonic Representation of Numeric Properties in Language Models
by: Heinzerling, Benjamin, et al.
Published: (2024)
by: Heinzerling, Benjamin, et al.
Published: (2024)
Primacy Effect of ChatGPT
by: Wang, Yiwei, et al.
Published: (2023)
by: Wang, Yiwei, et al.
Published: (2023)
Improving Personalisation in Valence and Arousal Prediction using Data Augmentation
by: Nwadike, Munachiso, et al.
Published: (2024)
by: Nwadike, Munachiso, et al.
Published: (2024)
Rectifying Belief Space via Unlearning to Harness LLMs' Reasoning
by: Niwa, Ayana, et al.
Published: (2025)
by: Niwa, Ayana, et al.
Published: (2025)
Cell-Based Representation of Relational Binding in Language Models
by: Dai, Qin, et al.
Published: (2026)
by: Dai, Qin, et al.
Published: (2026)
Representational Analysis of Binding in Language Models
by: Dai, Qin, et al.
Published: (2024)
by: Dai, Qin, et al.
Published: (2024)
Constrained Adversarial Perturbation
by: Nishad, Virendra, et al.
Published: (2025)
by: Nishad, Virendra, et al.
Published: (2025)
Exploiting Primacy Effect To Improve Large Language Models
by: Raimondi, Bianca, et al.
Published: (2025)
by: Raimondi, Bianca, et al.
Published: (2025)
Understanding and Controlling Repetition Neurons and Induction Heads in In-Context Learning
by: Doan, Nhi Hoai, et al.
Published: (2025)
by: Doan, Nhi Hoai, et al.
Published: (2025)
A Large Collection of Model-generated Contradictory Responses for Consistency-aware Dialogue Systems
by: Sato, Shiki, et al.
Published: (2024)
by: Sato, Shiki, et al.
Published: (2024)
The Curse of Popularity: Popular Entities have Catastrophic Side Effects when Deleting Knowledge from Language Models
by: Takahashi, Ryosuke, et al.
Published: (2024)
by: Takahashi, Ryosuke, et al.
Published: (2024)
Syntactic Learnability of Echo State Neural Language Models at Scale
by: Ueda, Ryo, et al.
Published: (2025)
by: Ueda, Ryo, et al.
Published: (2025)
TopK Language Models
by: Takahashi, Ryosuke, et al.
Published: (2025)
by: Takahashi, Ryosuke, et al.
Published: (2025)
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge
by: Zhang, Ying, et al.
Published: (2025)
by: Zhang, Ying, et al.
Published: (2025)
J-UniMorph: Japanese Morphological Annotation through the Universal Feature Schema
by: Matsuzaki, Kosuke, et al.
Published: (2024)
by: Matsuzaki, Kosuke, et al.
Published: (2024)
Analyzing Feed-Forward Blocks in Transformers through the Lens of Attention Maps
by: Kobayashi, Goro, et al.
Published: (2023)
by: Kobayashi, Goro, et al.
Published: (2023)
How Individual Traits and Language Styles Shape Preferences In Open-ended User-LLM Interaction: A Preliminary Study
by: Chevi, Rendi, et al.
Published: (2025)
by: Chevi, Rendi, et al.
Published: (2025)
FinchGPT: a Transformer based language model for birdsong analysis
by: Kobayashi, Kosei, et al.
Published: (2025)
by: Kobayashi, Kosei, et al.
Published: (2025)
On Psychology of AI -- Does Primacy Effect Affect ChatGPT and Other LLMs?
by: Hämäläinen, Mika
Published: (2025)
by: Hämäläinen, Mika
Published: (2025)
To Drop or Not to Drop? Predicting Argument Ellipsis Judgments: A Case Study in Japanese
by: Ishizuki, Yukiko, et al.
Published: (2024)
by: Ishizuki, Yukiko, et al.
Published: (2024)
LayerNorm Induces Recency Bias in Transformer Decoders
by: Kim, Junu, et al.
Published: (2025)
by: Kim, Junu, et al.
Published: (2025)
On Entity Identification in Language Models
by: Sakata, Masaki, et al.
Published: (2025)
by: Sakata, Masaki, et al.
Published: (2025)
Linear Representations of Hierarchical Concepts in Language Models
by: Sakata, Masaki, et al.
Published: (2026)
by: Sakata, Masaki, et al.
Published: (2026)
Why Mean Pooling Works: Quantifying Second-Order Collapse in Text Embeddings
by: Hara, Tomomasa, et al.
Published: (2026)
by: Hara, Tomomasa, et al.
Published: (2026)
Reducing the Cost: Cross-Prompt Pre-Finetuning for Short Answer Scoring
by: Funayama, Hiroaki, et al.
Published: (2024)
by: Funayama, Hiroaki, et al.
Published: (2024)
MQM-Chat: Multidimensional Quality Metrics for Chat Translation
by: Li, Yunmeng, et al.
Published: (2024)
by: Li, Yunmeng, et al.
Published: (2024)
ACORN: Aspect-wise Commonsense Reasoning Explanation Evaluation
by: Brassard, Ana, et al.
Published: (2024)
by: Brassard, Ana, et al.
Published: (2024)
How often do Answers Change? Estimating Recency Requirements in Question Answering
by: Piryani, Bhawna, et al.
Published: (2026)
by: Piryani, Bhawna, et al.
Published: (2026)
Similar Items
-
Mechanistic Insights into Grokking from the Embedding Layer
by: AlquBoj, H. V., et al.
Published: (2025) -
Number Representations in LLMs: A Computational Parallel to Human Perception
by: AlquBoj, H. V., et al.
Published: (2025) -
The Geometry of Numerical Reasoning: Language Models Compare Numeric Properties in Linear Subspaces
by: El-Shangiti, Ahmed Oumar, et al.
Published: (2024) -
Sycophancy Hides Linearly in the Attention Heads
by: Genadi, Rifo, et al.
Published: (2026) -
Measuring AI Reasoning: A Guide for Researchers
by: Nwadike, Munachiso Samuel, et al.
Published: (2026)