Analyzing Feed-Forward Blocks in Transformers through the Lens of Attention Maps
Fuente:
arXiv
Guardado en:
| Autores principales: | Kobayashi, Goro, Kuribayashi, Tatsuki, Yokoi, Sho, Inui, Kentaro |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Can Input Attributions Explain Inductive Reasoning in In-Context Learning?
por: Ye, Mengyu, et al.
Publicado: (2024)
por: Ye, Mengyu, et al.
Publicado: (2024)
Syntactic Learnability of Echo State Neural Language Models at Scale
por: Ueda, Ryo, et al.
Publicado: (2025)
por: Ueda, Ryo, et al.
Publicado: (2025)
To Drop or Not to Drop? Predicting Argument Ellipsis Judgments: A Case Study in Japanese
por: Ishizuki, Yukiko, et al.
Publicado: (2024)
por: Ishizuki, Yukiko, et al.
Publicado: (2024)
Large Language Models Are Human-Like Internally
por: Kuribayashi, Tatsuki, et al.
Publicado: (2025)
por: Kuribayashi, Tatsuki, et al.
Publicado: (2025)
Linear Representations of Hierarchical Concepts in Language Models
por: Sakata, Masaki, et al.
Publicado: (2026)
por: Sakata, Masaki, et al.
Publicado: (2026)
Why Mean Pooling Works: Quantifying Second-Order Collapse in Text Embeddings
por: Hara, Tomomasa, et al.
Publicado: (2026)
por: Hara, Tomomasa, et al.
Publicado: (2026)
On Entity Identification in Language Models
por: Sakata, Masaki, et al.
Publicado: (2025)
por: Sakata, Masaki, et al.
Publicado: (2025)
First Heuristic Then Rational: Dynamic Use of Heuristics in Language Model Reasoning
por: Aoki, Yoichi, et al.
Publicado: (2024)
por: Aoki, Yoichi, et al.
Publicado: (2024)
Does Vision Accelerate Hierarchical Generalization in Neural Language Learners?
por: Kuribayashi, Tatsuki, et al.
Publicado: (2023)
por: Kuribayashi, Tatsuki, et al.
Publicado: (2023)
LLMs Faithfully and Iteratively Compute Answers During CoT: A Systematic Analysis With Multi-step Arithmetics
por: Kudo, Keito, et al.
Publicado: (2024)
por: Kudo, Keito, et al.
Publicado: (2024)
FinchGPT: a Transformer based language model for birdsong analysis
por: Kobayashi, Kosei, et al.
Publicado: (2025)
por: Kobayashi, Kosei, et al.
Publicado: (2025)
On Representational Dissociation of Language and Arithmetic in Large Language Models
por: Kisako, Riku, et al.
Publicado: (2025)
por: Kisako, Riku, et al.
Publicado: (2025)
Psychometric Predictive Power of Large Language Models
por: Kuribayashi, Tatsuki, et al.
Publicado: (2023)
por: Kuribayashi, Tatsuki, et al.
Publicado: (2023)
Which Word Orders Facilitate Length Generalization in LMs? An Investigation with GCG-Based Artificial Languages
por: El-Naggar, Nadine, et al.
Publicado: (2025)
por: El-Naggar, Nadine, et al.
Publicado: (2025)
What Kind of Language is Easy to Language-Model Under Curriculum Learning?
por: El-Naggar, Nadine, et al.
Publicado: (2026)
por: El-Naggar, Nadine, et al.
Publicado: (2026)
Spelling-out is not Straightforward: LLMs' Capability of Tokenization from Token to Characters
por: Hiraoka, Tatsuya, et al.
Publicado: (2025)
por: Hiraoka, Tatsuya, et al.
Publicado: (2025)
Repetition Neurons: How Do Language Models Produce Repetitions?
por: Hiraoka, Tatsuya, et al.
Publicado: (2024)
por: Hiraoka, Tatsuya, et al.
Publicado: (2024)
Monotonic Representation of Numeric Properties in Language Models
por: Heinzerling, Benjamin, et al.
Publicado: (2024)
por: Heinzerling, Benjamin, et al.
Publicado: (2024)
Rethinking Attention: Exploring Shallow Feed-Forward Neural Networks as an Alternative to Attention Layers in Transformers
por: Bozic, Vukasin, et al.
Publicado: (2023)
por: Bozic, Vukasin, et al.
Publicado: (2023)
LLMs Can Compensate for Deficiencies in Visual Representations
por: Takishita, Sho, et al.
Publicado: (2025)
por: Takishita, Sho, et al.
Publicado: (2025)
J-UniMorph: Japanese Morphological Annotation through the Universal Feature Schema
por: Matsuzaki, Kosuke, et al.
Publicado: (2024)
por: Matsuzaki, Kosuke, et al.
Publicado: (2024)
Dual Alignment Between Language Model Layers and Human Sentence Processing
por: Kuribayashi, Tatsuki, et al.
Publicado: (2026)
por: Kuribayashi, Tatsuki, et al.
Publicado: (2026)
Merging Feed-Forward Sublayers for Compressed Transformers
por: Verma, Neha, et al.
Publicado: (2025)
por: Verma, Neha, et al.
Publicado: (2025)
Rectifying Belief Space via Unlearning to Harness LLMs' Reasoning
por: Niwa, Ayana, et al.
Publicado: (2025)
por: Niwa, Ayana, et al.
Publicado: (2025)
Cell-Based Representation of Relational Binding in Language Models
por: Dai, Qin, et al.
Publicado: (2026)
por: Dai, Qin, et al.
Publicado: (2026)
Representational Analysis of Binding in Language Models
por: Dai, Qin, et al.
Publicado: (2024)
por: Dai, Qin, et al.
Publicado: (2024)
An Existence Proof for Neural Language Models That Can Explain Garden-Path Effects via Surprisal
por: Yoshida, Ryo, et al.
Publicado: (2026)
por: Yoshida, Ryo, et al.
Publicado: (2026)
Sycophancy Hides Linearly in the Attention Heads
por: Genadi, Rifo, et al.
Publicado: (2026)
por: Genadi, Rifo, et al.
Publicado: (2026)
Understanding and Controlling Repetition Neurons and Induction Heads in In-Context Learning
por: Doan, Nhi Hoai, et al.
Publicado: (2025)
por: Doan, Nhi Hoai, et al.
Publicado: (2025)
Analyzing Wrap-Up Effects through an Information-Theoretic Lens
por: Meister, Clara, et al.
Publicado: (2022)
por: Meister, Clara, et al.
Publicado: (2022)
Can Language Models Learn Typologically Implausible Languages?
por: Xu, Tianyang, et al.
Publicado: (2025)
por: Xu, Tianyang, et al.
Publicado: (2025)
Layerwise Importance Analysis of Feed-Forward Networks in Transformer-based Language Models
por: Ikeda, Wataru, et al.
Publicado: (2025)
por: Ikeda, Wataru, et al.
Publicado: (2025)
Emergent Word Order Universals from Cognitively-Motivated Language Models
por: Kuribayashi, Tatsuki, et al.
Publicado: (2024)
por: Kuribayashi, Tatsuki, et al.
Publicado: (2024)
TopK Language Models
por: Takahashi, Ryosuke, et al.
Publicado: (2025)
por: Takahashi, Ryosuke, et al.
Publicado: (2025)
A Large Collection of Model-generated Contradictory Responses for Consistency-aware Dialogue Systems
por: Sato, Shiki, et al.
Publicado: (2024)
por: Sato, Shiki, et al.
Publicado: (2024)
Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge
por: Zhang, Ying, et al.
Publicado: (2025)
por: Zhang, Ying, et al.
Publicado: (2025)
Zipfian Whitening
por: Yokoi, Sho, et al.
Publicado: (2024)
por: Yokoi, Sho, et al.
Publicado: (2024)
Subspace Representations for Soft Set Operations and Sentence Similarities
por: Ishibashi, Yoichi, et al.
Publicado: (2022)
por: Ishibashi, Yoichi, et al.
Publicado: (2022)
Modality Bias in LVLMs: Analyzing and Mitigating Object Hallucination via Attention Lens
por: Zheng, Haohan, et al.
Publicado: (2025)
por: Zheng, Haohan, et al.
Publicado: (2025)
Quantifying Lexical Semantic Shift via Unbalanced Optimal Transport
por: Kishino, Ryo, et al.
Publicado: (2024)
por: Kishino, Ryo, et al.
Publicado: (2024)
Ejemplares similares
-
Can Input Attributions Explain Inductive Reasoning in In-Context Learning?
por: Ye, Mengyu, et al.
Publicado: (2024) -
Syntactic Learnability of Echo State Neural Language Models at Scale
por: Ueda, Ryo, et al.
Publicado: (2025) -
To Drop or Not to Drop? Predicting Argument Ellipsis Judgments: A Case Study in Japanese
por: Ishizuki, Yukiko, et al.
Publicado: (2024) -
Large Language Models Are Human-Like Internally
por: Kuribayashi, Tatsuki, et al.
Publicado: (2025) -
Linear Representations of Hierarchical Concepts in Language Models
por: Sakata, Masaki, et al.
Publicado: (2026)