Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions
Fuente:
arXiv
Salvato in:
| Autori principali: | Neo, Clement, Cohen, Shay B., Barez, Fazl |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Understanding Addition and Subtraction in Transformers
di: Quirke, Philip, et al.
Pubblicazione: (2024)
di: Quirke, Philip, et al.
Pubblicazione: (2024)
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
di: Lan, Michael, et al.
Pubblicazione: (2023)
di: Lan, Michael, et al.
Pubblicazione: (2023)
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
di: Chaudhary, Maheep, et al.
Pubblicazione: (2025)
di: Chaudhary, Maheep, et al.
Pubblicazione: (2025)
Same Question, Different Words: A Latent Adversarial Framework for Prompt Robustness
di: Fu, Tingchen, et al.
Pubblicazione: (2025)
di: Fu, Tingchen, et al.
Pubblicazione: (2025)
Understanding Addition in Transformers
di: Quirke, Philip, et al.
Pubblicazione: (2023)
di: Quirke, Philip, et al.
Pubblicazione: (2023)
Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
di: Oozeer, Narmeen, et al.
Pubblicazione: (2025)
di: Oozeer, Narmeen, et al.
Pubblicazione: (2025)
Scaling sparse feature circuit finding for in-context learning
di: Kharlapenko, Dmitrii, et al.
Pubblicazione: (2025)
di: Kharlapenko, Dmitrii, et al.
Pubblicazione: (2025)
Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders
di: Lan, Michael, et al.
Pubblicazione: (2024)
di: Lan, Michael, et al.
Pubblicazione: (2024)
Old Habits Die Hard: How Conversational History Geometrically Traps LLMs
di: Simhi, Adi, et al.
Pubblicazione: (2026)
di: Simhi, Adi, et al.
Pubblicazione: (2026)
JoMA: Demystifying Multilayer Transformers via JOint Dynamics of MLP and Attention
di: Tian, Yuandong, et al.
Pubblicazione: (2023)
di: Tian, Yuandong, et al.
Pubblicazione: (2023)
PoisonBench: Assessing Large Language Model Vulnerability to Data Poisoning
di: Fu, Tingchen, et al.
Pubblicazione: (2024)
di: Fu, Tingchen, et al.
Pubblicazione: (2024)
Large Language Models Relearn Removed Concepts
di: Lo, Michelle, et al.
Pubblicazione: (2024)
di: Lo, Michelle, et al.
Pubblicazione: (2024)
Best-of-N Jailbreaking
di: Hughes, John, et al.
Pubblicazione: (2024)
di: Hughes, John, et al.
Pubblicazione: (2024)
Towards Interpreting Visual Information Processing in Vision-Language Models
di: Neo, Clement, et al.
Pubblicazione: (2024)
di: Neo, Clement, et al.
Pubblicazione: (2024)
VAL-Bench: Belief Consistency as a measure for Value Alignment in Language Models
di: Gupta, Aman, et al.
Pubblicazione: (2025)
di: Gupta, Aman, et al.
Pubblicazione: (2025)
What can Large Language Models Capture about Code Functional Equivalence?
di: Maveli, Nickil, et al.
Pubblicazione: (2024)
di: Maveli, Nickil, et al.
Pubblicazione: (2024)
In-Context Learning with Transformers: Softmax Attention Adapts to Function Lipschitzness
di: Collins, Liam, et al.
Pubblicazione: (2024)
di: Collins, Liam, et al.
Pubblicazione: (2024)
MoBA: Mixture of Block Attention for Long-Context LLMs
di: Lu, Enzhe, et al.
Pubblicazione: (2025)
di: Lu, Enzhe, et al.
Pubblicazione: (2025)
HyperMLP: An Integrated Perspective for Sequence Modeling
di: Lu, Jiecheng, et al.
Pubblicazione: (2026)
di: Lu, Jiecheng, et al.
Pubblicazione: (2026)
CoMAT: Chain of Mathematically Annotated Thought Improves Mathematical Reasoning
di: Leang, Joshua Ong Jun, et al.
Pubblicazione: (2024)
di: Leang, Joshua Ong Jun, et al.
Pubblicazione: (2024)
Benchmark Inflation: Revealing LLM Performance Gaps Using Retro-Holdouts
di: Haimes, Jacob, et al.
Pubblicazione: (2024)
di: Haimes, Jacob, et al.
Pubblicazione: (2024)
PersonaLens: A Benchmark for Personalization Evaluation in Conversational AI Assistants
di: Zhao, Zheng, et al.
Pubblicazione: (2025)
di: Zhao, Zheng, et al.
Pubblicazione: (2025)
Jailbreak Defense in a Narrow Domain: Limitations of Existing Methods and a New Transcript-Classifier Approach
di: Wang, Tony T., et al.
Pubblicazione: (2024)
di: Wang, Tony T., et al.
Pubblicazione: (2024)
Visualizing Neural Network Imagination
di: Wichers, Nevan, et al.
Pubblicazione: (2024)
di: Wichers, Nevan, et al.
Pubblicazione: (2024)
Towards Understanding Subliminal Learning: When and How Hidden Biases Transfer
di: Schrodi, Simon, et al.
Pubblicazione: (2025)
di: Schrodi, Simon, et al.
Pubblicazione: (2025)
Spectral Editing of Activations for Large Language Model Alignment
di: Qiu, Yifu, et al.
Pubblicazione: (2024)
di: Qiu, Yifu, et al.
Pubblicazione: (2024)
CoCA: Fusing Position Embedding with Collinear Constrained Attention in Transformers for Long Context Window Extending
di: Zhu, Shiyi, et al.
Pubblicazione: (2023)
di: Zhu, Shiyi, et al.
Pubblicazione: (2023)
Incorporating Exponential Smoothing into MLP: A Simple but Effective Sequence Model
di: Chu, Jiqun, et al.
Pubblicazione: (2024)
di: Chu, Jiqun, et al.
Pubblicazione: (2024)
Can LLMs Compress (and Decompress)? Evaluating Code Understanding and Execution via Invertibility
di: Maveli, Nickil, et al.
Pubblicazione: (2026)
di: Maveli, Nickil, et al.
Pubblicazione: (2026)
Toeplitz MLP Mixers are Low Complexity, Information-Rich Sequence Models
di: Badger, Benjamin L., et al.
Pubblicazione: (2026)
di: Badger, Benjamin L., et al.
Pubblicazione: (2026)
SDMPrune: Self-Distillation MLP Pruning for Efficient Large Language Models
di: Zhu, Hourun, et al.
Pubblicazione: (2025)
di: Zhu, Hourun, et al.
Pubblicazione: (2025)
Self-Improving World Modelling with Latent Actions
di: Qiu, Yifu, et al.
Pubblicazione: (2026)
di: Qiu, Yifu, et al.
Pubblicazione: (2026)
ONNX-Net: Towards Universal Representations and Instant Performance Prediction for Neural Architectures
di: Qin, Shiwen, et al.
Pubblicazione: (2025)
di: Qin, Shiwen, et al.
Pubblicazione: (2025)
Used Car Salesbots? Honesty and Credulity of LLMs as Bargaining Agents under Partial Information
di: Miceli-Barone, Antonio Valerio, et al.
Pubblicazione: (2026)
di: Miceli-Barone, Antonio Valerio, et al.
Pubblicazione: (2026)
Overcoming Sparsity Artifacts in Crosscoders to Interpret Chat-Tuning
di: Minder, Julian, et al.
Pubblicazione: (2025)
di: Minder, Julian, et al.
Pubblicazione: (2025)
Decomposing Attention To Find Context-Sensitive Neurons
di: Gibson, Alex
Pubblicazione: (2025)
di: Gibson, Alex
Pubblicazione: (2025)
Which Attention Heads Matter for In-Context Learning?
di: Yin, Kayo, et al.
Pubblicazione: (2025)
di: Yin, Kayo, et al.
Pubblicazione: (2025)
Selective Attention Improves Transformer
di: Leviathan, Yaniv, et al.
Pubblicazione: (2024)
di: Leviathan, Yaniv, et al.
Pubblicazione: (2024)
Does Transformer Interpretability Transfer to RNNs?
di: Paulo, Gonçalo, et al.
Pubblicazione: (2024)
di: Paulo, Gonçalo, et al.
Pubblicazione: (2024)
SEAL: Scaling to Emphasize Attention for Long-Context Retrieval
di: Lee, Changhun, et al.
Pubblicazione: (2025)
di: Lee, Changhun, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Understanding Addition and Subtraction in Transformers
di: Quirke, Philip, et al.
Pubblicazione: (2024) -
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
di: Lan, Michael, et al.
Pubblicazione: (2023) -
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
di: Chaudhary, Maheep, et al.
Pubblicazione: (2025) -
Same Question, Different Words: A Latent Adversarial Framework for Prompt Robustness
di: Fu, Tingchen, et al.
Pubblicazione: (2025) -
Understanding Addition in Transformers
di: Quirke, Philip, et al.
Pubblicazione: (2023)