Understanding Addition and Subtraction in Transformers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Quirke, Philip, Neo, Clement, Barez, Fazl |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Understanding Addition in Transformers
von: Quirke, Philip, et al.
Veröffentlicht: (2023)
von: Quirke, Philip, et al.
Veröffentlicht: (2023)
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions
von: Neo, Clement, et al.
Veröffentlicht: (2024)
von: Neo, Clement, et al.
Veröffentlicht: (2024)
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2025)
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2025)
Same Question, Different Words: A Latent Adversarial Framework for Prompt Robustness
von: Fu, Tingchen, et al.
Veröffentlicht: (2025)
von: Fu, Tingchen, et al.
Veröffentlicht: (2025)
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
von: Lan, Michael, et al.
Veröffentlicht: (2023)
von: Lan, Michael, et al.
Veröffentlicht: (2023)
Towards Interpreting Visual Information Processing in Vision-Language Models
von: Neo, Clement, et al.
Veröffentlicht: (2024)
von: Neo, Clement, et al.
Veröffentlicht: (2024)
Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders
von: Lan, Michael, et al.
Veröffentlicht: (2024)
von: Lan, Michael, et al.
Veröffentlicht: (2024)
Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
von: Oozeer, Narmeen, et al.
Veröffentlicht: (2025)
von: Oozeer, Narmeen, et al.
Veröffentlicht: (2025)
Scaling sparse feature circuit finding for in-context learning
von: Kharlapenko, Dmitrii, et al.
Veröffentlicht: (2025)
von: Kharlapenko, Dmitrii, et al.
Veröffentlicht: (2025)
Interpreting Learned Feedback Patterns in Large Language Models
von: Marks, Luke, et al.
Veröffentlicht: (2023)
von: Marks, Luke, et al.
Veröffentlicht: (2023)
Do Sparse Autoencoders Generalize? A Case Study of Answerability
von: Heindrich, Lovis, et al.
Veröffentlicht: (2025)
von: Heindrich, Lovis, et al.
Veröffentlicht: (2025)
Towards Understanding Subliminal Learning: When and How Hidden Biases Transfer
von: Schrodi, Simon, et al.
Veröffentlicht: (2025)
von: Schrodi, Simon, et al.
Veröffentlicht: (2025)
Best-of-N Jailbreaking
von: Hughes, John, et al.
Veröffentlicht: (2024)
von: Hughes, John, et al.
Veröffentlicht: (2024)
TinySQL: A Progressive Text-to-SQL Dataset for Mechanistic Interpretability Research
von: Harrasse, Abir, et al.
Veröffentlicht: (2025)
von: Harrasse, Abir, et al.
Veröffentlicht: (2025)
Jailbreak Defense in a Narrow Domain: Limitations of Existing Methods and a New Transcript-Classifier Approach
von: Wang, Tony T., et al.
Veröffentlicht: (2024)
von: Wang, Tony T., et al.
Veröffentlicht: (2024)
VAL-Bench: Belief Consistency as a measure for Value Alignment in Language Models
von: Gupta, Aman, et al.
Veröffentlicht: (2025)
von: Gupta, Aman, et al.
Veröffentlicht: (2025)
Understanding In-context Learning of Addition via Activation Subspaces
von: Hu, Xinyan, et al.
Veröffentlicht: (2025)
von: Hu, Xinyan, et al.
Veröffentlicht: (2025)
Beyond Linear Probes: Dynamic Safety Monitoring for Language Models
von: Oldfield, James, et al.
Veröffentlicht: (2025)
von: Oldfield, James, et al.
Veröffentlicht: (2025)
Benchmark Inflation: Revealing LLM Performance Gaps Using Retro-Holdouts
von: Haimes, Jacob, et al.
Veröffentlicht: (2024)
von: Haimes, Jacob, et al.
Veröffentlicht: (2024)
Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders
von: Marks, Luke, et al.
Veröffentlicht: (2024)
von: Marks, Luke, et al.
Veröffentlicht: (2024)
DisEmbed: Transforming Disease Understanding through Embeddings
von: Faroz, Salman
Veröffentlicht: (2024)
von: Faroz, Salman
Veröffentlicht: (2024)
Generalization or Hallucination? Understanding Out-of-Context Reasoning in Transformers
von: Huang, Yixiao, et al.
Veröffentlicht: (2025)
von: Huang, Yixiao, et al.
Veröffentlicht: (2025)
Visualizing Neural Network Imagination
von: Wichers, Nevan, et al.
Veröffentlicht: (2024)
von: Wichers, Nevan, et al.
Veröffentlicht: (2024)
Theoretical Understanding of In-Context Learning in Shallow Transformers with Unstructured Data
von: Xing, Yue, et al.
Veröffentlicht: (2024)
von: Xing, Yue, et al.
Veröffentlicht: (2024)
Understanding Gated Neurons in Transformers from Their Input-Output Functionality
von: Gerstner, Sebastian, et al.
Veröffentlicht: (2025)
von: Gerstner, Sebastian, et al.
Veröffentlicht: (2025)
Learning Syntax Without Planting Trees: Understanding Hierarchical Generalization in Transformers
von: Ahuja, Kabir, et al.
Veröffentlicht: (2024)
von: Ahuja, Kabir, et al.
Veröffentlicht: (2024)
Feature Resemblance: Towards a Theoretical Understanding of Analogical Reasoning in Transformers
von: Xu, Ruichen, et al.
Veröffentlicht: (2026)
von: Xu, Ruichen, et al.
Veröffentlicht: (2026)
Understanding Contextual Recall in Transformers: How Finetuning Enables In-Context Reasoning over Pretraining Knowledge
von: Vasudeva, Bhavya, et al.
Veröffentlicht: (2026)
von: Vasudeva, Bhavya, et al.
Veröffentlicht: (2026)
Plain Transformers Can be Powerful Graph Learners
von: Ma, Liheng, et al.
Veröffentlicht: (2025)
von: Ma, Liheng, et al.
Veröffentlicht: (2025)
The Role of Logic and Automata in Understanding Transformers
von: Lin, Anthony W., et al.
Veröffentlicht: (2025)
von: Lin, Anthony W., et al.
Veröffentlicht: (2025)
Pre-trained Large Language Models Use Fourier Features to Compute Addition
von: Zhou, Tianyi, et al.
Veröffentlicht: (2024)
von: Zhou, Tianyi, et al.
Veröffentlicht: (2024)
Precise In-Parameter Concept Erasure in Large Language Models
von: Gur-Arieh, Yoav, et al.
Veröffentlicht: (2025)
von: Gur-Arieh, Yoav, et al.
Veröffentlicht: (2025)
Understanding Transformers via N-gram Statistics
von: Nguyen, Timothy
Veröffentlicht: (2024)
von: Nguyen, Timothy
Veröffentlicht: (2024)
Understanding Dynamic Compute Allocation in Recurrent Transformers
von: Moosa, Ibraheem Muhammad, et al.
Veröffentlicht: (2026)
von: Moosa, Ibraheem Muhammad, et al.
Veröffentlicht: (2026)
Quantifying the Effect of Test Set Contamination on Generative Evaluations
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2026)
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2026)
Understanding Factual Recall in Transformers via Associative Memories
von: Nichani, Eshaan, et al.
Veröffentlicht: (2024)
von: Nichani, Eshaan, et al.
Veröffentlicht: (2024)
Retrieval Backward Attention without Additional Training: Enhance Embeddings of Large Language Models via Repetition
von: Duan, Yifei, et al.
Veröffentlicht: (2025)
von: Duan, Yifei, et al.
Veröffentlicht: (2025)
Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer
von: Simhi, Adi, et al.
Veröffentlicht: (2025)
von: Simhi, Adi, et al.
Veröffentlicht: (2025)
Language Models Use Trigonometry to Do Addition
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2025)
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2025)
Beyond Monoliths: Expert Orchestration for More Capable, Democratic, and Safe Language Models
von: Quirke, Philip, et al.
Veröffentlicht: (2025)
von: Quirke, Philip, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Understanding Addition in Transformers
von: Quirke, Philip, et al.
Veröffentlicht: (2023) -
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions
von: Neo, Clement, et al.
Veröffentlicht: (2024) -
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2025) -
Same Question, Different Words: A Latent Adversarial Framework for Prompt Robustness
von: Fu, Tingchen, et al.
Veröffentlicht: (2025) -
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
von: Lan, Michael, et al.
Veröffentlicht: (2023)