Arithmetic in the Wild: Llama uses Base-10 Addition to Reason About Cyclic Concepts
Fuente:
arXiv
Salvato in:
| Autori principali: | Feucht, Sheridan, Haklay, Tal, Bhalla, Usha, Wurgaft, Daniel, Rager, Can, Sarfati, Raphaël, Merullo, Jack, McGrath, Thomas, Lewis, Owen, Lubana, Ekdeep Singh, Fel, Thomas, Geiger, Atticus |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Do Sparse Autoencoders Capture Concept Manifolds?
di: Bhalla, Usha, et al.
Pubblicazione: (2026)
di: Bhalla, Usha, et al.
Pubblicazione: (2026)
Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior
di: Wurgaft, Daniel, et al.
Pubblicazione: (2026)
di: Wurgaft, Daniel, et al.
Pubblicazione: (2026)
Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space
di: Bigelow, Eric, et al.
Pubblicazione: (2026)
di: Bigelow, Eric, et al.
Pubblicazione: (2026)
The Shape of Beliefs: Geometry, Dynamics, and Interventions along Representation Manifolds of Language Models' Posteriors
di: Sarfati, Raphaël, et al.
Pubblicazione: (2026)
di: Sarfati, Raphaël, et al.
Pubblicazione: (2026)
Features as Rewards: Scalable Supervision for Open-Ended Tasks via Interpretability
di: Prasad, Aaditya Vikram, et al.
Pubblicazione: (2026)
di: Prasad, Aaditya Vikram, et al.
Pubblicazione: (2026)
Are language models aware of the road not taken? Token-level uncertainty and hidden state dynamics
di: Zur, Amir, et al.
Pubblicazione: (2025)
di: Zur, Amir, et al.
Pubblicazione: (2025)
Uncovering Conceptual Blindspots in Generative Image Models Using Sparse Autoencoders
di: Bohacek, Matyas, et al.
Pubblicazione: (2025)
di: Bohacek, Matyas, et al.
Pubblicazione: (2025)
Reasoning Theater: Disentangling Model Beliefs from Chain-of-Thought
di: Boppana, Siddharth, et al.
Pubblicazione: (2026)
di: Boppana, Siddharth, et al.
Pubblicazione: (2026)
Projecting Assumptions: The Duality Between Sparse Autoencoders and Concept Geometry
di: Hindupur, Sai Sumedh R., et al.
Pubblicazione: (2025)
di: Hindupur, Sai Sumedh R., et al.
Pubblicazione: (2025)
From Flat to Hierarchical: Extracting Sparse Representations with Matching Pursuit
di: Costa, Valérie, et al.
Pubblicazione: (2025)
di: Costa, Valérie, et al.
Pubblicazione: (2025)
Evaluating Sparse Autoencoders: From Shallow Design to Matching Pursuit
di: Costa, Valérie, et al.
Pubblicazione: (2025)
di: Costa, Valérie, et al.
Pubblicazione: (2025)
Vector Arithmetic in Concept and Token Subspaces
di: Feucht, Sheridan, et al.
Pubblicazione: (2025)
di: Feucht, Sheridan, et al.
Pubblicazione: (2025)
In-Context Learning Strategies Emerge Rationally
di: Wurgaft, Daniel, et al.
Pubblicazione: (2025)
di: Wurgaft, Daniel, et al.
Pubblicazione: (2025)
How Causal Abstraction Underpins Computational Explanation
di: Geiger, Atticus, et al.
Pubblicazione: (2025)
di: Geiger, Atticus, et al.
Pubblicazione: (2025)
Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering
di: Bigelow, Eric, et al.
Pubblicazione: (2025)
di: Bigelow, Eric, et al.
Pubblicazione: (2025)
Priors in Time: Missing Inductive Biases for Language Model Interpretability
di: Lubana, Ekdeep Singh, et al.
Pubblicazione: (2025)
di: Lubana, Ekdeep Singh, et al.
Pubblicazione: (2025)
How Do LLMs Persuade? Linear Probes Can Uncover Persuasion Dynamics in Multi-Turn Conversations
di: Jaipersaud, Brandon, et al.
Pubblicazione: (2025)
di: Jaipersaud, Brandon, et al.
Pubblicazione: (2025)
Abrupt Learning in Transformers: A Case Study on Matrix Completion
di: Gopalani, Pulkit, et al.
Pubblicazione: (2024)
di: Gopalani, Pulkit, et al.
Pubblicazione: (2024)
Binomial confidence intervals for rare events: importance of defining margin of error relative to magnitude of proportion
di: McGrath, Owen, et al.
Pubblicazione: (2021)
di: McGrath, Owen, et al.
Pubblicazione: (2021)
A General Simulation-Based Optimisation Framework for Multipoint Constant-Stress Accelerated Life Tests
di: McGrath, Owen, et al.
Pubblicazione: (2025)
di: McGrath, Owen, et al.
Pubblicazione: (2025)
A General Simulation‐Based Optimisation Framework for Multipoint Constant‐Stress Accelerated Life Tests
di: Owen McGrath, et al.
Pubblicazione: (2025)
di: Owen McGrath, et al.
Pubblicazione: (2025)
Comparisons of the 1995 and 1998 coral bleaching events on the patch reefs of San Salvador Island, Bahamas
di: Thomas A. McGrath
Pubblicazione: (2003)
di: Thomas A. McGrath
Pubblicazione: (2003)
Evaluating Open-Source Sparse Autoencoders on Disentangling Factual Knowledge in GPT-2 Small
di: Chaudhary, Maheep, et al.
Pubblicazione: (2024)
di: Chaudhary, Maheep, et al.
Pubblicazione: (2024)
Archetypal SAE: Adaptive and Stable Dictionary Learning for Concept Extraction in Large Vision Models
di: Fel, Thomas, et al.
Pubblicazione: (2025)
di: Fel, Thomas, et al.
Pubblicazione: (2025)
Sparks of Explainability: Recent Advancements in Explaining Large Vision Models
di: Fel, Thomas
Pubblicazione: (2025)
di: Fel, Thomas
Pubblicazione: (2025)
Constructing Interpretable Features from Compositional Neuron Groups
di: Shafran, Or, et al.
Pubblicazione: (2025)
di: Shafran, Or, et al.
Pubblicazione: (2025)
Towards Reliable Evaluation of Behavior Steering Interventions in LLMs
di: Pres, Itamar, et al.
Pubblicazione: (2024)
di: Pres, Itamar, et al.
Pubblicazione: (2024)
Analyzing (In)Abilities of SAEs via Formal Languages
di: Menon, Abhinav, et al.
Pubblicazione: (2024)
di: Menon, Abhinav, et al.
Pubblicazione: (2024)
Language Models Implement Simple Word2Vec-style Vector Arithmetic
di: Merullo, Jack, et al.
Pubblicazione: (2023)
di: Merullo, Jack, et al.
Pubblicazione: (2023)
Into the Rabbit Hull: From Task-Relevant Concepts in DINO to Minkowski Geometry
di: Fel, Thomas, et al.
Pubblicazione: (2025)
di: Fel, Thomas, et al.
Pubblicazione: (2025)
Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention
di: Huang, Jing, et al.
Pubblicazione: (2026)
di: Huang, Jing, et al.
Pubblicazione: (2026)
A Percolation Model of Emergence: Analyzing Transformers Trained on a Formal Language
di: Lubana, Ekdeep Singh, et al.
Pubblicazione: (2024)
di: Lubana, Ekdeep Singh, et al.
Pubblicazione: (2024)
Competition Dynamics Shape Algorithmic Phases of In-Context Learning
di: Park, Core Francisco, et al.
Pubblicazione: (2024)
di: Park, Core Francisco, et al.
Pubblicazione: (2024)
Compositional Abilities Emerge Multiplicatively: Exploring Diffusion Models on a Synthetic Task
di: Okawa, Maya, et al.
Pubblicazione: (2023)
di: Okawa, Maya, et al.
Pubblicazione: (2023)
How Do Transformers Learn Variable Binding in Symbolic Programs?
di: Wu, Yiwei, et al.
Pubblicazione: (2025)
di: Wu, Yiwei, et al.
Pubblicazione: (2025)
Bucketing the Good Apples: A Method for Diagnosing and Improving Causal Abstraction
di: Puyin, Li, et al.
Pubblicazione: (2026)
di: Puyin, Li, et al.
Pubblicazione: (2026)
Discriminative Feature Attributions: Bridging Post Hoc Explainability and Inherent Interpretability
di: Bhalla, Usha, et al.
Pubblicazione: (2023)
di: Bhalla, Usha, et al.
Pubblicazione: (2023)
The Dual-Route Model of Induction
di: Feucht, Sheridan, et al.
Pubblicazione: (2025)
di: Feucht, Sheridan, et al.
Pubblicazione: (2025)
Erasing Conceptual Knowledge from Language Models
di: Gandikota, Rohit, et al.
Pubblicazione: (2024)
di: Gandikota, Rohit, et al.
Pubblicazione: (2024)
Token Erasure as a Footprint of Implicit Vocabulary Items in LLMs
di: Feucht, Sheridan, et al.
Pubblicazione: (2024)
di: Feucht, Sheridan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Do Sparse Autoencoders Capture Concept Manifolds?
di: Bhalla, Usha, et al.
Pubblicazione: (2026) -
Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior
di: Wurgaft, Daniel, et al.
Pubblicazione: (2026) -
Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space
di: Bigelow, Eric, et al.
Pubblicazione: (2026) -
The Shape of Beliefs: Geometry, Dynamics, and Interventions along Representation Manifolds of Language Models' Posteriors
di: Sarfati, Raphaël, et al.
Pubblicazione: (2026) -
Features as Rewards: Scalable Supervision for Open-Ended Tasks via Interpretability
di: Prasad, Aaditya Vikram, et al.
Pubblicazione: (2026)