Representation Shattering in Transformers: A Synthetic Study with Knowledge Editing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nishi, Kento, Ramesh, Rahul, Okawa, Maya, Khona, Mikail, Tanaka, Hidenori, Lubana, Ekdeep Singh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards an Understanding of Stepwise Inference in Transformers: A Synthetic Graph Navigation Model
von: Khona, Mikail, et al.
Veröffentlicht: (2024)
von: Khona, Mikail, et al.
Veröffentlicht: (2024)
Compositional Capabilities of Autoregressive Transformers: A Study on Synthetic, Interpretable Tasks
von: Ramesh, Rahul, et al.
Veröffentlicht: (2023)
von: Ramesh, Rahul, et al.
Veröffentlicht: (2023)
Compositional Abilities Emerge Multiplicatively: Exploring Diffusion Models on a Synthetic Task
von: Okawa, Maya, et al.
Veröffentlicht: (2023)
von: Okawa, Maya, et al.
Veröffentlicht: (2023)
ICLR: In-Context Learning of Representations
von: Park, Core Francisco, et al.
Veröffentlicht: (2024)
von: Park, Core Francisco, et al.
Veröffentlicht: (2024)
Emergence of Hidden Capabilities: Exploring Learning Dynamics in Concept Space
von: Park, Core Francisco, et al.
Veröffentlicht: (2024)
von: Park, Core Francisco, et al.
Veröffentlicht: (2024)
Swing-by Dynamics in Concept Learning and Compositional Generalization
von: Yang, Yongyi, et al.
Veröffentlicht: (2024)
von: Yang, Yongyi, et al.
Veröffentlicht: (2024)
A Percolation Model of Emergence: Analyzing Transformers Trained on a Formal Language
von: Lubana, Ekdeep Singh, et al.
Veröffentlicht: (2024)
von: Lubana, Ekdeep Singh, et al.
Veröffentlicht: (2024)
Emergence of Hierarchical Emotion Organization in Large Language Models
von: Zhao, Bo, et al.
Veröffentlicht: (2025)
von: Zhao, Bo, et al.
Veröffentlicht: (2025)
Abrupt Learning in Transformers: A Case Study on Matrix Completion
von: Gopalani, Pulkit, et al.
Veröffentlicht: (2024)
von: Gopalani, Pulkit, et al.
Veröffentlicht: (2024)
Competition Dynamics Shape Algorithmic Phases of In-Context Learning
von: Park, Core Francisco, et al.
Veröffentlicht: (2024)
von: Park, Core Francisco, et al.
Veröffentlicht: (2024)
In-Context Learning of Energy Functions
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2024)
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2024)
In-Context Learning Dynamics with Random Binary Sequences
von: Bigelow, Eric J., et al.
Veröffentlicht: (2023)
von: Bigelow, Eric J., et al.
Veröffentlicht: (2023)
In-Context Learning Strategies Emerge Rationally
von: Wurgaft, Daniel, et al.
Veröffentlicht: (2025)
von: Wurgaft, Daniel, et al.
Veröffentlicht: (2025)
From Flat to Hierarchical: Extracting Sparse Representations with Matching Pursuit
von: Costa, Valérie, et al.
Veröffentlicht: (2025)
von: Costa, Valérie, et al.
Veröffentlicht: (2025)
How Do LLMs Persuade? Linear Probes Can Uncover Persuasion Dynamics in Multi-Turn Conversations
von: Jaipersaud, Brandon, et al.
Veröffentlicht: (2025)
von: Jaipersaud, Brandon, et al.
Veröffentlicht: (2025)
Analyzing (In)Abilities of SAEs via Formal Languages
von: Menon, Abhinav, et al.
Veröffentlicht: (2024)
von: Menon, Abhinav, et al.
Veröffentlicht: (2024)
Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering
von: Bigelow, Eric, et al.
Veröffentlicht: (2025)
von: Bigelow, Eric, et al.
Veröffentlicht: (2025)
Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks
von: Jain, Samyak, et al.
Veröffentlicht: (2023)
von: Jain, Samyak, et al.
Veröffentlicht: (2023)
Evaluating Sparse Autoencoders: From Shallow Design to Matching Pursuit
von: Costa, Valérie, et al.
Veröffentlicht: (2025)
von: Costa, Valérie, et al.
Veröffentlicht: (2025)
Projecting Assumptions: The Duality Between Sparse Autoencoders and Concept Geometry
von: Hindupur, Sai Sumedh R., et al.
Veröffentlicht: (2025)
von: Hindupur, Sai Sumedh R., et al.
Veröffentlicht: (2025)
Aggregated Multi-output Gaussian Processes with Knowledge Transfer Across Domains
von: Tanaka, Yusuke, et al.
Veröffentlicht: (2022)
von: Tanaka, Yusuke, et al.
Veröffentlicht: (2022)
What Makes and Breaks Safety Fine-tuning? A Mechanistic Study
von: Jain, Samyak, et al.
Veröffentlicht: (2024)
von: Jain, Samyak, et al.
Veröffentlicht: (2024)
The Impact of Off-Policy Training Data on Probe Generalisation
von: Kirch, Nathalie, et al.
Veröffentlicht: (2025)
von: Kirch, Nathalie, et al.
Veröffentlicht: (2025)
From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts?
von: Mueller, Aaron, et al.
Veröffentlicht: (2025)
von: Mueller, Aaron, et al.
Veröffentlicht: (2025)
Detecting High-Stakes Interactions with Activation Probes
von: McKenzie, Alex, et al.
Veröffentlicht: (2025)
von: McKenzie, Alex, et al.
Veröffentlicht: (2025)
Uncovering Latent Memories: Assessing Data Leakage and Memorization Patterns in Frontier AI Models
von: Duan, Sunny, et al.
Veröffentlicht: (2024)
von: Duan, Sunny, et al.
Veröffentlicht: (2024)
Provable Low-Frequency Bias of In-Context Learning of Representations
von: Yang, Yongyi, et al.
Veröffentlicht: (2025)
von: Yang, Yongyi, et al.
Veröffentlicht: (2025)
Features as Rewards: Scalable Supervision for Open-Ended Tasks via Interpretability
von: Prasad, Aaditya Vikram, et al.
Veröffentlicht: (2026)
von: Prasad, Aaditya Vikram, et al.
Veröffentlicht: (2026)
Stable Minima of ReLU Neural Networks Suffer from the Curse of Dimensionality: The Neural Shattering Phenomenon
von: Liang, Tongtong, et al.
Veröffentlicht: (2025)
von: Liang, Tongtong, et al.
Veröffentlicht: (2025)
Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention
von: Huang, Jing, et al.
Veröffentlicht: (2026)
von: Huang, Jing, et al.
Veröffentlicht: (2026)
Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space
von: Bigelow, Eric, et al.
Veröffentlicht: (2026)
von: Bigelow, Eric, et al.
Veröffentlicht: (2026)
Meta-Learning for Neural Network-based Temporal Point Processes
von: Takimoto, Yoshiaki, et al.
Veröffentlicht: (2024)
von: Takimoto, Yoshiaki, et al.
Veröffentlicht: (2024)
Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior
von: Wurgaft, Daniel, et al.
Veröffentlicht: (2026)
von: Wurgaft, Daniel, et al.
Veröffentlicht: (2026)
Continuous-Time Analysis of Adaptive Optimization and Normalization
von: Gould, Rhys, et al.
Veröffentlicht: (2024)
von: Gould, Rhys, et al.
Veröffentlicht: (2024)
$\textit{New News}$: System-2 Fine-tuning for Robust Integration of New Knowledge
von: Park, Core Francisco, et al.
Veröffentlicht: (2025)
von: Park, Core Francisco, et al.
Veröffentlicht: (2025)
Shattered Compositionality: Counterintuitive Learning Dynamics of Transformers for Arithmetic
von: Zhao, Xingyu, et al.
Veröffentlicht: (2026)
von: Zhao, Xingyu, et al.
Veröffentlicht: (2026)
Model Zoo: A Growing "Brain" That Learns Continually
von: Ramesh, Rahul, et al.
Veröffentlicht: (2021)
von: Ramesh, Rahul, et al.
Veröffentlicht: (2021)
Cross-patient Seizure Onset Zone Classification by Patient-Dependent Weight
von: Zhao, Xuyang, et al.
Veröffentlicht: (2025)
von: Zhao, Xuyang, et al.
Veröffentlicht: (2025)
MCEL: Margin-Based Cross-Entropy Loss for Error-Tolerant Quantized Neural Networks
von: Yayla, Mikail, et al.
Veröffentlicht: (2026)
von: Yayla, Mikail, et al.
Veröffentlicht: (2026)
General Transform: A Unified Framework for Adaptive Transform to Enhance Representations
von: Budiutama, Gekko, et al.
Veröffentlicht: (2025)
von: Budiutama, Gekko, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Towards an Understanding of Stepwise Inference in Transformers: A Synthetic Graph Navigation Model
von: Khona, Mikail, et al.
Veröffentlicht: (2024) -
Compositional Capabilities of Autoregressive Transformers: A Study on Synthetic, Interpretable Tasks
von: Ramesh, Rahul, et al.
Veröffentlicht: (2023) -
Compositional Abilities Emerge Multiplicatively: Exploring Diffusion Models on a Synthetic Task
von: Okawa, Maya, et al.
Veröffentlicht: (2023) -
ICLR: In-Context Learning of Representations
von: Park, Core Francisco, et al.
Veröffentlicht: (2024) -
Emergence of Hidden Capabilities: Exploring Learning Dynamics in Concept Space
von: Park, Core Francisco, et al.
Veröffentlicht: (2024)