Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks
Fuente:
arXiv
Guardado en:
| Autores principales: | Jain, Samyak, Kirk, Robert, Lubana, Ekdeep Singh, Dick, Robert P., Tanaka, Hidenori, Grefenstette, Edward, Rocktäschel, Tim, Krueger, David Scott |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A Percolation Model of Emergence: Analyzing Transformers Trained on a Formal Language
por: Lubana, Ekdeep Singh, et al.
Publicado: (2024)
por: Lubana, Ekdeep Singh, et al.
Publicado: (2024)
Compositional Abilities Emerge Multiplicatively: Exploring Diffusion Models on a Synthetic Task
por: Okawa, Maya, et al.
Publicado: (2023)
por: Okawa, Maya, et al.
Publicado: (2023)
What Makes and Breaks Safety Fine-tuning? A Mechanistic Study
por: Jain, Samyak, et al.
Publicado: (2024)
por: Jain, Samyak, et al.
Publicado: (2024)
Compositional Capabilities of Autoregressive Transformers: A Study on Synthetic, Interpretable Tasks
por: Ramesh, Rahul, et al.
Publicado: (2023)
por: Ramesh, Rahul, et al.
Publicado: (2023)
How Do LLMs Persuade? Linear Probes Can Uncover Persuasion Dynamics in Multi-Turn Conversations
por: Jaipersaud, Brandon, et al.
Publicado: (2025)
por: Jaipersaud, Brandon, et al.
Publicado: (2025)
In-Context Learning Dynamics with Random Binary Sequences
por: Bigelow, Eric J., et al.
Publicado: (2023)
por: Bigelow, Eric J., et al.
Publicado: (2023)
Competition Dynamics Shape Algorithmic Phases of In-Context Learning
por: Park, Core Francisco, et al.
Publicado: (2024)
por: Park, Core Francisco, et al.
Publicado: (2024)
Towards Reliable Evaluation of Behavior Steering Interventions in LLMs
por: Pres, Itamar, et al.
Publicado: (2024)
por: Pres, Itamar, et al.
Publicado: (2024)
Analyzing (In)Abilities of SAEs via Formal Languages
por: Menon, Abhinav, et al.
Publicado: (2024)
por: Menon, Abhinav, et al.
Publicado: (2024)
Towards an Understanding of Stepwise Inference in Transformers: A Synthetic Graph Navigation Model
por: Khona, Mikail, et al.
Publicado: (2024)
por: Khona, Mikail, et al.
Publicado: (2024)
minimax: Efficient Baselines for Autocurricula in JAX
por: Jiang, Minqi, et al.
Publicado: (2023)
por: Jiang, Minqi, et al.
Publicado: (2023)
Emergence of Hidden Capabilities: Exploring Learning Dynamics in Concept Space
por: Park, Core Francisco, et al.
Publicado: (2024)
por: Park, Core Francisco, et al.
Publicado: (2024)
Representation Shattering in Transformers: A Synthetic Study with Knowledge Editing
por: Nishi, Kento, et al.
Publicado: (2024)
por: Nishi, Kento, et al.
Publicado: (2024)
Swing-by Dynamics in Concept Learning and Compositional Generalization
por: Yang, Yongyi, et al.
Publicado: (2024)
por: Yang, Yongyi, et al.
Publicado: (2024)
In-Context Learning Strategies Emerge Rationally
por: Wurgaft, Daniel, et al.
Publicado: (2025)
por: Wurgaft, Daniel, et al.
Publicado: (2025)
Abrupt Learning in Transformers: A Case Study on Matrix Completion
por: Gopalani, Pulkit, et al.
Publicado: (2024)
por: Gopalani, Pulkit, et al.
Publicado: (2024)
Investigating Non-Transitivity in LLM-as-a-Judge
por: Xu, Yi, et al.
Publicado: (2025)
por: Xu, Yi, et al.
Publicado: (2025)
Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering
por: Bigelow, Eric, et al.
Publicado: (2025)
por: Bigelow, Eric, et al.
Publicado: (2025)
Emergence of Hierarchical Emotion Organization in Large Language Models
por: Zhao, Bo, et al.
Publicado: (2025)
por: Zhao, Bo, et al.
Publicado: (2025)
Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models
por: Ruis, Laura, et al.
Publicado: (2024)
por: Ruis, Laura, et al.
Publicado: (2024)
ICLR: In-Context Learning of Representations
por: Park, Core Francisco, et al.
Publicado: (2024)
por: Park, Core Francisco, et al.
Publicado: (2024)
Detecting High-Stakes Interactions with Activation Probes
por: McKenzie, Alex, et al.
Publicado: (2025)
por: McKenzie, Alex, et al.
Publicado: (2025)
Are language models aware of the road not taken? Token-level uncertainty and hidden state dynamics
por: Zur, Amir, et al.
Publicado: (2025)
por: Zur, Amir, et al.
Publicado: (2025)
Uncovering Conceptual Blindspots in Generative Image Models Using Sparse Autoencoders
por: Bohacek, Matyas, et al.
Publicado: (2025)
por: Bohacek, Matyas, et al.
Publicado: (2025)
Infusion: Shaping Model Behavior by Editing Training Data via Influence Functions
por: Rosser, J, et al.
Publicado: (2026)
por: Rosser, J, et al.
Publicado: (2026)
Projecting Assumptions: The Duality Between Sparse Autoencoders and Concept Geometry
por: Hindupur, Sai Sumedh R., et al.
Publicado: (2025)
por: Hindupur, Sai Sumedh R., et al.
Publicado: (2025)
From Flat to Hierarchical: Extracting Sparse Representations with Matching Pursuit
por: Costa, Valérie, et al.
Publicado: (2025)
por: Costa, Valérie, et al.
Publicado: (2025)
Evaluating Sparse Autoencoders: From Shallow Design to Matching Pursuit
por: Costa, Valérie, et al.
Publicado: (2025)
por: Costa, Valérie, et al.
Publicado: (2025)
Scaling Opponent Shaping to High Dimensional Games
por: Khan, Akbir, et al.
Publicado: (2023)
por: Khan, Akbir, et al.
Publicado: (2023)
The Impact of Off-Policy Training Data on Probe Generalisation
por: Kirch, Nathalie, et al.
Publicado: (2025)
por: Kirch, Nathalie, et al.
Publicado: (2025)
From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts?
por: Mueller, Aaron, et al.
Publicado: (2025)
por: Mueller, Aaron, et al.
Publicado: (2025)
Reward Model Ensembles Help Mitigate Overoptimization
por: Coste, Thomas, et al.
Publicado: (2023)
por: Coste, Thomas, et al.
Publicado: (2023)
Force Dipole Interactions in Tubular Fluid Membranes
por: Jain, Samyak, et al.
Publicado: (2023)
por: Jain, Samyak, et al.
Publicado: (2023)
Nuclear stability and the Fold Catastrophe
por: Jain, Samyak, et al.
Publicado: (2023)
por: Jain, Samyak, et al.
Publicado: (2023)
Tunneling half-lives in macroscopic-microscopic picture
por: Jain, Samyak, et al.
Publicado: (2024)
por: Jain, Samyak, et al.
Publicado: (2024)
Catastrophe theoretic approach to the Higgs Mechanism
por: Jain, Samyak, et al.
Publicado: (2023)
por: Jain, Samyak, et al.
Publicado: (2023)
Understanding the Effects of RLHF on LLM Generalisation and Diversity
por: Kirk, Robert, et al.
Publicado: (2023)
por: Kirk, Robert, et al.
Publicado: (2023)
Interaction Dynamics as a Reward Signal for LLMs
por: Gooding, Sian, et al.
Publicado: (2025)
por: Gooding, Sian, et al.
Publicado: (2025)
Learning When to Plan: Efficiently Allocating Test-Time Compute for LLM Agents
por: Paglieri, Davide, et al.
Publicado: (2025)
por: Paglieri, Davide, et al.
Publicado: (2025)
Debating with More Persuasive LLMs Leads to More Truthful Answers
por: Khan, Akbir, et al.
Publicado: (2024)
por: Khan, Akbir, et al.
Publicado: (2024)
Ejemplares similares
-
A Percolation Model of Emergence: Analyzing Transformers Trained on a Formal Language
por: Lubana, Ekdeep Singh, et al.
Publicado: (2024) -
Compositional Abilities Emerge Multiplicatively: Exploring Diffusion Models on a Synthetic Task
por: Okawa, Maya, et al.
Publicado: (2023) -
What Makes and Breaks Safety Fine-tuning? A Mechanistic Study
por: Jain, Samyak, et al.
Publicado: (2024) -
Compositional Capabilities of Autoregressive Transformers: A Study on Synthetic, Interpretable Tasks
por: Ramesh, Rahul, et al.
Publicado: (2023) -
How Do LLMs Persuade? Linear Probes Can Uncover Persuasion Dynamics in Multi-Turn Conversations
por: Jaipersaud, Brandon, et al.
Publicado: (2025)