Transferring Linear Features Across Language Models With Model Stitching
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Chen, Alan, Merullo, Jack, Stolfo, Alessandro, Pavlick, Ellie |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Circuit Component Reuse Across Tasks in Transformer Language Models
par: Merullo, Jack, et autres
Publié: (2023)
par: Merullo, Jack, et autres
Publié: (2023)
Language Models Implement Simple Word2Vec-style Vector Arithmetic
par: Merullo, Jack, et autres
Publié: (2023)
par: Merullo, Jack, et autres
Publié: (2023)
Dual Process Learning: Controlling Use of In-Context vs. In-Weights Strategies with Weight Forgetting
par: Anand, Suraj, et autres
Publié: (2024)
par: Anand, Suraj, et autres
Publié: (2024)
Talking Heads: Understanding Inter-layer Communication in Transformer Language Models
par: Merullo, Jack, et autres
Publié: (2024)
par: Merullo, Jack, et autres
Publié: (2024)
How Do Vision-Language Models Process Conflicting Information Across Modalities?
par: Hua, Tianze, et autres
Publié: (2025)
par: Hua, Tianze, et autres
Publié: (2025)
$100K or 100 Days: Trade-offs when Pre-Training with Academic Resources
par: Khandelwal, Apoorv, et autres
Publié: (2024)
par: Khandelwal, Apoorv, et autres
Publié: (2024)
Does Training on Synthetic Data Make Models Less Robust?
par: Zhang, Lingze, et autres
Publié: (2025)
par: Zhang, Lingze, et autres
Publié: (2025)
Groundedness in Retrieval-augmented Long-form Generation: An Empirical Study
par: Stolfo, Alessandro
Publié: (2024)
par: Stolfo, Alessandro
Publié: (2024)
Improving Instruction-Following in Language Models through Activation Steering
par: Stolfo, Alessandro, et autres
Publié: (2024)
par: Stolfo, Alessandro, et autres
Publié: (2024)
I Have No Mouth, and I Must Rhyme: Uncovering Internal Phonetic Representations in LLaMA 3.2
par: McLaughlin, Oliver, et autres
Publié: (2025)
par: McLaughlin, Oliver, et autres
Publié: (2025)
Confidence Regulation Neurons in Language Models
par: Stolfo, Alessandro, et autres
Publié: (2024)
par: Stolfo, Alessandro, et autres
Publié: (2024)
From Memorization to Reasoning in the Spectrum of Loss Curvature
par: Merullo, Jack, et autres
Publié: (2025)
par: Merullo, Jack, et autres
Publié: (2025)
How Do Language Models Compose Functions?
par: Khandelwal, Apoorv, et autres
Publié: (2025)
par: Khandelwal, Apoorv, et autres
Publié: (2025)
Dense SAE Latents Are Features, Not Bugs
par: Sun, Xiaoqing, et autres
Publié: (2025)
par: Sun, Xiaoqing, et autres
Publié: (2025)
Do Language Models Exhibit the Same Cognitive Biases in Problem Solving as Human Learners?
par: Opedal, Andreas, et autres
Publié: (2024)
par: Opedal, Andreas, et autres
Publié: (2024)
Born a Transformer -- Always a Transformer? On the Effect of Pretraining on Architectural Abilities
par: Jobanputra, Mayank, et autres
Publié: (2025)
par: Jobanputra, Mayank, et autres
Publié: (2025)
What is an "Abstract Reasoner"? Revisiting Experiments and Arguments about Large Language Models
par: Yun, Tian, et autres
Publié: (2025)
par: Yun, Tian, et autres
Publié: (2025)
Feature Alignment and Representation Transfer in Knowledge Distillation for Large Language Models
par: Yang, Junjie, et autres
Publié: (2025)
par: Yang, Junjie, et autres
Publié: (2025)
Does CLIP Bind Concepts? Probing Compositionality in Large Image Models
par: Lewis, Martha, et autres
Publié: (2022)
par: Lewis, Martha, et autres
Publié: (2022)
Source-Modality Monitoring in Vision-Language Models
par: Hua, Etha Tianze, et autres
Publié: (2026)
par: Hua, Etha Tianze, et autres
Publié: (2026)
mOthello: When Do Cross-Lingual Representation Alignment and Cross-Lingual Transfer Emerge in Multilingual Models?
par: Hua, Tianze, et autres
Publié: (2024)
par: Hua, Tianze, et autres
Publié: (2024)
R-Stitch: Dynamic Trajectory Stitching for Efficient Reasoning
par: Chen, Zhuokun, et autres
Publié: (2025)
par: Chen, Zhuokun, et autres
Publié: (2025)
Reasoning Theater: Disentangling Model Beliefs from Chain-of-Thought
par: Boppana, Siddharth, et autres
Publié: (2026)
par: Boppana, Siddharth, et autres
Publié: (2026)
Atlas-Alignment: Making Interpretability Transferable Across Language Models
par: Puri, Bruno, et autres
Publié: (2025)
par: Puri, Bruno, et autres
Publié: (2025)
On Linear Representations and Pretraining Data Frequency in Language Models
par: Merullo, Jack, et autres
Publié: (2025)
par: Merullo, Jack, et autres
Publié: (2025)
Sparse Autoencoder Features for Classifications and Transferability
par: Gallifant, Jack, et autres
Publié: (2025)
par: Gallifant, Jack, et autres
Publié: (2025)
Probing for Arithmetic Errors in Language Models
par: Sun, Yucheng, et autres
Publié: (2025)
par: Sun, Yucheng, et autres
Publié: (2025)
Linear Dynamics in the RLVR Training of Large Language Models
par: Wang, Tianle, et autres
Publié: (2026)
par: Wang, Tianle, et autres
Publié: (2026)
Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders
par: Lan, Michael, et autres
Publié: (2024)
par: Lan, Michael, et autres
Publié: (2024)
Replacing Language Model for Style Transfer
par: Cheng, Pengyu, et autres
Publié: (2022)
par: Cheng, Pengyu, et autres
Publié: (2022)
On the Origins of Linear Representations in Large Language Models
par: Jiang, Yibo, et autres
Publié: (2024)
par: Jiang, Yibo, et autres
Publié: (2024)
Transfer Learning for Finetuning Large Language Models
par: Strangmann, Tobias, et autres
Publié: (2024)
par: Strangmann, Tobias, et autres
Publié: (2024)
Generalizing Large Language Model Usability Across Resource-Constrained
par: Tsai, Yun-Da
Publié: (2025)
par: Tsai, Yun-Da
Publié: (2025)
Can Large Language Models Generalize Procedures Across Representations?
par: Lin, Fangru, et autres
Publié: (2026)
par: Lin, Fangru, et autres
Publié: (2026)
Invariant Features in Language Models: Geometric Characterization and Model Attribution
par: Dasgupta, Agnibh, et autres
Publié: (2026)
par: Dasgupta, Agnibh, et autres
Publié: (2026)
Persona Vectors: Monitoring and Controlling Character Traits in Language Models
par: Chen, Runjin, et autres
Publié: (2025)
par: Chen, Runjin, et autres
Publié: (2025)
Localize-and-Stitch: Efficient Model Merging via Sparse Task Arithmetic
par: He, Yifei, et autres
Publié: (2024)
par: He, Yifei, et autres
Publié: (2024)
LoLCATs: On Low-Rank Linearizing of Large Language Models
par: Zhang, Michael, et autres
Publié: (2024)
par: Zhang, Michael, et autres
Publié: (2024)
Rotary Offset Features in Large Language Models
par: Jonasson, André
Publié: (2025)
par: Jonasson, André
Publié: (2025)
Can LLMs subtract numbers?
par: Jobanputra, Mayank, et autres
Publié: (2025)
par: Jobanputra, Mayank, et autres
Publié: (2025)
Documents similaires
-
Circuit Component Reuse Across Tasks in Transformer Language Models
par: Merullo, Jack, et autres
Publié: (2023) -
Language Models Implement Simple Word2Vec-style Vector Arithmetic
par: Merullo, Jack, et autres
Publié: (2023) -
Dual Process Learning: Controlling Use of In-Context vs. In-Weights Strategies with Weight Forgetting
par: Anand, Suraj, et autres
Publié: (2024) -
Talking Heads: Understanding Inter-layer Communication in Transformer Language Models
par: Merullo, Jack, et autres
Publié: (2024) -
How Do Vision-Language Models Process Conflicting Information Across Modalities?
par: Hua, Tianze, et autres
Publié: (2025)