Transcoder Adapters for Reasoning-Model Diffing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hu, Nathan, Ward, Jake, Icard, Thomas, Potts, Christopher |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Internal Causal Mechanisms Robustly Predict Language Model Out-of-Distribution Behaviors
von: Huang, Jing, et al.
Veröffentlicht: (2025)
von: Huang, Jing, et al.
Veröffentlicht: (2025)
Simple LLM Baselines are Competitive for Model Diffing
von: Kempf, Elias, et al.
Veröffentlicht: (2026)
von: Kempf, Elias, et al.
Veröffentlicht: (2026)
Reasoning-Finetuning Repurposes Latent Representations in Base Models
von: Ward, Jake, et al.
Veröffentlicht: (2025)
von: Ward, Jake, et al.
Veröffentlicht: (2025)
Reading the Finetuning Prior: Verbatim Content Recovery via Contrastive Decoding Diffing
von: Brzozowski, Michał, et al.
Veröffentlicht: (2026)
von: Brzozowski, Michał, et al.
Veröffentlicht: (2026)
Delta-Crosscoder: Robust Crosscoder Model Diffing in Narrow Fine-Tuning Regimes
von: Kassem, Aly, et al.
Veröffentlicht: (2026)
von: Kassem, Aly, et al.
Veröffentlicht: (2026)
Transcoders Beat Sparse Autoencoders for Interpretability
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2025)
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2025)
Rank-1 LoRAs Encode Interpretable Reasoning Signals
von: Ward, Jake, et al.
Veröffentlicht: (2025)
von: Ward, Jake, et al.
Veröffentlicht: (2025)
Cross-Architecture Model Diffing with Crosscoders: Unsupervised Discovery of Differences Between LLMs
von: Jiralerspong, Thomas, et al.
Veröffentlicht: (2026)
von: Jiralerspong, Thomas, et al.
Veröffentlicht: (2026)
Beyond Dense States: Elevating Sparse Transcoders to Active Operators for Latent Reasoning
von: Wang, Yadong, et al.
Veröffentlicht: (2026)
von: Wang, Yadong, et al.
Veröffentlicht: (2026)
A Reply to Makelov et al. (2023)'s "Interpretability Illusion" Arguments
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2024)
von: Wu, Zhengxuan, et al.
Veröffentlicht: (2024)
Transcoder-based Circuit Analysis for Interpretable Single-Cell Foundation Models
von: Hosokawa, Sosuke, et al.
Veröffentlicht: (2025)
von: Hosokawa, Sosuke, et al.
Veröffentlicht: (2025)
Transcoders Find Interpretable LLM Feature Circuits
von: Dunefsky, Jacob, et al.
Veröffentlicht: (2024)
von: Dunefsky, Jacob, et al.
Veröffentlicht: (2024)
FedDifRC: Unlocking the Potential of Text-to-Image Diffusion Models in Heterogeneous Federated Learning
von: Wang, Huan, et al.
Veröffentlicht: (2025)
von: Wang, Huan, et al.
Veröffentlicht: (2025)
How Causal Abstraction Underpins Computational Explanation
von: Geiger, Atticus, et al.
Veröffentlicht: (2025)
von: Geiger, Atticus, et al.
Veröffentlicht: (2025)
DifFaiRec: Generative Fair Recommender with Conditional Diffusion Model
von: Jiang, Zhenhao, et al.
Veröffentlicht: (2024)
von: Jiang, Zhenhao, et al.
Veröffentlicht: (2024)
Transcoders Trace Visual Grounding and Hallucinations in Vision-Language Models
von: Damianos, Dimitrios, et al.
Veröffentlicht: (2026)
von: Damianos, Dimitrios, et al.
Veröffentlicht: (2026)
Protein Circuit Tracing via Cross-layer Transcoders
von: Tsui, Darin, et al.
Veröffentlicht: (2026)
von: Tsui, Darin, et al.
Veröffentlicht: (2026)
A Parametric Rate-Distortion Model for Video Transcoding
von: Jamali, Maedeh, et al.
Veröffentlicht: (2024)
von: Jamali, Maedeh, et al.
Veröffentlicht: (2024)
Shifting the Gradient: Understanding How Defensive Training Methods Protect Language Model Integrity
von: Grant, Satchel, et al.
Veröffentlicht: (2026)
von: Grant, Satchel, et al.
Veröffentlicht: (2026)
Scaling-Aware Adapter for Structure-Grounded LLM Reasoning
von: Jing, Zihao, et al.
Veröffentlicht: (2026)
von: Jing, Zihao, et al.
Veröffentlicht: (2026)
DifCluE: Generating Counterfactual Explanations with Diffusion Autoencoders and modal clustering
von: Jain, Suparshva, et al.
Veröffentlicht: (2025)
von: Jain, Suparshva, et al.
Veröffentlicht: (2025)
ScoNe: Benchmarking Negation Reasoning in Language Models With Fine-Tuning and In-Context Learning
von: She, Jingyuan Selena, et al.
Veröffentlicht: (2023)
von: She, Jingyuan Selena, et al.
Veröffentlicht: (2023)
Reviving Your MNEME: Predicting The Side Effects of LLM Unlearning and Fine-Tuning via Sparse Model Diffing
von: Kassem, Aly M., et al.
Veröffentlicht: (2025)
von: Kassem, Aly M., et al.
Veröffentlicht: (2025)
CLT-Forge: A Scalable Library for Cross-Layer Transcoders and Attribution Graphs
von: Draye, Florent, et al.
Veröffentlicht: (2026)
von: Draye, Florent, et al.
Veröffentlicht: (2026)
Demystifying Verbatim Memorization in Large Language Models
von: Huang, Jing, et al.
Veröffentlicht: (2024)
von: Huang, Jing, et al.
Veröffentlicht: (2024)
GIO: Gradient Information Optimization for Training Dataset Selection
von: Everaert, Dante, et al.
Veröffentlicht: (2023)
von: Everaert, Dante, et al.
Veröffentlicht: (2023)
Structural Priors and Modular Adapters in the Composable Fine-Tuning Algorithm of Large-Scale Models
von: Wang, Yuxiao, et al.
Veröffentlicht: (2025)
von: Wang, Yuxiao, et al.
Veröffentlicht: (2025)
Asymmetry in Low-Rank Adapters of Foundation Models
von: Zhu, Jiacheng, et al.
Veröffentlicht: (2024)
von: Zhu, Jiacheng, et al.
Veröffentlicht: (2024)
How does My Model Fail? Automatic Identification and Interpretation of Physical Plausibility Failure Modes with Matryoshka Transcoders
von: Tang, Yiming, et al.
Veröffentlicht: (2025)
von: Tang, Yiming, et al.
Veröffentlicht: (2025)
Do Language Models Use Their Depth Efficiently?
von: Csordás, Róbert, et al.
Veröffentlicht: (2025)
von: Csordás, Róbert, et al.
Veröffentlicht: (2025)
Dif4FF: Leveraging Multimodal Diffusion Models and Graph Neural Networks for Accurate New Fashion Product Performance Forecasting
von: Avogaro, Andrea, et al.
Veröffentlicht: (2024)
von: Avogaro, Andrea, et al.
Veröffentlicht: (2024)
ONNXPruner: ONNX-Based General Model Pruning Adapter
von: Ren, Dongdong, et al.
Veröffentlicht: (2024)
von: Ren, Dongdong, et al.
Veröffentlicht: (2024)
Federated Adapter on Foundation Models: An Out-Of-Distribution Approach
von: Yang, Yiyuan, et al.
Veröffentlicht: (2025)
von: Yang, Yiyuan, et al.
Veröffentlicht: (2025)
Adapter-Based Multi-Agent AVSR Extension for Pre-Trained ASR Models
von: Simic, Christopher, et al.
Veröffentlicht: (2025)
von: Simic, Christopher, et al.
Veröffentlicht: (2025)
Rapid Switching and Multi-Adapter Fusion via Sparse High Rank Adapters
von: Bhardwaj, Kartikeya, et al.
Veröffentlicht: (2024)
von: Bhardwaj, Kartikeya, et al.
Veröffentlicht: (2024)
Hadamard Adapter: An Extreme Parameter-Efficient Adapter Tuning Method for Pre-trained Language Models
von: Chen, Yuyan, et al.
Veröffentlicht: (2024)
von: Chen, Yuyan, et al.
Veröffentlicht: (2024)
HG-Adapter: Improving Pre-Trained Heterogeneous Graph Neural Networks with Dual Adapters
von: Mo, Yujie, et al.
Veröffentlicht: (2024)
von: Mo, Yujie, et al.
Veröffentlicht: (2024)
Recursive Models for Long-Horizon Reasoning
von: Yang, Chenxiao, et al.
Veröffentlicht: (2026)
von: Yang, Chenxiao, et al.
Veröffentlicht: (2026)
ELLA: Efficient Lifelong Learning for Adapters in Large Language Models
von: Biswas, Shristi Das, et al.
Veröffentlicht: (2026)
von: Biswas, Shristi Das, et al.
Veröffentlicht: (2026)
Improving Robustness of Foundation Models in Domain Adaptation with Soup-Adapters
von: Roschkowski, Marco
Veröffentlicht: (2025)
von: Roschkowski, Marco
Veröffentlicht: (2025)
Ähnliche Einträge
-
Internal Causal Mechanisms Robustly Predict Language Model Out-of-Distribution Behaviors
von: Huang, Jing, et al.
Veröffentlicht: (2025) -
Simple LLM Baselines are Competitive for Model Diffing
von: Kempf, Elias, et al.
Veröffentlicht: (2026) -
Reasoning-Finetuning Repurposes Latent Representations in Base Models
von: Ward, Jake, et al.
Veröffentlicht: (2025) -
Reading the Finetuning Prior: Verbatim Content Recovery via Contrastive Decoding Diffing
von: Brzozowski, Michał, et al.
Veröffentlicht: (2026) -
Delta-Crosscoder: Robust Crosscoder Model Diffing in Narrow Fine-Tuning Regimes
von: Kassem, Aly, et al.
Veröffentlicht: (2026)