ReFT: Representation Finetuning for Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Wu, Zhengxuan, Arora, Aryaman, Wang, Zheng, Geiger, Atticus, Jurafsky, Dan, Manning, Christopher D., Potts, Christopher |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
di: Wu, Zhengxuan, et al.
Pubblicazione: (2025)
di: Wu, Zhengxuan, et al.
Pubblicazione: (2025)
PreFT: Prefill-only finetuning for efficient inference
di: Lanpouthakoun, Andrew, et al.
Pubblicazione: (2026)
di: Lanpouthakoun, Andrew, et al.
Pubblicazione: (2026)
A Reply to Makelov et al. (2023)'s "Interpretability Illusion" Arguments
di: Wu, Zhengxuan, et al.
Pubblicazione: (2024)
di: Wu, Zhengxuan, et al.
Pubblicazione: (2024)
Bayesian scaling laws for in-context learning
di: Arora, Aryaman, et al.
Pubblicazione: (2024)
di: Arora, Aryaman, et al.
Pubblicazione: (2024)
pyvene: A Library for Understanding and Improving PyTorch Models via Interventions
di: Wu, Zhengxuan, et al.
Pubblicazione: (2024)
di: Wu, Zhengxuan, et al.
Pubblicazione: (2024)
HyperSteer: Activation Steering at Scale with Hypernetworks
di: Sun, Jiuding, et al.
Pubblicazione: (2025)
di: Sun, Jiuding, et al.
Pubblicazione: (2025)
RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations
di: Huang, Jing, et al.
Pubblicazione: (2024)
di: Huang, Jing, et al.
Pubblicazione: (2024)
CausalGym: Benchmarking causal interpretability methods on linguistic tasks
di: Arora, Aryaman, et al.
Pubblicazione: (2024)
di: Arora, Aryaman, et al.
Pubblicazione: (2024)
Recurrent Neural Networks Learn to Store and Generate Sequences using Non-Linear Representations
di: Csordás, Róbert, et al.
Pubblicazione: (2024)
di: Csordás, Róbert, et al.
Pubblicazione: (2024)
Nested-ReFT: Efficient Reinforcement Learning for Large Language Model Fine-Tuning via Off-Policy Rollouts
di: Heuillet, Maxime, et al.
Pubblicazione: (2025)
di: Heuillet, Maxime, et al.
Pubblicazione: (2025)
Improved Representation Steering for Language Models
di: Wu, Zhengxuan, et al.
Pubblicazione: (2025)
di: Wu, Zhengxuan, et al.
Pubblicazione: (2025)
MrT5: Dynamic Token Merging for Efficient Byte-level Language Models
di: Kallini, Julie, et al.
Pubblicazione: (2024)
di: Kallini, Julie, et al.
Pubblicazione: (2024)
HyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks
di: Sun, Jiuding, et al.
Pubblicazione: (2025)
di: Sun, Jiuding, et al.
Pubblicazione: (2025)
How Do Transformers Learn Variable Binding in Symbolic Programs?
di: Wu, Yiwei, et al.
Pubblicazione: (2025)
di: Wu, Yiwei, et al.
Pubblicazione: (2025)
Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations
di: Geiger, Atticus, et al.
Pubblicazione: (2023)
di: Geiger, Atticus, et al.
Pubblicazione: (2023)
How Causal Abstraction Underpins Computational Explanation
di: Geiger, Atticus, et al.
Pubblicazione: (2025)
di: Geiger, Atticus, et al.
Pubblicazione: (2025)
Mechanistic evaluation of Transformers and state space models
di: Arora, Aryaman, et al.
Pubblicazione: (2025)
di: Arora, Aryaman, et al.
Pubblicazione: (2025)
Do Language Models Use Their Depth Efficiently?
di: Csordás, Róbert, et al.
Pubblicazione: (2025)
di: Csordás, Róbert, et al.
Pubblicazione: (2025)
Language Model Circuits Are Sparse in the Neuron Basis
di: Arora, Aryaman, et al.
Pubblicazione: (2026)
di: Arora, Aryaman, et al.
Pubblicazione: (2026)
ScoNe: Benchmarking Negation Reasoning in Language Models With Fine-Tuning and In-Context Learning
di: She, Jingyuan Selena, et al.
Pubblicazione: (2023)
di: She, Jingyuan Selena, et al.
Pubblicazione: (2023)
ReCOGS: How Incidental Details of a Logical Form Overshadow an Evaluation of Semantic Interpretation
di: Wu, Zhengxuan, et al.
Pubblicazione: (2023)
di: Wu, Zhengxuan, et al.
Pubblicazione: (2023)
Mission: Impossible Language Models
di: Kallini, Julie, et al.
Pubblicazione: (2024)
di: Kallini, Julie, et al.
Pubblicazione: (2024)
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
di: Rafailov, Rafael, et al.
Pubblicazione: (2023)
di: Rafailov, Rafael, et al.
Pubblicazione: (2023)
Internal Causal Mechanisms Robustly Predict Language Model Out-of-Distribution Behaviors
di: Huang, Jing, et al.
Pubblicazione: (2025)
di: Huang, Jing, et al.
Pubblicazione: (2025)
Large Language Models to Diffusion Finetuning
di: Cetin, Edoardo, et al.
Pubblicazione: (2025)
di: Cetin, Edoardo, et al.
Pubblicazione: (2025)
Reasoning Theater: Disentangling Model Beliefs from Chain-of-Thought
di: Boppana, Siddharth, et al.
Pubblicazione: (2026)
di: Boppana, Siddharth, et al.
Pubblicazione: (2026)
Delta Activations: A Representation for Finetuned Large Language Models
di: Xu, Zhiqiu, et al.
Pubblicazione: (2025)
di: Xu, Zhiqiu, et al.
Pubblicazione: (2025)
Combining Causal Models for More Accurate Abstractions of Neural Networks
di: Pîslar, Theodora-Mara, et al.
Pubblicazione: (2025)
di: Pîslar, Theodora-Mara, et al.
Pubblicazione: (2025)
Transfer Learning of Tabular Data by Finetuning Large Language Models
di: Rabbani, Shourav B., et al.
Pubblicazione: (2025)
di: Rabbani, Shourav B., et al.
Pubblicazione: (2025)
Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability
di: Geiger, Atticus, et al.
Pubblicazione: (2023)
di: Geiger, Atticus, et al.
Pubblicazione: (2023)
Vanishing Gradients in Reinforcement Finetuning of Language Models
di: Razin, Noam, et al.
Pubblicazione: (2023)
di: Razin, Noam, et al.
Pubblicazione: (2023)
Evaluating Open-Source Sparse Autoencoders on Disentangling Factual Knowledge in GPT-2 Small
di: Chaudhary, Maheep, et al.
Pubblicazione: (2024)
di: Chaudhary, Maheep, et al.
Pubblicazione: (2024)
Diffusion Language Models Can Perform Many Tasks with Scaling and Instruction-Finetuning
di: Ye, Jiasheng, et al.
Pubblicazione: (2023)
di: Ye, Jiasheng, et al.
Pubblicazione: (2023)
Fine-Tuning and Prompt Optimization: Two Great Steps that Work Better Together
di: Soylu, Dilara, et al.
Pubblicazione: (2024)
di: Soylu, Dilara, et al.
Pubblicazione: (2024)
LoLCATs: On Low-Rank Linearizing of Large Language Models
di: Zhang, Michael, et al.
Pubblicazione: (2024)
di: Zhang, Michael, et al.
Pubblicazione: (2024)
Updating CLIP to Prefer Descriptions Over Captions
di: Zur, Amir, et al.
Pubblicazione: (2024)
di: Zur, Amir, et al.
Pubblicazione: (2024)
In-Context Sharpness as Alerts: An Inner Representation Perspective for Hallucination Mitigation
di: Chen, Shiqi, et al.
Pubblicazione: (2024)
di: Chen, Shiqi, et al.
Pubblicazione: (2024)
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model
di: Liang, Jing, et al.
Pubblicazione: (2025)
di: Liang, Jing, et al.
Pubblicazione: (2025)
Cooking Up Creativity: Enhancing LLM Creativity through Structured Recombination
di: Mizrahi, Moran, et al.
Pubblicazione: (2025)
di: Mizrahi, Moran, et al.
Pubblicazione: (2025)
Unfamiliar Finetuning Examples Control How Language Models Hallucinate
di: Kang, Katie, et al.
Pubblicazione: (2024)
di: Kang, Katie, et al.
Pubblicazione: (2024)
Documenti analoghi
-
AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
di: Wu, Zhengxuan, et al.
Pubblicazione: (2025) -
PreFT: Prefill-only finetuning for efficient inference
di: Lanpouthakoun, Andrew, et al.
Pubblicazione: (2026) -
A Reply to Makelov et al. (2023)'s "Interpretability Illusion" Arguments
di: Wu, Zhengxuan, et al.
Pubblicazione: (2024) -
Bayesian scaling laws for in-context learning
di: Arora, Aryaman, et al.
Pubblicazione: (2024) -
pyvene: A Library for Understanding and Improving PyTorch Models via Interventions
di: Wu, Zhengxuan, et al.
Pubblicazione: (2024)